0% found this document useful (0 votes)
27 views5 pages

Data Science Homework 2: Perceptron & Neural Networks

Uploaded by

xunzhang
Copyright
© All Rights Reserved
We take content rights seriously. If you suspect this is your content, claim it here.
Available Formats
Download as PDF, TXT or read online on Scribd
0% found this document useful (0 votes)
27 views5 pages

Data Science Homework 2: Perceptron & Neural Networks

Uploaded by

xunzhang
Copyright
© All Rights Reserved
We take content rights seriously. If you suspect this is your content, claim it here.
Available Formats
Download as PDF, TXT or read online on Scribd

CS M148: Introduction to Data Science Submit to Gradescope

Released: Feb. 28th, 2025 Homework 2 Due: March 10th, 2PM PST

1 Perceptron [15 points]


Consider training a Perceptron model y = sgn(w⊤ x), w ∈ Rd on a dataset D = {(xi , yi )}, i = 1 . . . 5.
Both w and xi are vectors of dimension d, and y ∈ {+1, −1} is binary. Assume that the bias term
is already augmented in xi : xi = [1, x1 , . . . , xd−1 ]. The activation function is a sign function where
sgn(x) = 1 for all x > 0 and sgn(x) = −1 otherwise. The Perceptron algorithm is given below,

Algorithm 1 Perceptron
Initialize w = 0
for i = 1 . . . N do
if yi ̸= sgn(w⊤ xi ) then
w ← w + yi xi
end if
end for
return w

(a) (5 points) The Perceptron model is trained for 1 epoch, i.e. iterated over the entire dataset
once, and the weight vector after this training epoch, w, is expressed as w = y1 x1 + y2 x2 +
y4 x4 + y5 x5 . What can we infer about the training process from this?
Solution: x1 , x2 , x4 , x5 are all misclassified.

(b) (5 points) Let d = 3 and the data points be given as follows

i xi,1 xi,2 xi,3 yi


1 1 1 0 +1
2 1 2 -1 +1
3 1 2 -3 -1
4 1 3 -1 +1
5 1 1 -1 +1

Following the formulation of your answer in (a), what is of w given the values of the data
points? Express with a vector of numbers this time. Furthermore, if we iterate through the
dataset again, what will be the model’s prediction on x2 ?
Solution: w = [4, 7, −3]. It will not make mistake again and will predict +1.

(c) (5 points) How does the Perceptron handle cases where the data is not linearly separable?
How does this compare to logistic regression?
Solution: The Perceptron cannot handle non-linearly separable data because it only updates
weights based on misclassified examples and does not minimize any continuous loss. If the
data is not separable, it will oscillate or fail to converge. Logistic regression uses probabilistic
approach, so the loss can converge even when the data is not separable.
.
CS M148: Introduction to Data Science Submit to Gradescope
Released: Feb. 28th, 2025 Homework 2 Due: March 10th, 2PM PST

2 Neural Networks [15 points]


(a) (4 points) Refer to Lecture 13 for the activation functions of neural networks. Considering a
binary classification problem, what are possible activations choices for the hidden and output
layers repectively? Explain why.
Solution: Hidden layers: Any activations are ok. ReLU, Leaky ReLU, ELU and variants
are popular choices. Activation functions in the hidden layers are meant to make our model
sparse and address the gradient vanish or exploding issues.
Output layers: sigmoid, softmax, or tanh activations to generate probabilities of the input
being in each class. We need specific activation functions to map the output of our neural
network to the desired format.
(b) (3 points) Consider a binary classification problem where y ∈ {0, 1}. We consider the neural
network in Figure 2 with 2 inputs, 2 hidden neurons, and 1 output. We let neuron 1 use
Sigmoid activation and neuron 2 use ReLU activation, respectively. The other layers use
the linear activation function. Suppose we have an input X1 = 3.1 and X2 = −9.8, with
label y = 0. And the weights are initialized as W11 = −0.8, W12 = −0.1, W21 = 3.8, W22 =
0.8, W31 = −2, W32 = 0.2, and the bias term W10 , W20 , W30 are all initialized to be 0. Compute
the output of the network rounded to two decimal places.
Solution:
z1 = W1T X = W11 X1 + W12 X2 + W10 = 1.2 × 3.1 + 0.3 × (−9.8) + 0 = −2.48 + 0.98 = −1.5
h1 = Sigmoid(z1 ) = 0.182425
z2 = W2T X = W21 X2 + W22 X2 + W20 = (3.8) × 3.1 + (0.8) × (−9.8) + 0 = 11.78 − 7.84 = 3.94
h2 = ReLU(z2 ) = 3.94
ŷ = W31 h1 + W32 h2 + W30 = (−2) × 0.182425 + 0.2 × 3.94 = −0.36485 + 0.788 ≈ 0.42.
(c) (4 points) We consider the binary cross entropy loss function. What is the loss of the
network on the given data point in (b)? What is ∂L∂ ŷ ? Report both rounded to three decimal
places. (Hint. Refer to the lecture slides for defining a binary cross-entropy loss function.
For simplicity, use your rounded answer from part b.)
Solution:
L = −y ln(ŷ) − (1 − y) ln(1 − ŷ) = −1 × ln(1 − 0.42) = −1 × ln(0.58) = 0.544.
∂L ∂ y 1−y 1
∂ ŷ = ∂ ŷ (−y ln(ŷ) − (1 − y) ln(1 − ŷ)) = − ŷ + 1−ŷ = 0.58 = 1.724.

(d) (6 points) We now consider the backward pass. Given the same initialized weights and input
as in (b), write the formula and calculate the derivative of the loss w.r.t W12 to the nearest
∂L
decimal point, i.e. ∂W 12
. Should the bias at Neuron 1 be increased or decreased?
Solution:
∂L ∂L ∂ ŷ ∂L ∂ ŷ ∂h1 ∂z1 ∂L
∂W12 = ∂ ŷ ∂W12 = ∂ ŷ ∂h1 ∂z1 ∂W12 = ∂ ŷ W31 · σ(z1 ) · (1 − σ(z1 )) · X2 = 1.724 × −2 × 0.149146 ×
−9.8 = 5.0.
The gradient of the bias can be calculated as:
∂L ∂L ∂ ŷ ∂L ∂ ŷ ∂h1 ∂z1 ∂L
∂W10 = ∂ ŷ ∂W10 = ∂ ŷ ∂h1 ∂z1 ∂W10 = ∂ ŷ W31 ·σ(z1 )·(1−σ(z1 ))·1 = 1.724×−2×0.149146×1 =
−0.514.
∂L
We can save time by reusing the intermediate values we got while calculating ∂W 12
.
The gradient at the bias is negative, so we should increase it.
CS M148: Introduction to Data Science Submit to Gradescope
Released: Feb. 28th, 2025 Homework 2 Due: March 10th, 2PM PST

Figure 1: Neural Network

(e) (4 points) Given the neural network as in (b), how many parameters does the network have?
(Hint. Each weight unit counts as a parameter, and we also consider the bias terms (W10 , . . .)
as parameters.)
Solution: 3 × 2 + 3 = 9 parameters
CS M148: Introduction to Data Science Submit to Gradescope
Released: Feb. 28th, 2025 Homework 2 Due: March 10th, 2PM PST

3 Multi-class Classification [10 points]

Figure 2: Multiclass Logistic Regression

(a) (5 points) Consider a multi-class classification problem with 4 classes and 25 features. We
will use the logistic regression model to first build our binary classifier. Then, what will be
the total number of parameters for using One vs. Rest (ovr) strategies for the multi-class
classification task using logistic regression? (Hint. Refer to lecture 11 for multi-class logistic
regression model.)
Solution: For each classes we need to train a classification model, and each model will have
25 + 1 parameters, so in total is 4 × 26 = 104

(b) (5 points) Consider a new multi-class classification problem with 3 classes. The distribution
of the points is shown in the figure (Square - Class 1, Circle - Class 2, Star - Class 3). Draw
the linear classifiers used for classifying the three classes, using (i) One vs. Rest (OvR), and
(ii) Multinomial approaches.
Solution: 3 lines for OvR where each line separates a class with all others. 2 lines for
Multinomial, where each line separates a class with the reference class; any reference class is
fine.
CS M148: Introduction to Data Science Submit to Gradescope
Released: Feb. 28th, 2025 Homework 2 Due: March 10th, 2PM PST

4 Decision Boundary [10 points]


Consider the classification problems with two classes, which are illustrated by circles and crosses in
the plots below. In each of the plots, one of the following classification methods has been used, and
the resulting decision boundary is shown:

(1) Neural Network (1 hidden layer with 10 ReLU)


Solution: (b) It should be piecewise linear due to the ReLU activiation function

(2) Neural Network (1 hidden layer with 10 tanh units)


Solution: (a) It should be curved due to the tanh activiation function

Assign each of the previous methods to exactly one of the following plots (in a one to one corre-
spondence) by annotating the plots with the respective letters, and explain briefly why did you
make each assignment.

(a) (b)

Common questions

Powered by AI

ReLU and its variants are popular for hidden layers because they contribute to sparsity in the network and address issues like the vanishing gradient problem, enhancing the training efficiency. For binary classification output layers, sigmoid, softmax, or tanh are suitable choices, as they convert the output into probabilities, aligning with the binary classification requirements .

To compute the gradient of the binary cross-entropy loss concerning a specific weight, one multiplies the gradient of the loss with respect to the predicted output, the gradient of the output with respect to the activation, the gradient of the activation with respect to the pre-activation value (z1), and finally the gradient of z1 with respect to the weight. For example, the gradient with respect to W12 is calculated using these chain rule derivatives, resulting in a value such as 5.0 .

With a neural network using ReLU activation in the hidden layer, the decision boundary is piecewise linear due to the linearity of the ReLU function. In contrast, a network employing tanh units would have a curved decision boundary, reflecting the non-linear nature of the tanh function .

When a Perceptron training epoch results in a weight vector expressed as w = y1x1 + y2x2 + y4x4 + y5x5, it indicates that data points x1, x2, x4, and x5 were misclassified during the training. This is because the Perceptron only updates its weights for misclassified instances .

The total number of parameters in a neural network is calculated by counting each weight and bias term. For a neural network with a structure as described, having weights like W11 and W12, and biases like W10, the total parameters include both weights and biases. In the example network from the document, there are 3 biases and weights contributing to the term count, totaling 9 parameters .

For a One vs. Rest (OvR) logistic regression model with 4 classes and 25 features, each class requires a separate binary classifier with 26 parameters (25 features + 1 bias term). Therefore, the total number of parameters is 4 × 26 = 104 parameters .

The Perceptron algorithm struggles with non-linearly separable data, as it does not converge and may oscillate due to lack of a continuous loss function. In contrast, logistic regression, which uses a probabilistic approach, can still converge by minimizing a differentiable loss function, even if the data is not linearly separable .

Given a neural network with specified weights and inputs X1 = 3.1 and X2 = -9.8, with activations such as Sigmoid and ReLU, you calculate each layer's pre-activation value, apply the activation function, and feed forward to the next layer. For instance, with specified weights and biases, the output was calculated as approximately 0.42 .

A bias term in a neural network should be adjusted based on the sign of its gradient. If the gradient of the bias is negative, it indicates that increasing the bias could reduce the loss, thus the bias should be increased. In the provided case, since the gradient was negative, increasing the bias was suggested .

A linearly classifying model like the Perceptron might oscillate and fail to converge in scenarios where the data is not linearly separable. This happens because the Perceptron model does not have a mechanism like a loss function to manage and settle on optimal weights, which can lead to continuous misclassification and weight updates based on differing subsets of the data .

You might also like