0% found this document useful (0 votes)
3 views5 pages

ProblemSet 1 Solutions

The document outlines an exercise on multiple linear regression analysis using a given data table. It includes tasks such as writing the regression model in matrix form, filling in missing values from OLS output, calculating the variance of error terms, discussing the model's goodness-of-fit, and evaluating the significance of independent variables x1 and x2. The analysis concludes that x1 does not significantly influence y, while x2 does, and it also assesses the overall significance of the regression model using an F-test.

Uploaded by

Mathilde
Copyright
© All Rights Reserved
We take content rights seriously. If you suspect this is your content, claim it here.
Available Formats
Download as PDF, TXT or read online on Scribd
0% found this document useful (0 votes)
3 views5 pages

ProblemSet 1 Solutions

The document outlines an exercise on multiple linear regression analysis using a given data table. It includes tasks such as writing the regression model in matrix form, filling in missing values from OLS output, calculating the variance of error terms, discussing the model's goodness-of-fit, and evaluating the significance of independent variables x1 and x2. The analysis concludes that x1 does not significantly influence y, while x2 does, and it also assesses the overall significance of the regression model using an F-test.

Uploaded by

Mathilde
Copyright
© All Rights Reserved
We take content rights seriously. If you suspect this is your content, claim it here.
Available Formats
Download as PDF, TXT or read online on Scribd

Big Data Analysis

PGE L3
Academic Year 2023-2024
Exercise 1.

Given the following data table

y 0 24 12 8 12 16
x1 -2 -1 0 0 1 2
x2 -5 4 0 -2 2 1

consider the multiple linear regression model

𝑦! = 𝛽" + 𝛽# 𝑥#! + 𝛽$ 𝑥$! + 𝜀! , 𝑖 = 1,2, … ,6,

where OLS assumptions are satisfied.

1) Write down the model in matrix form and describe all the vectors and matrices.

You can answer this by studying slide 34 in the slide deck from lecture 1:
𝑌 = 𝑋𝛽 + 𝜀,
where
• 𝑌 is a 6 × 1 matrix representing the values in data for the dependent variable,
• 𝑋 is a 6 × 3 matrix with all 1s in the first column for the intercept and the values in data
for the independent variables 𝑥# and 𝑥$ ,
• 𝛽 is a 3× 1 matrix representing the coefficients, and
• 𝜀 is a 6× 1 matrix representing the error terms
such that, when 6 = 𝑛, we have
3 2
<latexit sha1_base64="c/WD1gcAkivKm4b7Gd+KyMmZX68=">AAACLXicdVDLSsNAFJ34Nr6qLt0MFsFVSaK2dSGIunBZwfqgCWUyua2Dk0mYmRRL8Ifc+CsiuFDErb/hpK34QA8M93DuY+49YcqZ0o7zbI2NT0xOTc/M2nPzC4tLpeWVM5VkkkKTJjyRFyFRwJmApmaaw0UqgcQhh/Pw+rDIn/dAKpaIU91PIYhJV7AOo0QbqV068lv2Jd7DfghdJvIwJlqym1u733ax75vgFcHvRYlWQ0HYPojoq9IP7Hap7FR261Vvx8NOxXFq3la1IF5t29vCrlEKlNEIjXbp0Y8SmsUgNOVEqZbrpDrIidSMcjAzMwUpodekCy1DBYlBBfng2lu8YZQIdxJpntB4oH7vyEmsVD8OTaXZ8Ur9zhXiX7lWpjv1IGcizTQIOvyok3GsE1xYhyMmgWreN4RQycyumF4RSag2BhcmfF6K/ydnXsWtVnZOvPL+wciOGbSG1tEmclEN7aNj1EBNRNEdekDP6MW6t56sV+ttWDpmjXpW0Q9Y7x9b+qZ0</latexit>

2 <latexit sha1_base64="5l4AxJR8ChyX1hVNW3AV/rPL1mg=">AAACcHicdZFNT9wwEIadtFBI+djCBalFNaxAiMMqMbALh0qovXCkUhdW2kQrx5ldLBwnsh3EKtpz/x83fgQXfgHOfhSKYCRLr56Z8Yxfx7ng2vj+veN++Dg3/2lh0fu8tLyyWvuydqGzQjFos0xkqhNTDYJLaBtuBHRyBTSNBVzG17+q/OUNKM0z+ccMc4hSOpC8zxk1FvVqf8Ou18E/cBjDgMsyTqlR/HbkBXgX3/bKIBhNBLEiDP9hMsNkjMObJDPaotfiRYuctVjhhSCT52Fh5PVqdb9xctwkRwT7Dd9vkYNmJUjrkBzgwJIq6mga573aXZhkrEhBGiao1t3Az01UUmU4E2DvLDTklF3TAXStlDQFHZVjw0Z4x5IE9zNljzR4TF92lDTVepjGttLueKVf5yr4Vq5bmP5xVHKZFwYkmwzqFwKbDFfu44QrYEYMraBMcbsrZldUUWbsH1UmzF6K3xcXpBE0G0e/Sf3059SOBfQVbaM9FKAWOkVn6By1EUMPzrrzzdl0Ht0N97u7NSl1nWnPOvov3P0ndF21Xw==</latexit>

3 3
<latexit sha1_base64="9r5mfJdH7YBh9dh+SzbPLKl+gyc=">AAACVXicdVFdS8MwFE3r9/ya+uhLcAg+jbbqNh8E0RcfFdwmrGWk6d0MpklJ0uEo+5N7Ef+JL4KpmzBFDwQO59x7c3MSZ5xp43lvjru0vLK6tr5R2dza3tmt7u13tMwVhTaVXKrHmGjgTEDbMMPhMVNA0phDN36+Kf3uCJRmUjyYcQZRSoaCDRglxkr9Kg97lXBEFGSacSnwJQ5jGDJRxCkxir1MFt2+j8PwhxDMhEQa/dsSlRBEsjAnqvSrNa9+0WoE5wH26p7XDE4bJQmaZ8Ep9q1SoobmuOtXp2EiaZ6CMJQTrXu+l5moIMowysHOzDVkhD6TIfQsFSQFHRVfqUzwsVUSPJDKHmHwl7rYUZBU63Ea20q745P+7ZXiX14vN4NWVDCR5QYEnV00yDk2EpcR44QpoIaPLSFUMbsrpk9EEWrsR5QhfL8U/086Qd1v1M/vg9rV9TyOdXSIjtAJ8lETXaFbdIfaiKIpenccx3VenQ932V2dlbrOvOcA/YC7+wld4rPB</latexit>

2
y1 1 x11 x21 "1
6 y2 7 61 x12 x22 7 2 3 6 "2 7
6 7 6 7 <latexit sha1_base64="1nGGJdaw1H4JR81lmxmCnwDA2Fo=">AAACNHicdVDLSgMxFM34rOOr6tJNsAiuysxUW10IRTeCGwVrhc5QMultG8xkhiQjlsGPcuOHuBHBhSJu/QYztmIVPRByOOfe5N4TJpwp7TiP1sTk1PTMbGHOnl9YXFourqyeqziVFBo05rG8CIkCzgQ0NNMcLhIJJAo5NMPLw9xvXoFULBZnepBAEJGeYF1GiTZSu3jst2w/BE3wPjZ3j4ksjIiW7PpmqLcd7Psj6n5Tz/ZBdMZqA9tuF0tOeW+36u142Ck7Ts2rVHPi1ba9CnaNkqOERjhpF+/9TkzTCISmnCjVcp1EBxmRmlEO5tFUQULoJelBy1BBIlBB9rn0Dd40Sgd3Y2mO0PhTHe/ISKTUIApNpRmyr357ufiX10p1dzfImEhSDYIOP+qmHOsY5wniDpNANR8YQqhkZlZM+0QSqk3OeQhfm+L/yblXdqvlnVOvVD8YxVFA62gDbSEX1VAdHaET1EAU3aIH9IxerDvryXq13oalE9aoZw39gPX+AegLqTQ=</latexit>

6 7
Y = 6 . 7 X = 6. .. .. 7 0 "=6 . 7
4 .. 5 4 .. . . 5 =4 5 4 .. 5
1
yn , 1 x1n x2n , 2 , and "n .
2) The OLS estimation output from Excel is provided below. Fill in the missing values.

Regression Statistics
Multiple R 0,952
R2 0,906
Adjusted R2 0,844
Std Err 3,162
Observations 6

ANOVA
Significance
df SS MS F F
Regression 2 290 145 14,5 0,029
Residual 3 30 10
Total 5 320

Standard Upper
Coefficients Error t Stat P value Lower 95% 95%
Intercept 12 1,291 9,295 0,003 7,891 16,109
X1 -0,5 1,118 -0,447 0,685 -4,058 3,058
X2 2,5 0,500 5,000 0,015 0,909 4,091

The table could be completed by studying the slides: 46, 48, 54, 55 and 59 of the slide deck
from Lecture 1.

3) What is the value of 𝜎4 $ , the unbiased estimator of 𝜎 $ , the variance of 𝜀?

See slides 35 in the slide deck from Lecture 1.

%
1
𝜎4 $ = : 𝜖̂!$
𝑁 − (𝑘 + 1)
!&#
Thus, we have
𝜎4 $ = 30/(6 − (2 + 1)) = 10.

4) Discuss the goodness-of-fit (quality) of the model using R2.

The OLS output table from question 2 gives you the value of R2. However, you should be able to
calculate this also based on slide 46.

𝑆𝑆 𝑅𝑒𝑔𝑟𝑒𝑠𝑠𝑖𝑜𝑛 𝑆𝑆 𝑅𝑒𝑠𝑖𝑑𝑢𝑎𝑙 (𝐸𝑟𝑟𝑜𝑟)


𝑅$ = =1−
𝑆𝑆 𝑇𝑜𝑡𝑎𝑙 𝑆𝑆 𝑇𝑜𝑡𝑎𝑙
290 30
𝑅$ = =1− = 0.906
320 320

The (estimated) model explains 90.6% of the variance of Y. This model has a high explanatory
power.

5) Discuss whether x1 and x2 significantly influence y.

a) The question implicitly asks us to evaluate whether the coefficient for x1 (𝛽# ) is different from
zero. In order to evaluate this, we need to conduct a statistical test with following hypotheses:

H0: 𝛽# = 0
H1: 𝛽# ≠ 0

We can calculate the statistics for such a test based on slide 59:
−0.5
𝑡'()(!'(!* = = −0.447
1.118

And then make the judgement of whether this statistic (t) is large enough (in absolute value) to
conclude that 𝛽# is really different from zero.
However, there are simpler ways of doing this by looking at the OLS output from question 2.
- This output directly calculates the p-value of our t-statistic (which is also calculated in the
table, by the way). It is equal to 0.685. This means that if we reject the null hypothesis
and say that coefficient is different from zero, the probability that we will be making an
error is 0.685. This is tremendously high. With 95% significance, we can only reject null
hypotheses is p-values are less than 0.05.
- Another way of evaluating whether x1 affects y, is to look at the error bounds around the
coefficient which are also given in the table from question 2. We see that the error band
contains 0, which means that our coefficient (-0.5) cannot be statistically distinguished
from zero.
These two methods indicate that x1 does not influence y.

b) For the influence of x2, we need to study coefficient 𝛽$ . The same two methods indicate that
p-value of the relevant test is 0.015. This is a small probability. It is less than 0.05. Which means
that we have statistical evidence that the coefficient is different from zero. The error bands
confirm this, because the band does not contain 0 (meaning coefficient is statistically
distinguishable from zero). Therefore, we can conclude that x2 does affect y.

Note: By looking at the output from OLS regression we can also evaluate whether the intercept
of our regression line is different from zero (for this you will need to evaluate the p-value or error
bands corresponding to intercept coefficient in the regression output).
If the confidence interval for a coefficient contains zero, it implies that there is a possibility that the true value of the coefficient is
zero. In other words, there is no statistically significant evidence that the corresponding independent variable has an effect on the
dependent variable.
6) Using F-test, what can you say about the model’s significance?

We can calculate the F statistic of the regression by studying slide 55.


𝑀𝑆 𝑅𝑒𝑔𝑟𝑒𝑠𝑠𝑖𝑜𝑛 145
𝐹'()(!'(!* = = = 14.5
𝑀𝑆 𝑅𝑒𝑠𝑖𝑑𝑢𝑎𝑙 10

However, it is also given in the OLS table. The table also gives the p-value of this statistics. It is
labeled “Significance F”, and is equal to 0.029. This is smaller than 0.05, thus we conclude that
the estimated regression model as a whole is statistically significant. This means that the model
fits the data better than the (hypothetical) model with no predictor variables.

Exercise 2.

An OLS estimation using a quarterly dataset yields the following regression equation

𝑦4( = 2.2 + 0.104𝑥#( + 3.4𝑥$( + 0.34𝑥+( , 𝑡 = 1,2, … ,64

𝑠𝑒V𝛽W" X = 3.4 𝑠𝑒(𝛽W# ) = 0.0005 𝑠𝑒(𝛽W$ ) = 2.2 𝑠𝑒(𝛽W+ ) = 0.15


𝑆𝑆𝑅 = 112 𝑆𝑆𝐸 = 19.5,

where SSR stands for the regression sum of squares and SSE stands for the error sum of squares.

1) Which independent variables influence the dependent variable significantly?

Evaluating this is somewhat more challenging as we do not have p-values readily available to
us. However, we can calculate t-statistics as we have estimated coefficients and their standard
errors. Then we can use the rule of thumb (for the number of observations >30) that the critical
value of t-statistic is approximately equal to 2 (for 95% confidence).

𝑡 ∗ = 𝑡*-!(!*). = 𝑡/,12&3 – (78#) = 𝑡".";,<+" = 2.000

Then we can calculate t-statistics for appropriate tests (again, see slide 59) and compare each
value to 2. If the calculated statistic is bigger than 2 (in absolute value), we can reject the null
hypothesis (which is that given variable does not influence the dependent variable). Otherwise,
we cannot reject the null hypothesis.

For 𝛽" , 𝑡'()(!'(!* = 2.2/3.4 < 𝑡 ∗ . Thus, fail to reject H0.


For 𝛽# , 𝑡'()(!'(!* = 0.104/0.0005 > 𝑡 ∗ . Thus, reject H0.
For 𝛽$ , 𝑡'()(!'(!* = 3.4/2.2 < 𝑡 ∗ . Thus, fail to reject H0.
For 𝛽+ , 𝑡'()(!'(!* = 3.4/0.15 > 𝑡 ∗ . Thus, reject H0.

In conclusion, only 𝑥# and 𝑥# effect the dependent variable significantly.


2) Compute R2 and discuss the goodness-of-fit of the model.

𝑆𝑆 𝑅𝑒𝑔𝑟𝑒𝑠𝑠𝑖𝑜𝑛 𝑆𝑆𝑅 112


𝑅$ = = = = 0.852
𝑆𝑆 𝑇𝑜𝑡𝑎𝑙 𝑆𝑆𝑅 + 𝑆𝑆𝐸 112 + 19.5

The model explains 85.2% of variation in 𝑦, which is quite high.

3) Compute the adjusted R2, denoted with 𝑅"! , and discuss its comparison with R2.

𝑁 − 1
𝑅"! = 1 − '1 − 𝑅2 ( ) *
𝑁 − (𝑘 + 1)
Thus, we have 𝑅"! = 1 − (1 − 0.852)(63/60) = 0.845.
When comparing models with differing numbers of independent variables, we can use 𝑅"! but it
is wrong to say that the model explains 84.5% of the variation in 𝑦. That is given by 𝑅 $ .

4) Implement the F-test and conclude.

𝑀𝑆 𝑅𝑒𝑔𝑟𝑒𝑠𝑠𝑖𝑜𝑛 𝑅$ 𝑁 − (𝑘 + 1) 0.852 60
𝐹'()(!'(!* = = $
[ \= [ \ = 115.13
𝑀𝑆 𝐸𝑟𝑟𝑜𝑟 1 − 𝑅 𝑘 0.148 3

This is a very large number (critical value of F statistic with more than 10 observations is not
higher than 5), so we reject the null hypothesis that the model is not significant.

p-value < alpha : we reject the null hypothesis


F statistic < F critical : we fail to reject the null hypothesis

You might also like