DATA ANALYSIS FOR ECONOMICS: PS3
PROBLEM SET 3: HYPOTHESIS TESTING
1 We have a sample of 45 workers employed in a company. We ask to each worker to
evaluate her/his satisfaction level at work (𝑥) from 0 to 10. We also know, for each worker,
the number of labour absenteeism days (𝑦) last year. A linear regression line is estimated such
Assuming all assumptions hold, on
that average when the satisfaction level is 0,
the number of absentisms is 12.6 days.
When the satisfection level improves by
The satisfaction one day, the number of absentisms
level explains 𝑦̂𝑖 = 12.6 − 1.2𝑥𝑖 decreases by 1.2 days on average.
32.1% of
absentism days (0.112) (0.088) H0= b1=0
H1= b1 diff 0
2
𝑅 = 0.321
t=1.2/0.088=13.64
13.64>2.69, therefore we reject the null
hypothesis.
There is significant evidence to prove that
a- Interpret the estimated regression model and the value of the R-squared satisfaction influences absentism
b- Test the null hypothesis that work satisfaction does not produce any significant effect
on labour absenteeism at a 1% significance level.
c- The level of work satisfaction of a different worker is 6. Find the predicted labour
absenteeism days per year for this worker. 5,4 days
d- In your opinion, explain one application of the above model from the perspective of
They can use this to evaluate the overall
the Human Resources department of the company. satisfaction level, based on how much
absentism is done.
2 Consider a SLRM relating the annual number of crimes on college campuses (crime)
to student enrollment (enroll) with the following estimation results:
̂ 𝑖 = −6.63 + 1.27log(𝑒𝑛𝑟𝑜𝑙𝑙)𝑖
log(𝑐𝑟𝑖𝑚𝑒) 𝑛 = 97 𝑅2 = 0.585
H0= b1=0
H1= b1diff 0
(1.03) (0.11)
t=1.27/0.11= 11.55 Considering all assumptions hold, when students
vs = 2.63 enrolled increase by 1%, on average the crimes
Therefore we reject H0, a- Interpret the estimated slope coefficient. on campuses increase by 1.27%
there is sufficent level to
claim that enroll has effect b- Calculate two-tailed test to find whether the variable enroll should be included in the
on crime with 99%
certainty. regression model (at 1% significance level).
c) vs t = (1.27 - 0) / 0.11
t = 11.55 c- Test whether the model is useful at 5% significance level.
once again we reject H0
3 Consider the model: 𝑦𝑖 = 𝛽0 + 𝛽1 𝑥1𝑖 + 𝛽2 𝑥2𝑖 + 𝛽3 𝑥3𝑖 + 𝛽4 𝑥4𝑖 + 𝛽5 𝑥5𝑖 + 𝑢𝑖
Explain how you would test the following null hypothesis:
a- 𝛽2 = 0 t test
b- 𝛽3 = 𝛽4 = 𝛽5 = 0 f test
c- 𝛽1 = 𝛽2 = 𝛽3 = 𝛽4 = 𝛽5 = 0 global t test
1
DATA ANALYSIS FOR ECONOMICS: PS3
4 A consultancy firm has been commissioned by a property investment company to
develop a model that will help their managers assess the value of real estate in the London
area. They have information on the following variables: price (property price in pounds);
floorm2 (size of the property in square meters); dholborn (distance in kilometers from the
property to the city center – Holborn tube station); age (age of the property in years) and
buyage (age of the buyer). The estimation results are given in the next Table.
OLS estimates - Dependent variable: log(price)
Model 1 Model 2 Model 3
const 10.89 11.68 11.53
(0.036) (0.056) (0.072)
floorm2 0.012 0.0104 0.009
(0.0005) (0.0005) (0.0005)
log(dholborn) -0.291 -0.251
(0.017) (0.018)
age 0.002
(0.0002)
buyage -0.0007
a) if all assumptions are met, the expected
(0.001) value of the house will increase by 1.2%
b) assuming all assumptions hold, and
for every expra square meter, on average.
keeping all else fixed, when the distance
from the center increases by 1%, the
n 1199 1199 1199
2
0.314 0.451 0.472 To test if it is statistically significant at
expected price will decreace by 0.291%. R 1%, we see that
use t test because testing indipendent SSR 147.188 117.803 113.431
t=0.012/0.0005=24
significance
t=-0.291/0.017=-17.12
Note: Standard errors in parentheses.
therefore, we have enough proof that it is
vs the 2.58 Conduct all hypothesis tests at 1 % significance level statistically significant at a 99%
confidence level
We reject H0
e) f test needed
h0= b3=0, b4=0 a- Interpret the OLS slope coefficient of the SLRM. Is size statistically significant?
h1=not H0
b- Interpret the estimated coefficient on the variable log(dholborn) in Model 2. Is
unresrticted= 11.53
+0.009(floorm2)-0.251(log(dh distance statistically significant?
olborn))
+0.002(age)-0.0007(buyage) c- Predict the price of a 120 square meter property that is 7 km away from the Holborn
+u
log P= 11.68+0,0104*120-0,281*ln(7)= 12.37
station using Model 2. exp (12.37)= 236026,65 pounds
restricted=11.53
+0.009(floorm2)-0.251(log(dh
olborn))+u
d- How does the effect of size on property prices change respect to the estimation result
It slightly decreases, this could be because we add the
in the SLRM (section a)? Why? log(dholborn) variable, so now not everything is only explained by
f=((117.803-113.431)/2)/ the first variable- it lowers its importance. This could be because
(113.431/(1199-4-1))=
e- Test whether age and buyage are jointly significant. the two variables are correlated
23.01 - most likely the 4 assumtpion is violated, hence Bsize is biased
f- Test the overall significance of Model 3. and contains error terms
fscore=4.61
g- If you were part of the team in the consultancy firm and had to choose a model,
Therefore we reject H0, which
means there is sufficient which one would you choose and why? f) Global test
evidence to claim that age and
buyage are jointly significant at h0=b1,b2,b3,b4=0
a 1% significance level h1 not h0
unrestricted: 11.53 +0.009(floorm2)-0.251(log(dholborn))+0.002(age)-0.0007(buyage)+u
restricted= 11.53 + u
F=((0.472/4))/((1-0.472)/(1199-4-1))= 266.84
we reject H0, there is enough evidence to claim the model as a whole is significant
2
DATA ANALYSIS FOR ECONOMICS: PS3
5 We have information about families below poverty level (POVRATE=percentage of
families with income below the poverty level) in a specific year for 58 counties in California
combined with information about potential determinants: UNEMP (percentage of
unemployment rate), FAMSIZE (persons per household), EDU (percent that completed
four years of college or higher), URBAN (percentage of urban population). Estimation
results are presented in the following table:
OLS Estimation Results - Dependent variable: POVRATE
Variable Model 1 Model 2 Model 3
Constant 2.637 1.906 4.309
(0.987) (4.292) (4.535)
Unemp 0.731 0.721 0.424
(0.092) (0.106) (0.166)
Famsize 0.305 2.388
(1.742) (1.871)
Edu -0.177
(0.081)
Urban -0.051
b) individual of famsize
(0.022)
0.305/1.742= 0.17 n 58 58 58
we accept h0
Adjusted R squared 0.518 0.510 0.548
There is not enough evidence to when unemployment increases by 1%
prove that SSR 421.692 421.457 374.675 pint, on avaerage it is assumed that the
Size of the fam influences poverty poverty rate will increase by 0.731
rate Note: Standard errors are in parenthesis. percentage points
Conduct all the hypothesis tests at 1% significance level.
GLOBAL t= (0.731)/(0.092)=7.94
F = 18.96 we reject H0, there is enough evidence
Therefore we reject H0, which to prove unemployment effects poverty
means that the model overall has rate at 99% certainty
enough evidence to prove it
influences poverty rate, which a- Interpret the slope coefficient in Model 1 and test its individual significance. c) we saw in the
means unemp and famsize are previous question that
jointly significant b- Test the individual and global significance in Model 2. famsize is very likely
not influencing the rate
c- Comment on the effect of FAMSIZE on POVRATE in the second model. Why do of poverty, hence
showing the
d) we can compere the
you think is a positive and insignificant effect? Does this effect affect the goodness insignificant changes.
restricted and unrestricted We see that adjusted R
model—> after computing f
of fit of model 2 if compared with model 1? Why? squared is slightly
test then we see that we lower in model 2,
reject H0, meaning they do
d- In Model 3 we add two new explanatory variables: EDU and URBAN. Test whether wadding FAMSIZE to
have effect at a 99% the model does not
confidence level.
this inclusion helps to improve the quality of the model. Is model 3 the best in terms improve the
explanatory power of
of goodness-of-fit? Why? the model, and that it
may even introduce
e- Are the effects of these two new variables the expected ones? some noise or
multicollinearity.
f- What about the individual significance of UNEMP in model 3 if compared with the difference in the
adjusted R-squared is
model 2? Explain. very small, and it may
not be statistically
significant. Therefore,
f) If we compare this result with model 2, we can see that the coefficient of UNEMP in model 2 is larger (0.721) and more we cannot conclude
significant (t = 6.79) than in model 3 (t=2.55). This suggests that UNEMP has a stronger and more robust effect on POVRATE that model 2 is worse
when EDU and URBAN are not included in the model. One possible explanation for this difference is that EDU and URBAN are than model 1 based on
correlated with both UNEMP and POVRATE, and thus they act as confounding variables that affect the relationship between this criterion alone.
UNEMP and POVRATE. By including EDU and URBAN in model 3, we control for their effects and isolate the direct effect of
UNEMP on POVRATE, which turns out to be weaker and less significant than in model 2.
we overestimate the effect of unemployment on poverty rate Because education and urbanization are confounding variables
that affect both unemployment and poverty rate —> By including education and urbanization in model 3, we control for
their effects and isolate the direct effect of unemployment on poverty rate