ECONOMETRICS: Problem Set 5
ECONOMETRICS:
Problem Set 5
ACADEMIC YEAR: 2025-2026
CAMPUS: Segovia & Madrid
Professor: Rodrigo Alegría
Email: ralegria@[Link]
1
ECONOMETRICS: Problem Set 5
PREFACE
This document contains exercises for you to practice with the content of the course in a
practical way. Some of them will be included in the so-called Problem Sets which will be
solved in class. Students are required to work by themselves on these Problem Sets. Exercises
to be solved within each Problem Set will be announced in advance to the due date. We will
solve five Problem Sets during the course.
Important: All rights reserved. No part of this document may be reproduced, in any form
or by any means, without the permission in writing from the author.
Any errors in this document are the responsibility of the author. Corrections and comments
regarding any material in this text are welcomed and appreciated.
Author: Rodrigo Alegría
e-mail: ralegria@[Link]
2
ECONOMETRICS: Problem Set 5
3
ECONOMETRICS: Problem Set 5
PS5
Estimation
Problems
COURSE CONTENT
-Chapter 7: Estimation Problems
An econometrician is an expert who will know tomorrow why the things
he predicted yesterday did not happen today.
4
ECONOMETRICS: Problem Set 5
1 Answer to the following three questions:
a- The first empirical studies aimed at measuring the impact of class size on education
performance were based on data comparing the grades in comprehensive tests
achieved by students from different schools and different class sizes. If we aimed
at measuring the relationship between class size and academic performance with
such data, could we infer that size has a causal effect on performance? Justify.
b- The presence of more policemen to fight crime is a matter of controversy. Suppose
that we have data for all the capital cities in France about crime incidence per 10,000
inhabitants and number of police units per 10,000 inhabitants. With such data,
could we obtain the causal effect of police surveillance on crime incidence?
Explain.
c- Suppose that there is a positive and strong correlation between the number of
children´s books within a home and the academic performance of the children at
that home. Could you say that the number of children´s book at home has a
positive causal effect on the academic performance of children at such home.
Justify.
2 Suppose you are interested in estimating the effect of hours spent in a SAT
preparation course (hours) on total SAT score (sat). The population is all college-bound high
school seniors for a particular year.
a- Suppose you are given a grant to run a controlled experiment. Explain how you would
structure the experiment to estimate the causal effect of hours on sat.
b- Consider the more realistic case where students choose how much time to in a
preparation course, and you can only randomly simply sat and hours from the
population. Write the population model as:
𝑠𝑎𝑡 = 𝛽 + 𝛽 ℎ𝑜𝑢𝑟𝑠 + 𝑢
List, at least, two factors contained in the random perturbance term. Are these likely
to have positive or negative correlation with hours? Explain.
3 The following equation describes the number of hours of television watched per
week by a child as a function of his age, his education, his mother´s education, his father´s
education and the number of siblings:
𝑡𝑣ℎ𝑜𝑢𝑟𝑠 ∗ = 𝛽 + 𝛽 𝑎𝑔𝑒 + 𝛽 𝑎𝑔𝑒 + 𝛽 𝑚𝑜𝑡ℎ𝑒𝑑𝑢 + 𝛽 𝑓𝑎𝑡ℎ𝑒𝑑𝑢 + 𝛽 𝑠𝑖𝑏𝑠 + 𝑢
We suspect the dependent variable contains a certain error of measurement. Explain the
consequences in your estimation results.
5
ECONOMETRICS: Problem Set 5
4 We have the following variables:
Y: Food expenditure in USA.
X: Family income.
P: Price index.
Two different regressions are estimated with the following estimation results (standard errors
are in brackets and sample size is 500):
Coefficient for Coefficient for Adjusted Determination
Regression X P coefficient
Y/P 2.462 0.614
(0.407)
Y / X; P 0.112 -0.739 0.978
(0.003) (0.114)
Find and discuss the specifation error the first model is suffering. Explain it using the estimation
results of the above table.
5 There is an econometric study at IE University which relates the average grade in
Econometrics with the time students employ in different activities during the week. Some
students are asked about how many hours they employ in four different activities: study,
sleep, work and leisure. Any activity must be included in one of these four categories such
that the time spent in the four activities is 168 hours for each student.
The model is the following:
𝐴𝐺𝐸 = 𝛽 + 𝛽1 𝑠𝑡𝑢𝑑𝑦 + 𝛽2 𝑠𝑙𝑒𝑒𝑝 + 𝛽3 𝑤𝑜𝑟𝑘 + 𝛽4 𝑙𝑒𝑖𝑠𝑢𝑟𝑒 + 𝑢
a- Find the assumption that does not hold in this model and explain why.
b- How would you rewrite the model to solve the problem?
6
ECONOMETRICS: Problem Set 5
6 In a university campus someone argues that going to class by car lowers grades. In
order to test this hypothesis, we have estimated a model, using 141 students, explaining GPA
with a dummy variable taking a value of 1 if student goes to class by class and 0 otherwise
obtaining the following regression results:
GPA = 3.43 − 0.151drive
Knowing that the standard error associated to the effect of drive is 0.065:
a.- Interpret the effect of drive variable and test its individual significance at 5%
significance level. Do you think is it realistic?
Given that the correlation between students age and drive is 0.57 and the correlation between
GPA and age is -0.48:
b.- Explain the potential problem the above model may suffer.
A new regression model is estimated such that:
GPA = 3.43 − 0.151drive − 0.056age
c.- Knowing that in this model drive is statistically insignificant and age is individually
significant, is your answer in question b consistent with the second model estimation
results? Explain.
7 We have estimated a SLRM explaining office rental prices in the city of Madrid (Y)
with the information contained in distance to the city center (X). The following two graphs:
Figure 1( Y versus X) and Figure 2 (residuals versus fitted values of Y) are related to the above
model.
Figure 1 Figure 2
7
ECONOMETRICS: Problem Set 5
a- Discus, according to the two graphs if the model may suffer a non linearity problem
b- Provide an economic reason explaining the possible non-linearity in the above
relationship.
c- How should Figure 2 be if the relationship between office rental prices and distance
was a linear relationship?
8 The following table shows two different samples with two explanatory variables each
of them in order to study the behavior of Y (dependent variable):
Sample 1 Sample 2
Observation Y X1 X2 Z1 Z2
1 1 2 4 2 4
2 4 6 12 6 12
3 2 4 11 4 8
a- Can you detect a multicollinearity problem in any of the two samples?
b- If yes, please explain the consequences in your OLS estimations in each sample.
c- If yes, please explain the strategies that you would use in order to solve the problem
in each sample.
9 Consider the regression of country level GDP per capita on percentage urban
population in several countries (1995) obtaining a determination coefficient of 0.457 and
obtaining the following graph when plotting the data (Figure 1):
Figure 1: GDPpc versus % urban pop
8
ECONOMETRICS: Problem Set 5
a- Can you detect a non-linear relationship between the two variables? Why?
b- Can you explain solutions to be implemented to solve the non-linearity problem?
Suppose now that we estimate the same model but using a semilog transformation obtaining
the following estimation results:
log (𝐺𝐷𝑃𝑝𝑐) = 4.631 + 0.052𝑢𝑟𝑏𝑎𝑛 𝑅 = 0.549
and obtaining Figure 2 when plotting the data:
Figure 2: logGDPpc versus % urban pop.
c- Compare the determination coefficients and the graphs between the two models. Do
you think the semilog transformation might be a good solution for the nonlinearity
problem? Explain your answer.
10 We have data for a sample of high schools in Vietnam where the variable math
denotes the percentage of students who passed a maths test. We want to estimate the effect
that spending per student has on the outcomes of this test and propose the following model:
𝑚𝑎𝑡ℎ = 𝛽 + 𝛽 log(𝑠𝑝𝑒𝑛𝑑) + 𝛽 log(𝑒𝑛𝑟𝑜𝑙𝑙) + 𝛽 𝑝𝑜𝑣𝑒𝑟𝑡𝑦 + 𝑢
10
ECONOMETRICS: Problem Set 5
Where poverty describes the percentage of students living below the poverty line, spend denotes
spending per student and enroll is the number of students enrolled in the high school.
a- We do not have data for poverty variable but the variable lnchprg describes the
percentage of students eligible for a program subsidizing school lunches. Why is this
variable a sensible proxy variable for poverty?
b- The table below shows the OLS estimates with and without the inclusion of lnchprg
as an explanatory variable:
Explanatory variables (1) (2)
log(spend) 11.13 7.75
(3.30) (3.04)
log(enroll) 0.022 -1.26
(0.615) (0.58)
lnchprg - -0.324
(0.036)
intercept -69.24 -23.14
(26.74) (24.99)
n 408 408
Determination coefficient 0.0293 0.1893
Explain why the effect of spending and enroll are greater in the first model than in
the second one? What about if we compare standard errors between the two models?
c- What conclusions can you derive when comparing both models?
10
ECONOMETRICS: Problem Set 5
11 We want to estimate a regression model explaining the behavior of property prices
in the city of Barcelona in 2015 (cross sectional analysis). We are provided with a dataset
containing information about property, neighborhood, and buyer´s characteristics that can
be used as explanatory variables. The following table describes those variables:
VARIABLE DESCRIPTION
NAME
advance Loan amount when buying the property
age Age of property
bathroom Number of bathroms
bedroom Number of bedrooms
buyage Age of main buyer
chnone No central heating (dummy)
dcitycenter Distance to city center (km)
floorm2 Floor area of dwelling (m2)
ftbuyer First time buyer (dummy)
lagood Dwelling is in neighborhood with higher-status social housing
labad Dwelling is in neighborhood with lower-status social housing
pflat Flat/maisonnette dwelling (dummy)
psemi Semi-detached dwelling (dummy)
pdetach Detached dwelling (dummy)
pterrace Terraced dwelling (dummy)
11
ECONOMETRICS: Problem Set 5
a- In order to avoid specification errors, which variables would you keep in your analysis
according to practical significance? Justify your choices.
b- Explain, the process you would follow to specify your final model and to choose the
final variables in your model.
c- Explain the difference between practical and statistical significance.
12
ECONOMETRICS: Problem Set 5
12 We have the following information for the annual growth rates (%) in different
countries about stock prices (Y) and in consumer prices (X):
Stock prices Predicted Estimation
Country (Y) Consumer prices (X) Y Residuals
Australia 5 4.3
Austria 11.1 4.6
Belgium 3.2 2.4
Canada 7.9 2.4
Denmark 3.8 4.2
Finland 11.1 5.5
France 9.9 4.7
Germany 13.5 2.2
India 1.5 4
Ireland 6.4 4
Israel 8.9 8.4
Italy 8.1 3.3
Japan 13.5 4.7
Mexico 4.7 5.2
Netherlands 7.5 3.6
New Zealand 4.7 3.6
Sweden 8 4
UK 7.5 3.9
USA 9 2.1
Knowing that: 𝑦 = 6.83 + 0.201𝑥
Answer to the following questions:
a- Complete the missing values in the above table.
b- Show both graphically and formally if the above data suffers from an outlier problem.
c- If the answer to b is positive, please explain any strategy you would perform in order
to solve the problem.
13 Imagine that you are interested in analyzing the determinants of infant mortality rates
worldwide. Using the Development Reports from the World Bank in 2013, you get the
following information for 248 countries:
IMR Infant Mortality rate - is the number of deaths of infants per 1,000 live births.
GDP GDP per capita (constant 2005 US$)
Source: World Bank Development Reports, 2013.
13
ECONOMETRICS: Problem Set 5
And construct the following figure:
a- Have a look at the graph above, why Angola and Guinea might be considered as
outliers in this regression model? Comment on the implications of the inclusion of
these two countries in the analysis.
b- Angola presents one of the highest infant mortality rates in this sample (103 per 1,000
live births). Compute the residual for this country given that our model predicts for
Angola an infant mortality rate of 28.6 per 1,000 live births.
c- Knowing that the standard deviation of the estimation residuals (using all the
observations) is 26.22, is Angola a significant outlier?
d- What about Guinea? Note that the estimation residual associated to Guinea
observation is 52.
14 We have representative data for 30 years old for the US. Levine, Gustafson and
Velenchik (1997) estimated a wage equation using the following variables:
Y = ln(wage)
F = a dummy variable that takes a value of 1 for smokers and 0, otherwise
ED = years of education
14
ECONOMETRICS: Problem Set 5
Two specifications are considered:
MODEL 1: Y = -0.176F omitting education
(se=0.031)
Coefficient of determination = 0.35
MODEL 2: Y = -0.080F + 0.070ED including education
(se=0.021) (se=0.0004)
Coefficient of determination = 0.68
Compare the two fitted models and explain what happens when we omit one relevant variable
(in this case, years of education).
15 Consider the following regression model to analyze salaries respect to age groups in a sample
with 150 individuals:
𝑤 =𝛽 𝑑 +𝛽 𝑑 +𝛽 𝑑 +𝑢
Such that d1 refers to individuals between 20-30 years old, d2 accounts for individuals between
30-40 years old and finally, d3 captures individuals between 40-50 years old.
In order to test for heteroscedasticity, the second group of age is eliminated and we run two
different regressions models (with the same specification). The first model, taking into
consideration the youngest individuals group (with 50 observations and with SSR1=234). The
second regression model accounting for the oldest individuals group (with 50 observations
and SSR2=387).
a-Test for heteroscedasticity at 1% significance level and interpret your result.
15
ECONOMETRICS: Problem Set 5
16 A time series analysis is conducted to see the relationship between Consumption and income
in a country. The analysis is carry out using yearly observations from 1959 until 1988. Since
the researcher suspects heteroscedasticity issues, she decides to perform the GQT by running
the same simple linear regression model, first using yearly data from 1959 until 1971 and
secondly using yearly observations from 1976 until 1988. She obtains that in the first
subsample the SSR is 2532224 while the SSR in the second subsample is 10339356.
a- Help her by testing heteroscedasticity at 5% significance level.
b- Explain the implications in the estimation results of the model of your result in the above
question.
17 We have data regarding profits and sales for 20 companies and are interested in estimating a
models explaining profits with sales. The following graph shows the existing relationship
between the two variable.
Profits vs. Sales
200
150
100
50
0
0 5 10 15 20 25
a- Interpret the above diagram. Could you infer any possible potential problem?
b- Ordering the sample from the lowest to the highest sales and dropping the six central
observations, we obtain SSR1=1.134 for the first observation when running a SLRM
explaining profits with sales and a SSR2=32.934 when estimating the same model for the
last observations. Test for heteroscedasticity at 5% significance level.
20
ECONOMETRICS: Problem Set 5
18 Answer to the following multiple choice questions about estimation problems:
A- Autocorrelation refers to a situation in which:
a- Successive error terms derived from the application of regression analysis to time
series data are correlated.
b- There is a high degree of correlation between two or more of the independent
variables included in a multiple regression model.
c- The dependent variable is highly correlated with the independent variable(s) in a
regression analysis.
d- The application of a multiple regression model yields estimates that are nonlinear in
form.
B- A situation in which measures of two or more variables are statistically related but
are not in fact causally linked because the statistical relationship is caused by a third
omitted variable is called:
a- Partial correlation
b- Linear correlation
c- Spurious correlation
d- Marginal correlation
C- Step-wise regression is the most widely used search procedure of developing the
……….. regression model without examining all possible models.
a- worst
b- best
c- medium
d- least
D- If there is measurement error in both dependent and explanatory variables of your
simple linear regression model, then
a- OLS is unbiased but inefficient.
b- OLS is unbiased but inconsistent.
c- OLS is biased and inefficient.
d- OLS is biased but efficient.
21
ECONOMETRICS: Problem Set 5
E- A non-formal way to detect a non-linearity problem is plotting your model fitted
values versus the
a- Values of your independent variables
b- Values of your explanatory variables
c- Model residuals
d- Model predictions
F- Multicollinearity refers to a situation in which
a- The dependent variable is highly correlated with the explanatory variables included
in the regression model.
b- There is a high degree of correlation between the explanatory variables included in a
multiple regression model.
c- The application of a multiple regression model yields estimates that are nonlinear
form.
d- None of the above.
G- If your dataset has heteroscedasticity, but you completely ignore the problem and
use OLS, you will
a- Get biased estimates of the parameters.
b- Get parameter standard errors that could be either too large or too small.
c- Get t-statistics that make you too optimistic about your parameters being statistically
different from zero.
d- Get t-statistics that make you too pessimistic about your parameters being statistically
different from zero.
H- A useful graphical method for detecting the presence of heteroscedasticity is
a- Plot against each variable in turn
b- Plot the residuals from a preliminary regression against the variables, each in turn
c- Plot the squared residuals from a preliminary regression against the variables, each
in turn
d- Plot the logarithm of the squared residuals from a preliminary regression against the
variables, each in turn
22
ECONOMETRICS: Problem Set 5
23