0% found this document useful (0 votes)
18 views3 pages

Regression Analysis Problem Set Exercises

The document contains a series of exercises related to regression analysis, exploring the relationship between various factors such as education, salary, smoking habits, and study hours on outcomes like fertility, starting salary, and birth weight. It includes questions about model interpretation, causal relationships, and statistical assumptions. The exercises encourage critical thinking about the implications of regression results and the importance of controlling for confounding variables.

Uploaded by

lyukantaku
Copyright
© All Rights Reserved
We take content rights seriously. If you suspect this is your content, claim it here.
Available Formats
Download as PDF, TXT or read online on Scribd
0% found this document useful (0 votes)
18 views3 pages

Regression Analysis Problem Set Exercises

The document contains a series of exercises related to regression analysis, exploring the relationship between various factors such as education, salary, smoking habits, and study hours on outcomes like fertility, starting salary, and birth weight. It includes questions about model interpretation, causal relationships, and statistical assumptions. The exercises encourage critical thinking about the implications of regression results and the importance of controlling for confounding variables.

Uploaded by

lyukantaku
Copyright
© All Rights Reserved
We take content rights seriously. If you suspect this is your content, claim it here.
Available Formats
Download as PDF, TXT or read online on Scribd

Problem Set 1

Tutorial Exercises

Exercise 2.1
Let kids denote the number of children ever born to a woman, and let educ denote years of
education for the woman. A simple model relating fertility to years of education is
kids=β 0+ β1 educ +u
where u is the unobserved error.

1) What kinds of factors are contained in u? Are these likely to be correlated with level of
education?

2 Will a simple regression analysis uncover the ceteris paribus effect of education on
fertility? Explain.

Exercise 2.2
In the simple linear regression model y=β 0 + β 1 x +u , suppose that E(u)≠ 0. Letting α o=E (u)
, show that the model can always be rewritten with the same slope, but a new intercept and
error, where the new error has a zero expected value.

Exercise 3.4
The median starting salary for new law school graduates is determined by
log ( salary ) =β 0+ β1 LSAT + β 2 GPA+ β 3 log ( libvol )+ β 4 log ( cost )+ β 5 rank +u
where LSAT is the median LSAT score for the graduating class, GPA is the median college GPA
for the class, libvol is the number of volumes in the law school library, cost is the annual cost
of attending law school, and rank is a law school ranking (with rank=1 being the best).

1) Explain why we expect β 5 ≤ 0

2) What signs do you expect for the other slope parameters? Justify your answers.

3) Using the data in LAWSCH85, the estimated equation is


^
log ( salary )=8.34+ 0.047 LSAT +.248 GPA+ .095 log ( libvol ) +.038 log ( cost )−.0033 rank
n=136 , R 2=.842
What is the predicted ceteris paribus difference in salary for schools with a median GPA
different by one point? (Report your answer as a percentage.)

4) Interpret the coefficient on the variable log(libvol).

5) Would you say it is better to attend a higher ranked law school? How much is a difference
in ranking of 20 worth in terms of predicted starting salary?

Additional Exercises
(these will not be discussed in the tutorials, but solutions will be provided on Learn or via
Cengage)

Exercise 2.4
The data set BWGHT contains data on births to women in the United States. Two variables of
interest are the dependent variable, infant birth weight in ounces (bwght), and an
explanatory variable, average number of cigarettes the mother smoked per day during
pregnancy (cigs). The following simple regression was estimated using data on
n=1,388 births:
^
bwght =119.77−0.514 cigs

1) What is the predicted birth weight when cigs=0? What about when cigs=20 (one pack
per day)? Comment on the difference.

2) Does this simple regression necessarily capture a causal relationship between the child’s
birth weight and the mother’s smoking habits? Explain.

3) To predict a birth weight of 125 ounces, what would cigs have to be? Comment.

4) The proportion of women in the sample who do not smoke while pregnant is about .85.
Does this help reconcile your finding from part (3)?

Exercise 2.11
Suppose you are interested in estimating the effect of hours spent in an SAT preparation
course (hours) on total SAT score (sat). The population is all college-bound high school
seniors for a particular year.

1) Suppose you are given a grant to run a controlled experiment. Explain how you would
structure the experiment in order to estimate the causal effect of hours on sat.

2) Consider the more realistic case where students choose how much time to spend in a
preparation course, and you can only randomly sample sat and hours from the population.
Write the population model as
sat=β 0+ β1 hours+ u
where, as usual in a model with an intercept, we can assume E ( u )=0. List at least two
factors contained in u. Are these likely to have positive or negative correlation with hours?

3) In the equation from part (ii), what should be the sign of β 1if the preparation course is
effective?

4) In the equation from part (ii), what is the interpretation of β 0?

Exercise 3.5
In a study relating college grade point average to time spent in various activities, you
distribute a survey to several students. The students are asked how many hours they spend
each week in four activities: studying, sleeping, working, and leisure. Any activity is put into
one of the four categories, so that for each student, the sum of hours in the four activities
must be 168.

1) In the model
GPA=β 0 + β 1 study+ β 2 sleep+ β 3 work + β 4 leisure+u
does it make sense to hold sleep, work, and leisure fixed, while changing study?
2) Explain why this model violates Assumption MLR.3.

3) How could you reformulate the model so that its parameters have a useful interpretation
and it satisfies Assumption MLR.3?

Common questions

Powered by AI

The error term (u) includes unobserved factors affecting fertility such as personal preferences, cultural influences, economic conditions, and access to healthcare. These factors could be correlated with education; for example, higher education levels might be associated with economic advantages or cultural attitudes that influence fertility decisions .

The predicted ceteris paribus difference in salary for law schools with a median GPA different by one point is approximately 24.8%, as indicated by the coefficient of the GPA variable in the regression equation .

Self-selection can introduce bias if students more motivated or already inclined to perform well on SATs choose more hours of preparation. This creates endogeneity in estimating the causal effect; the unobserved factors influencing self-selection could correlate with both preparation hours and SAT scores, skewing results .

A simple regression may fail to capture the ceteris paribus effect of education on fertility due to omitted variable bias. Factors correlated with both education and fertility, like income or access to family planning, if left unaccounted for, could skew the results, making it difficult to isolate the effect of education alone .

To establish causality, I would randomly assign students to different levels of preparation course hours while controlling for previous academic performance, motivation, and test history. This ensures any observed effect on SAT scores can be attributed to the course hours. Randomization helps reduce the influence of confounding variables .

The coefficient on log(libvol), 0.095, indicates a positive relationship between the number of law library volumes and median starting salary. Specifically, a 1% increase in library volumes is associated with a 0.095% increase in salary, suggesting that greater library resources enhance perceived quality and better prepare students for high-paying jobs .

The regression may not capture a causal relationship due to potential confounding factors such as maternal health, nutrition, and socio-economic status, which affect birth weight and are correlated with smoking habits. This limits the validity of attributing changes in birth weight directly to smoking .

The rank coefficient (β₅) is expected to be non-positive because better-ranked (lower numerical rank) schools typically have higher prestige, resources, and networking opportunities, leading to better salary prospects. Hence, an increase in numerical rank (lower school quality) should negatively impact salaries .

This setup violates the assumption of no perfect multicollinearity (Assumption MLR.3) because the sum of study, sleep, work, and leisure hours is constrained to total 168, creating a dependent relationship among the variables. The proper model should account for these constraints to allow for meaningful interpretation .

If E(u)≠0, redefine the intercept by setting α₀=E(u). The model can be rewritten as y=(β₀+α₀)+β₁x+v, where v is the new error term with E(v)=0, thus adjusting the intercept while keeping the slope, β₁, unchanged .

You might also like