0% found this document useful (0 votes)
10 views1 page

Tutorial 2

The document outlines a series of exercises based on a dataset from the HighSchool and Beyond survey, focusing on the relationship between education and various demographic factors. It includes tasks such as running regressions, investigating measurement errors, and discussing external validity. The exercises require careful data analysis and consideration of potential biases in regression results.

Uploaded by

Meryem Balamou
Copyright
© All Rights Reserved
We take content rights seriously. If you suspect this is your content, claim it here.
Available Formats
Download as PDF, TXT or read online on Scribd
0% found this document useful (0 votes)
10 views1 page

Tutorial 2

The document outlines a series of exercises based on a dataset from the HighSchool and Beyond survey, focusing on the relationship between education and various demographic factors. It includes tasks such as running regressions, investigating measurement errors, and discussing external validity. The exercises require careful data analysis and consideration of potential biases in regression results.

Uploaded by

Meryem Balamou
Copyright
© All Rights Reserved
We take content rights seriously. If you suspect this is your content, claim it here.
Available Formats
Download as PDF, TXT or read online on Scribd

This is an adapted version of an exercise from the book by Stock & Watson.

Do you expect the parameter of interest to change in a regression of education on Dist? And
Use the data set CollegeDistance_adapted.dta. its standard error?
f. Now do exactly the same as in the previous exercise, but this time with the same classical
These data are taken from the HighSchool and Beyond survey conducted by the Department of measurement error in dist (i.e., in the X-variable) instead of in ed (i.e., in the Y-variable).
Education in 1980, with a follow-up in 1986. The data used here were supplied by Cecilia Rouse of g. Discuss the external validity of the results you found in question d. For simplicity, we assume that
Princeton University and were used in her paper “Democratization or Diversion? The Effect of your results are causal.
Community Colleges on Educational Attainment,” Journal of Business and Economic Statistics, April
1995, Vol. 12, No. 2, pp 217-224. h. For each type of classical and nonclassical measurement error in the independent, and in the
dependent variable, give an example of how it could occur in the following example, and whether
The original survey included students from approximately 1100 high schools, but the dataset you are and how this would bias regressions results.
working with does not include information on the high school someone went to.
This dataset excludes students in the western US states and contains some edits by myself (RvE). The example:
The effect of the average time someone spends for creating a post on social media platform Y on
All of the variables in this dataset have been measured in 1980, except for the central variable “ed” the average number of likes they get for each of these posts.
“ed” = years of education, as well as the variable incomehi, which were both measured in 1986.
Everyone in the dataset was first observed at the end of high school. This corresponds to 12 years of Both numbers are self-reported by individual respondents.
education. Each additional year of secondary education counted as a one year. Student's with
vocational degrees were assigned 13 years, associate degrees 14 years, bachelor degrees 16 years, (N.B. “Y” does not refer to a specific social medium here. I would have taken “X”, but that was
those with some graduate education were assigned 17 years, and those with a graduate (master’s) already taken.)
degree were assigned 18 years.

Carry out the following exercises. Always use a do- and a log-file for your analyses.

a. Investigate the relation between completed education (ED) and the demographic variables sex and
race.
 Investigate whether the 4 OLS assumptions have been met.
 Explain which variables you include.
Make sure to carefully check the data that you are using. You should from now on always
do that when working with data in this course! There may be mistakes in the data. Check
which variables are in the dataset and what they mean / how they are measured. How
would you go about finding out if there is anything strange with these variables? Identify
mistakes, outliers, etc., and find an appropriate way of dealing with them!

b. In the following regression, limit the sample to black & white people only. Is the black-white
education gap larger for males or for females?

c. Run a regression of years of completed education (ED) on distance to the nearest college (Dist).
What is the estimated slope?
 Investigate whether the 4 OLS assumptions have been met.

d. Run a regression of ED on Dist, but include some additional regressors to control for characteristics
of the student, the student’s family, and the local labor market. What is the estimated effect of
Dist on ED?
Is the estimated effect of Dist on ED in the regression in (d) substantively different from the
regression in (c)? Based on this, does the regression in (c) seem to suffer from important omitted
variable bias?

e. We now go back to a simple regression of ED on Dist without any additional controls.


Run this regression.
Now create a new variable, ed_new in which you add classical measurement error to the variable
ed. Create a new variable that is randomly normally distributed with mean 0 and standard
deviation 2. (tip: gen noise = rnormal(mean,var)). Add this variable to ed, to create ed_new.

You might also like