Econometrics Problem Set on Regression Analysis
Econometrics Problem Set on Regression Analysis
E(Ui|X1i) = 0 does not hold in the regression Yi = β0 + β1X1i + Ui if there are other variables that influence Yi and are correlated with X1i. The omitted factors could bias the OLS estimator, making the estimator for β1 potentially biased and inconsistent if E(Ui|X1i) ≠ 0, indicating that certain assumptions of the classical linear regression model might be violated .
When X1 and X2 are correlated, the variance of the estimator of β1 increases compared to when they are uncorrelated. This is because the precision of the estimator decreases due to multicollinearity among the predictors. If highly correlated variables are included in the regression, it can lead to inflated standard errors making the estimates less reliable. Therefore, if the interest is solely in β1 and there is concern about multicollinearity, it might be better to leave X2 out of the regression if it is correlated with X1 .
The negative sign on the coefficient for STWMFG80 (State average manufacturing hourly wage) aligns with economic theory, as higher potential wages increase the opportunity cost of attending college, thus reducing college attendance. Conversely, the positive sign on the coefficient for CUE80 (county unemployment rate) suggests that higher unemployment reduces the opportunity cost, thereby increasing college attendance since it is harder to find a job. Both coefficients are consistent with the theory that economic conditions influence educational decisions through opportunity cost considerations .
The relationship between γ1 in the linear model Y = γ0 + γ1X + ε and β1 in the log-transformed model Y = β0 + β1 log X + U arises from the different assumptions about the nature of the relationship between Y and X. The linear model implies a constant change in Y for a unit change in X, while the log-transformed model implies a constant percentage change in Y for a percentage change in X. The choice of functional form thus significantly influences parameter interpretation and estimation. Estimating a linear relationship when the true relationship is logarithmic can lead to misspecification bias, affecting accuracy and validity of inferred causality .
E(Ui|X1i, X2i) may depend on X2i if there are unobserved factors affecting the outcome (Yi) that are correlated with the incoming status of students (X2i). This dependence could lead to biased estimates of β2. If E(Ui|X1i, X2i) ≠ 0, this indicates that the omission of variables influencing both the test score and the incoming status has resulted in omitted variable bias. Hence, the β2 OLS estimator may not provide an unbiased and consistent estimation of the causal effect of being an incoming student unless all relevant confounders are included in the model .
A positive coefficient on DadColl indicates that students whose fathers attended college complete, on average, more years of education compared to those whose fathers did not. Specifically, in this study, "dadcoll = 1" (father attended college) leads to an additional 0.696 years of education compared to "dadcoll = 0" .
Omitted variables in the police on crime rate study might include socioeconomic factors, education levels, neighborhood characteristics, and existing crime deterrence programs. Their omission can lead to bias if these factors are correlated with police size and crime rates but are not included in the model. For instance, areas with higher crime rates may also have more police presence and different socioeconomic conditions, leading to endogeneity issues where the estimated effect does not accurately reflect the causal impact of police alone .
The regression of ED on Dist without additional control variables appears to suffer from omitted variable bias. This is supported by the fact that the coefficient of Dist falls by more than 50% when additional control variables are included, indicating that the initial estimate was biased due to the omission of important student, family, and local labor market characteristics .
The estimator of β1 does not suffer from omitted variable bias when X1 and X2 are uncorrelated. This is because omitted variable bias arises when the omitted variable (here X2) is correlated with the included explanatory variable (X1). Since X1 and X2 are uncorrelated, the omission of X2 does not introduce bias in the estimate of β1 .
The R2 and adjusted R2 are similar in the extended regression model (b) because of the large sample size (n = 3796). In large samples, the adjusted R2, which accounts for the number of predictors, converges towards the R2 because the adjustment factor becomes less influential as the sample size increases. Thus, the similarity indicates a well-fitting model where the inclusion of additional predictors improves the fit without overly penalizing the model complexity .