Econometrics Practice Test: 300 Questions
Econometrics Practice Test: 300 Questions
56. The Total Sum of Squares (SST) measures: A. The variation in x. B. The total variation in the
dependent variable y . C. The variation explained by the model. D. The variation due to error.
Answer: B
57. The R2 of a regression is defined as: A. SSR/SST B. SSE/SST C. SST /SSE D.
1 + (SSR/SST ) Answer: B
58. If R2 = 0.65, this means: A. 65% of the variation in y is explained by x. B. The slope is 0.65. C.
The correlation between x and y is 0.65. D. The error is 65%. Answer: A
59. Which assumption states E(u∣x) = 0? A. Homoskedasticity B. Zero Conditional Mean C.
Random Sampling D. No Perfect Collinearity Answer: B
60. If E(u∣x) = 0, then: A. u and x are correlated. B. u and x are uncorrelated. C. u is constant. D. x
is constant. Answer: B
61. The assumption Var(u∣x) = σ 2 is known as: A. Heteroskedasticity B. Homoskedasticity C.
Consistency D. Normality Answer: B
62. If the variance of the error term increases as income increases, the data exhibits: A.
Homoskedasticity B. Heteroskedasticity C. Perfect collinearity D. Bias Answer: B
63. The variance of the OLS slope estimator β^1 is given by: A. σ 2 / ∑(xi − x
ˉ)2 B. σ 2 C.
^
64. To decrease the variance of β1 , one could: A. Decrease the sample size. B. Increase the variation
in the independent variable x. C. Increase the error variance σ 2 . D. Remove the intercept. Answer:
B
65. An estimator is "unbiased" if: A. Its variance is zero. B. It always equals the true parameter. C. Its
expected value equals the true population parameter. D. It has the smallest variance. Answer: C
66. In the regression log(wage) = β0 + β1 educ + u, the coefficient β1 is interpreted as: A. The
dollar change in wage for one year of education. B. The percentage change in wage for one year
of education (multiplied by 100). C. The percentage change in wage for a 1% change in education.
D. The elasticity of wage with respect to education. Answer: B
67. In the regression log(y) = β0 + β1 log(x) + u, β1 is: A. The semi-elasticity. B. The elasticity of
SSR/n D. ^2i
∑u Answer: B
70. If the correlation coefficient between x and y is r, then R2 in a simple regression is: A. r B. r2
C. r D. 1 − r Answer: B
71. Regression through the origin forces the intercept β0 to be: A. 1 B. yˉ C. 0 D. β1 Answer: C
72. The "Explained Sum of Squares" (SSE) is: A. ∑(yi − yˉ)2 B. ∑(y^i − yˉ)2 C. ∑ u
^2i D.
2
∑(y^i − yi ) Answer: B
73. Assumptions SLR.1 through SLR.4 ensure that OLS estimators are: A. Normal B. Unbiased C.
Efficient D. Positive Answer: B
74. Ideally, we want the error variance σ 2 to be: A. As large as possible. B. As small as possible. C.
Equal to 1. D. Equal to the sample size. Answer: B
75. If we change the units of measurement of the dependent variable y (e.g., dollars to cents),
what happens to R2 ? A. It increases by 100. B. It decreases by 100. C. It stays the same. D. It
becomes zero. Answer: C
~
76. If we multiply the independent variable x by a constant c, the new slope estimate β1 will be: A.
77. The fraction of the total variation in y that is not explained by the regression is: A. R2 B.
1 − R2 C. SSR/SSE D. σ 2 Answer: B
78. SLR.3 requires that: A. x is constant. B. x varies in the sample. C. y varies in the sample. D. The
sample size is infinite. Answer: B
79. Which term represents the predicted change in y given a one-unit change in x? A. β^0 B. u
^ C.
^ 2
β1 D. R Answer: C
80. If the covariance between x and y is positive, the slope of the regression line will be: A.
Negative B. Zero C. Positive D. Undefined Answer: C
81. The sample average of the OLS residuals is: A. yˉ B. x ^ 2 Answer: C
ˉ C. Zero D. σ
^ is: A. 1 B. -1 C. Zero D. σ 2
82. The sample covariance between the regressor x and the residuals u
Answer: C
83. When estimating β0 and β1 , we lose how many degrees of freedom? A. 0 B. 1 C. 2 D. n Answer:
C
84. The standard error of the estimate (SER) is measured in the same units as: A. The independent
variable x. B. The dependent variable y . C. The slope β1 . D. It is unitless. Answer: B
85. Which of the following is an algebraic property of OLS? A. Unbiasedness B. Consistency C. The
regression line passes through the sample means. D. Efficiency Answer: C
86. If the true relationship is linear but you fit a curved line, you violate: A. SLR.1 (Linear in
parameters) B. SLR.2 (Random sampling) C. SLR.5 (Homoskedasticity) D. None of the above
Answer: A
87. In a wage equation, if u includes "ability" and high-ability people tend to get more education (
0, causing bias. C. OLS is still unbiased. D. R2 will
x), then: A. E(u∣x) = 0 holds. B. E(u∣x) =
be 1. Answer: B
88. The variance of β^0 depends on: A. Only σ 2 . B. σ 2 , sample size, and sample values of x. C. Only
93. In the model cons = β0 + β1 inc + u, β1 is: A. The Average Propensity to Consume. B. The
^
102. If we add a constant to every xi , what happens to the slope β1 ? A. It increases. B. It decreases.
of OLS residuals: A. Is always zero. B. Is not necessarily zero. C. Is always 1. D. Minimizes SSE.
Answer: B
105. If the slope β^1 is positive, it means: A. x causes y . B. x and y are positively correlated in the
110. If y is temperature in Fahrenheit and x is year, converting y to Celsius will: A. Change the R2 . B.
Change the t-statistics. C. Linearly transform the coefficients. D. Make the slope zero. Answer: C
111. In a simple regression, if the sample size n increases, the variance of β^1 : A. Increases. B.
113. The "population regression function" (PRF) is: A. E(y∣x) = β0 + β1 x. B. y^ = β^0 + β^1 x. C. yˉ.
random sample data. C. It is always changing over time. D. It is not calculated correctly. Answer: B
116. Which OLS algebraic property ensures the regression line passes through the means? A.
∑u ^i = 0 C. The formula for β^0 = yˉ − β^1 x
^i = 0 B. ∑ xi u ˉ. D. SST = SSE + SSR. Answer:
C
117. For unbiasedness, do we need the error term u to be normally distributed? A. Yes. B. No. C.
Only if n is small. D. Only if x is constant. Answer: B
118. If x is measured with error, the OLS estimator of the slope is generally: A. Unbiased. B. Biased
towards zero (attenuation bias). C. Biased away from zero. D. Consistent. Answer: B
119. The term σ (sigma) represents: A. The variance of x. B. The standard deviation of the error term
u. C. The standard error of β^1 . D. The slope. Answer: B
122. The "method of moments" derivation of OLS relies on sample analogs of: A. E(u) = 0 and
E(xu) = 0. B. Var(u) = σ 2 . C. u ∼ N ormal. D. y = mx + b. Answer: A
123. If all data points lie exactly on a line: A. R2 = 0. B. SSR = 0. C. β^1 = 0. D. SST = 0. Answer:
B
124. Which of the following increases the standard error of the slope? A. Larger n. B. Larger
variation in x. C. Larger error variance σ 2 . D. Zero covariance between x and u. Answer: C
125. Homoskedasticity means Var(u∣x) is: A. A function of x. B. Constant. C. Zero. D. Infinite.
Answer: B
126. If we run regression of y on x and get slope 2, and regression of x on y and get slope 0.5, then
R2 is: A. 0.25 B. 0.5 C. 1.0 D. Undefined. Answer: C
127. Changing the unit of x from years to months (multiplying x by 12) divides the slope coefficient
by 12. A. True. B. False. Answer: A
128. In the model y = β0 + β1 x + u, if we drop β0 , we force the line to: A. Be horizontal. B. Be
Answer: A
136. If x is binary (0 or 1), the slope β1 is: A. The difference in means of y between the two groups. B.
139. Which plot is best for detecting heteroskedasticity in simple regression? A. Histogram of x. B.
Plot of residuals vs. x. C. Plot of y vs. time. D. Pie chart. Answer: B
140. If the error term u is correlated with x, we say x is: A. Exogenous. B. Endogenous. C.
Homoskedastic. D. Significant. Answer: B
Chapter 3: Multiple Regression Analysis: Estimation
141. The multiple regression model allows us to: A. Eliminate the error term. B. Control for multiple
factors simultaneously. C. Avoid using computers. D. Guarantee causality. Answer: B
142. In the model y = β0 + β1 x1 + β2 x2 + u, β1 measures: A. The effect of x1 on y ignoring x2 . B.
The effect of x1 on y holding x2 fixed (ceteris paribus). C. The correlation between x1 and x2 . D.
A. The part of x1 that is uncorrelated with the other regressors. B. The total variation in x1 . C. The
147. Omitted Variable Bias occurs when: A. We include an irrelevant variable. B. We exclude a variable
that affects y and is correlated with included regressors. C. The sample size is too small. D. The
error term is heteroskedastic. Answer: B
148. If we omit a variable x2 that has a positive effect on y (β2 > 0) and is positively correlated
with x1 , the bias in β^1 will be: A. Positive (Upward). B. Negative (Downward). C. Zero. D.
Undefined. Answer: A
149. If x1 and x2 are uncorrelated, omitting x2 from the model will: A. Bias β^1 . B. Not bias β^1 . C.
150. Multicollinearity refers to: A. High (but not perfect) correlation among independent variables. B.
Correlation between x and y . C. Correlation between u and y . D. Perfect correlation between x1
and x2 . Answer: A
151. High multicollinearity leads to: A. Biased coefficients. B. Large variances (standard errors) for the
estimators. C. Small confidence intervals. D. High t-statistics. Answer: B
152. The Variance Inflation Factor (VIF) is used to detect: A. Heteroskedasticity. B. Serial correlation.
C. Multicollinearity. D. Non-normality. Answer: C
153. A VIF of 1 indicates: A. Perfect multicollinearity. B. No correlation between that regressor and
others. C. Extreme bias. D. A calculation error. Answer: B
154. The Gauss-Markov Theorem states that under assumptions MLR.1–MLR.5, OLS is: A. BLUE
(Best Linear Unbiased Estimator). B. Biased but consistent. C. Nonlinear. D. Normally distributed.
Answer: A
155. "Best" in BLUE means: A. Smallest bias. B. Smallest variance among linear unbiased estimators. C.
Highest R2 . D. Easiest to calculate. Answer: B
156. In multiple regression, the degrees of freedom is: A. n − 1 B. n − k C. n − k − 1 D. n Answer:
C
2
157. The unbiased estimator of the error variance σ 2 in multiple regression is: A. SSR/n B.
SSR/(n − k − 1) C. SSE/n D. SST /(n − 1) Answer: B
158. Including an irrelevant variable (one with β = 0) in a regression: A. Creates bias. B. Does not
create bias but increases the variance of other estimators. C. Decreases R2 . D. Violates MLR.4.
Answer: B
159. If a variable is highly significant, it means: A. It is economically important. B. The coefficient is
large. C. It is unlikely the true coefficient is zero. D. The data is perfect. Answer: C
160. The intercept β0 in a multiple regression represents: A. The value of y when all x's are zero. B.
(where β2 < 0) causes: A. Upward bias in β^1 . B. Downward (negative) bias in β^1 . C. No bias. D.
171. In a log-level model log(y) = β0 + β1 x + u, a 1 unit increase in x leads to: A. 100β1 % change
172. The Adjusted R2 penalizes for: A. Adding variables. B. Removing variables. C. High correlation. D.
Low variance. Answer: A
173. Adjusted R2 can be: A. Greater than 1. B. Negative. C. Greater than R2 . D. Equal to 100. Answer:
B
174. When choosing between non-nested models with different numbers of variables, one should
compare: A. SSR B. R2 C. Adjusted R2 D. The intercepts Answer: C
175. To minimize the variance of the OLS estimator, we prefer: A. Small sample size. B. Small variation
in x. C. Large variation in x. D. Large error variance. Answer: C
2
176. Perfect collinearity happens if: A. x2 = 2x1 . B. x2 = log(x1 ). C. x2 = x21 . D. Correlation is 0.8.
Answer: A
177. Which of the following is an example of nonlinear relationship handled by OLS? A.
y = β0 + β1 x + β2 x2 + u B. y = β0 xβ1 C. y = 1/(β0 + β1 x) D. y = β0 + β1γ x Answer: A
178. If we run a regression on y , x1 , and x2 , and find x1 has a positive coefficient, but a simple
regression of y on x1 gives a negative coefficient, this is likely due to: A. Calculation error. B.
violation. Answer: B
179. The assumption E(u) = 0 basically defines: A. The slope. B. The intercept. C. The sample size.
D. The variance. Answer: B
180. Under Gauss-Markov assumptions, Var(β^j ) decreases as: A. σ 2 increases. B. n increases. C.
SSTj decreases. D.
Rj2 increases. Answer: B
181. If V IFj = 10, then Rj2 is: A. 0.1 B. 0.9 C. 0.5 D. 0.0 Answer: B
182. The error term u represents: A. Measurement error in y . B. Omitted variables. C. Random noise.
D. All of the above. Answer: D
183. If we estimate price = β0 + β1 sqrf t + β2 bdrms + u, and bdrms is highly correlated with
sqrf t, then: A. Estimates will be biased. B. Standard errors for β^1 and β^2 will be large. C. R2 will
188. If a variable is constant for all observations: A. It explains everything. B. It violates the "No
Perfect Collinearity" assumption (if intercept is included). C. It has a large variance. D. It should be
the dependent variable. Answer: B
189. The primary cost of including an irrelevant variable is: A. Bias. B. Inconsistency. C. Lower
efficiency (higher variances) for other estimators. D. Lower R2 . Answer: C
190. Which is true about the Intercept estimate? A. It is usually the most important parameter. B. It is
often not meaningful if x = 0 is impossible. C. It is always zero. D. It has no variance. Answer: B
191. In the partialling out process, the residuals r^i1 represent: A. The part of x1 explained by x2 . B.
The part of x1 uncorrelated with x2 . C. The error term of the main regression. D. The predicted
values of y . Answer: B
192. If Rj2 = 0 (regressor xj is uncorrelated with all other regressors), then the VIF is: A. 0. B. 1. C.
regression) is _____ than the variance from the multiple regression. A. Larger. B. Smaller. C. The
same. D. Unknown. Answer: B
203. When estimating the ceteris paribus effect of a variable, control variables are included to: A.
Increase R2 . B. Reduce error variance. C. Avoid omitted variable bias. D. All of the above. Answer:
D
204. In the model log(y) = β0 + β1 x1 + β2 x2 + u, if x1 is measured in logs, β1 is: A. Elasticity. B.
correlation between included and omitted variables (δ1 ). C. Both A and B. D. Neither. Answer: C
omitted variable causing bias. C. There is multicollinearity. D. The sample size is too small. Answer:
B
212. A variable should be removed from a model if: A. It is insignificant (t < 1). B. It induces perfect
collinearity. C. It is correlated with y . D. It reduces Adjusted R2 . Answer: B
213. The "Best" in Gauss-Markov refers to: A. Efficiency (minimum variance). B. Unbiasedness. C.
Consistency. D. Linearity. Answer: A
214. If x1 and x2 are uncorrelated, omitting x2 : A. Biases β^1 . B. Does not bias β^1 . C. Makes β^1 zero.
2
D. Increases R . Answer: B
215. With heteroskedasticity, OLS is: A. Biased. B. Inconsistent. C. No longer BLUE. D. Impossible to
calculate. Answer: C
216. The sample regression line always passes through (x ˉ2 , ..., yˉ). A. True. B. False. Answer: A
ˉ1 , x
217. If the intercept is excluded, R2 : A. Is always higher. B. Is always lower. C. Can be negative or
misleading. D. Is unchanged. Answer: C
218. The variance of the intercept β^0 is minimized when: A. x
ˉ = 0. B. x
ˉ is large. C. Sample size is
2
small. D. σ is large. Answer: A
219. Standard errors of coefficients measure: A. Sampling variability of the estimates. B. The error
term variance. C. The spread of x. D. The bias. Answer: A
220. If x1 is measured in thousands and x2 in units, this causes: A. Multicollinearity. B. Bias. C.
B. x3 C. u D. Parameters. Answer: C
(SSR): A. Increases or stays same. B. Decreases or stays same. C. Must stay same. D. Becomes
zero. Answer: B
224. The term "linear" in multiple linear regression means: A. Linear in variables. B. Linear in
parameters. C. Straight line graph. D. No exponentials. Answer: B
225. If we use Generalized Least Squares (GLS) instead of OLS when there is heteroskedasticity: A.
We get biased estimates. B. We get BLUE estimates. C. We get inconsistent estimates. D. It makes
no difference. Answer: B
226. In a model with k regressors and intercept, the number of estimated parameters is: A. k B.
k + 1 C. n D. n − k Answer: B
227. The sample average of fitted values yˉ^ equals: A. 0. B. yˉ. C. x
^ . Answer: B
ˉ. D. σ
228. If V IF > 10, it is a common rule of thumb for: A. High heteroskedasticity. B. High
multicollinearity. C. Significance. D. Autocorrelation. Answer: B
229. Dropping a collinear variable to fix multicollinearity might lead to: A. Omitted variable bias. B.
Higher variance. C. Lower R2 . D. Nonlinearity. Answer: A
230. The variance of β^j is proportional to: A. The sample size n. B. The inverse of the sample size 1/n
234. If the absolute value of the t-statistic is greater than the critical value, we: A. Fail to reject the
null hypothesis. B. Reject the null hypothesis. C. Accept the null hypothesis. D. Change the model.
Answer: B
235. The null hypothesis H0 : βj = 0 typically implies: A. The variable xj has a large effect on y . B.
The variable xj has no ceteris paribus effect on y . C. The intercept is zero. D. The variable xj is
negative. Answer: B
236. A p-value is: A. The probability the null hypothesis is true. B. The probability of observing a test
statistic as extreme as the one computed, assuming the null is true. C. The significance level α. D.
The probability the alternative is true. Answer: B
237. If the p-value is 0.03 and the significance level is 0.05, we: A. Reject the null. B. Fail to reject the
null. C. Accept the null. D. Do nothing. Answer: A
238. The 95% confidence interval roughly corresponds to: A. β^ ± 1se(β^) B. β^ ± 2se(β^) C.
239. To test joint significance of multiple variables (e.g., β1 = 0 and β2 = 0), we use: A. The t-test.
SSRur
SSRr
Answer: A
241. In the F-test, q represents: A. The sample size. B. The number of regressors. C. The number of
restrictions (hypotheses). D. The error variance. Answer: C
242. The "Restricted Model" in an F-test is the model where: A. All variables are included. B. The
restrictions imposed by the null hypothesis are assumed true (variables removed). C. Only the
intercept is included. D. No intercept is included. Answer: B
243. If the F-statistic is very large, we likely: A. Reject the joint null hypothesis. B. Fail to reject the
null. C. Have a calculation error. D. Have multicollinearity. Answer: A
244. Testing the "Overall Significance" of the regression means testing: A. β0 = 0 B.
2 2
β1 = β2 = ... = βk = 0 C. R = 1 D. σ = 0 Answer: B
245. Which relationship holds between the t-stat and F-stat for a single restriction (q = 1)? A.
F = t B. F = t2 C. F = t D. F = 2t Answer: B
246. A "statistically significant" coefficient means: A. It is economically large. B. It is not zero. C. We
have strong evidence to reject the hypothesis that it is zero. D. It explains all variation. Answer: C
247. Economic significance refers to: A. The size of the t-statistic. B. The magnitude and practical
importance of the coefficient estimate. C. The p-value. D. The F-statistic. Answer: B
248. If a confidence interval includes zero, we usually: A. Reject the null hypothesis that β = 0. B.
Fail to reject the null hypothesis that β = 0. C. Conclude the variable is highly significant. D.
Assume the data is bad. Answer: B
249. To test H0 : β1 = β2 , we can rewrite the model using: A. θ = β1 − β2 B. θ = β1 + β2 C.
0.5, 2.5
Variances and the covariance between β^1 and β^2 . C. The sum of standard errors. D. The
D. 1/se(β^1 ) Answer: B
261. Non-normality of errors in large samples: A. Invalidates t-tests. B. Does not matter for t-tests
(due to CLT). C. Makes OLS biased. D. Requires non-parametric methods. Answer: B
262. The p-value for a one-sided test is _____ the p-value for a two-sided test (assuming the sign
matches). A. Half B. Double C. The same as D. Square root of Answer: A
263. If we reject the null hypothesis, we prove the alternative is true. A. True. B. False (we only find
evidence for it). C. True, if p < 0.01. D. True, if n > 1000. Answer: B
264. The power of a test is: A. Probability of rejecting a true null (Type I error). B. Probability of
rejecting a false null. C. Probability of accepting a true null. D. Probability of accepting a false null.
Answer: B
265. Standard errors reported by software assume: A. Homoskedasticity. B. Heteroskedasticity. C.
Multicollinearity. D. Zero mean. Answer: A
266. A variable can be statistically significant but economically insignificant if: A. Sample size is very
small. B. Sample size is very large. C. Variance is high. D. It is never possible. Answer: B
267. A variable can be economically significant but statistically insignificant if: A. Sample size is very
large. B. Sample size is small (large standard error). C. The coefficient is zero. D. t-stat is 10.
Answer: B
268. The confidence interval formula β^ ± c ⋅ se(β^) assumes: A. A symmetric distribution (like t or
confidence interval: A. Contains 0. B. Does not contain 0. C. Contains 1. D. Is very wide. Answer: B
272. If we test 20 independent hypotheses at the 5% level, and all nulls are true, we expect to
reject roughly: A. 0. B. 1. C. 5. D. 20. Answer: B
273. The F-distribution has _____ degrees of freedom parameters. A. 1 B. 2 C. 3 D. n Answer: B
274. If the standard error of a coefficient doubles, the t-statistic: A. Doubles. B. Halves. C. Squares.
D. Stays the same. Answer: B
275. If the sample size n increases, the critical value for the t-test (at the same α): A. Increases. B.
Decreases (converges to Normal critical value). C. Stays constant. D. Goes to zero. Answer: B
276. Hypothesis testing is about making inferences on: A. Sample statistics. B. Population
parameters. C. Residuals. D. Data points. Answer: B
277. A Type I error occurs when: A. We reject a true null hypothesis. B. We fail to reject a false null
hypothesis. C. We reject a false null hypothesis. D. We fail to reject a true null hypothesis. Answer:
A
278. A Type II error occurs when: A. We reject a true null hypothesis. B. We fail to reject a false null
hypothesis. C. We reject a false null hypothesis. D. We fail to reject a true null hypothesis. Answer:
B
279. Increasing the significance level (α) from 0.01 to 0.05: A. Increases the probability of Type I
error. B. Decreases the probability of Type I error. C. Increases the probability of Type II error. D.
Has no effect. Answer: A
280. The p-value depends on: A. The sample data. B. The significance level. C. The population
parameter (true value). D. The alternative hypothesis only. Answer: A
281. Testing H0 : β1 = β2 = β3 = 0 requires: A. Three separate t-tests. B. One F-test with q = 3. C.
n D. n2 Answer: B
308. The Lagrange Multiplier (LM) test is an alternative to the F-test that requires estimating: A.
Only the unrestricted model. B. Only the restricted model. C. Both models. D. No models. Answer:
B
2
309. The LM statistic is calculated as: A. SSR ⋅ n B. n ⋅ Raux (from auxiliary regression of residuals).
2
C. F ⋅ q . D. t . Answer: B
310. The LM statistic follows which distribution asymptotically? A. Normal B. t C. Chi-square (χ2q ) D.
F Answer: C
311. Which assumption is NOT needed for consistency? A. Zero conditional mean (E(u∣x) = 0). B.
Normality of errors (u ∼ N ormal). C. Random sampling. D. No perfect collinearity. Answer: B
312. Ideally, for large sample inference, we want the error variance σ 2 to be: A. Small. B. Large. C.
Infinite. D. Negative. Answer: A
313. If plim(β^) = β , the estimator is: A. Unbiased. B. Efficient. C. Consistent. D. Normal. Answer: C
314. The term "plim" stands for: A. Probability Limit. B. Polynomial Limit. C. Partial Limit. D. Parameter
Linear. Answer: A
315. Asymptotic efficiency of OLS requires: A. Only MLR.1-4. B. MLR.1-5 (Gauss-Markov assumptions).
C. Normality. D. Small samples. Answer: B
316. In large samples, the t-distribution looks like: A. The Chi-square distribution. B. The Standard
Normal (Z) distribution. C. The F distribution. D. The Uniform distribution. Answer: B
317. Does heteroskedasticity affect consistency? A. Yes, it makes OLS inconsistent. B. No, OLS
remains consistent. C. Yes, if sample is small. D. Only if errors are non-normal. Answer: B
318. The auxiliary regression for the LM test regresses: A. y on x. B. Residuals from the restricted
model on all independent variables. C. Residuals on predicted values. D. x1 on x2 . Answer: B
319. For the LM test, if n ⋅ R2 exceeds the critical value, we: A. Reject the null hypothesis. B. Fail to
reject the null. C. Accept the null. D. Re-calculate. Answer: A
320. Large sample properties justify using OLS even when: A. Errors are not normal. B. Estimators are
biased in finite samples. C. Variables are correlated with errors. D. Both A and B. Answer: D
321. If correlation between x and u decreases, inconsistency: A. Increases. B. Decreases. C. Stays
same. D. Becomes zero. Answer: B
322. The "asymptotic variance" is: A. The variance limit as n → 0. B. The variance approximation for
large n. C. Always smaller than finite sample variance. D. The variance of the population. Answer:
B
323. Which test is typically easier to compute if the unrestricted model is hard to estimate? A. F-
test B. LM test C. t-test D. Wald test Answer: B
324. If a variable is omitted that is correlated with x, β^1 converges to: A. β1 B. β1 + Bias term C. 0
D. ∞ Answer: B
325. The Law of Large Numbers (LLN) generally supports: A. Unbiasedness. B. Consistency. C.
Homoskedasticity. D. Multicollinearity. Answer: B
326. The Central Limit Theorem (CLT) generally supports: A. Consistency. B. Asymptotic Normality. C.
Zero mean. D. Efficiency. Answer: B
327. If OLS is inconsistent, collecting more data: A. Fixes the problem. B. Makes the estimate
converge to the wrong value more tightly. C. Increases variance. D. Makes errors normal. Answer:
B
328. Asymptotic standard errors are obtained by plugging σ
^ into the variance formulas: A. True. B.
False. Answer: A
329. The degrees of freedom adjustment (n − k − 1) matters less as: A. n gets larger. B. n gets
smaller. C. k gets larger. D. R2 gets larger. Answer: A
330. Which assumption allows us to treat regressors as fixed in repeated samples for finite sample
analysis? A. Random sampling. B. Zero conditional mean. C. Fixed regressors (alternative to
random sampling). D. Normality. Answer: C
2
337. A quadratic model y = β0 + β1 x + β2 x2 + u allows for: A. Constant marginal effects. B.
β1 + β2 x Answer: C
339. The turning point of the quadratic y = β0 + β1 x + β2 x2 occurs at: A. x = −β1 /(2β2 ) B.
340. An interaction term β3 (x1 ⋅ x2 ) allows: A. The effect of x1 to depend on x2 . B. The effect of x1
β1 + β2 Answer: C
342. To compare a log-linear model (logy ) and a linear model (y ), we can use: A. Standard R2 . B.
Adjusted R2 . C. The squared correlation between y and y^ (after exponentiating predicted log
348. Residual analysis helps to: A. Identify outliers or influential observations. B. Prove causality. C.
Increase R2 . D. Calculate coefficients. Answer: A
349. If a variable is "over-controlled" (e.g., controlling for beer consumption when testing beer tax
effects on fatalities), we mask: A. The total effect of the policy. B. The error term. C. The
intercept. D. The sample size. Answer: A
350. Centering variables around their means before creating interactions: A. Changes the model fit (
R2 ). B. Makes the coefficients on the main effects interpretable as effects at the mean. C. Biases
the estimates. D. Removes multicollinearity completely. Answer: B
351. Can Adjusted R2 be negative? A. Yes. B. No. C. Only in time series. D. Only in simple regression.
Answer: A
352. Log-log models yield coefficients interpreted as: A. Slopes. B. Elasticities. C. Semi-elasticities. D.
Probabilities. Answer: B
353. Semi-elasticity is the interpretation for: A. Log-Log models. B. Level-Level models. C. Log-Level
or Level-Log models. D. Quadratic models. Answer: C
354. Using log(1 + y) allows us to: A. Handle cases where y = 0. B. Increase R2 . C. Reduce sample
size. D. Eliminate bias. Answer: A
355. "Goodness of fit" should not be the only criterion for model selection because: A. It ignores
statistical significance. B. It encourages overfitting (including irrelevant variables). C. It is hard to
calculate. D. It is always low. Answer: B
356. Confidence intervals for predictions are narrowest at: A. The extreme values of x. B. The sample
mean of x. C. x = 0. D. x = ∞. Answer: B
357. A standardized variable has a mean of __ and variance of __. A. 0, 1 B. 1, 0 C. 0, 0 D. 1, 1 Answer:
A
358. If the interaction term β3 in y = β0 + β1 x1 + β2 x2 + β3 x1 x2 is statistically significant: A. The
correlated. Answer: A
359. Non-nested models are models where: A. One is a special case of the other. B. Neither is a
special case of the other. C. They have the same variables. D. They have no error term. Answer: B
360. When comparing non-nested models, we prefer the one with: A. Higher Adjusted R2 . B. Lower
Adjusted R2 . C. More variables. D. Fewer variables. Answer: A
361. Polynomial regression involves: A. Powers of x (e.g., x2 , x3 ). B. Logs of x. C. Lags of x. D.
Dummy variables. Answer: A
362. If the marginal effect of x diminishes as x increases, the coefficient on x2 should be: A.
Positive. B. Negative. C. Zero. D. One. Answer: B
363. The "optimal" level of x in a quadratic model is found by: A. Setting the derivative w.r.t x to zero.
B. Setting x = 0. C. Setting y = 0. D. Maximizing R2 . Answer: A
364. Standard errors of predictions _____ as we move away from the data center. A. Decrease B.
Increase C. Stay constant D. Become zero Answer: B
365. The 95% prediction interval contains approximately 95% of: A. The true population means. B.
The future observations. C. The coefficients. D. The residuals. Answer: B
366. Data scaling affects: A. t-statistics. B. F-statistics. C. Coefficient estimates. D. R2 . Answer: C
367. Beta coefficients are useful for: A. Comparing relative strength of regressors. B. Fixing
heteroskedasticity. C. Removing outliers. D. Increasing n. Answer: A
368. If y is salary and x is sales, log(salary) = 4 + 0.3log(sales) implies a 1% increase in sales
leads to: A. 0.3% increase in salary. B. 30% increase in salary. C. $0.3 increase in salary. D. 3%
increase in salary. Answer: A
369. The adjusted R2 formula is: A. 1 − (1 − R2 )(n − 1)/(n − k − 1) B. 1 − R2 C. R2 /k D.
(SSR/n)/(SST /n) Answer: A
370. Choosing variables solely to maximize Adjusted R2 : A. Is the best strategy. B. Can lead to bias if
theoretical controls are omitted. C. Guarantees causality. D. Eliminates errors. Answer: B
difference in intercepts between females and males (base group). C. The predicted wage for
males. D. The error variance. Answer: B
373. The "Dummy Variable Trap" occurs when: A. You use too many dummies. B. You include dummies
for all groups AND the intercept (perfect collinearity). C. You use binary variables. D. The
coefficients are zero. Answer: B
374. To avoid the dummy variable trap with g groups, you should include: A. g dummies. B. g − 1
dummies (plus intercept). C. 1 dummy. D. No dummies. Answer: B
375. The "Base Group" (or Benchmark Group) is the group: A. Represented by a dummy variable. B.
Without a dummy variable (captured by the intercept). C. With the highest value. D. With the
lowest value. Answer: B
376. If the dependent variable is log(y) and the dummy coeff is 0.20, the approximate percentage
difference is: A. 20% B. 0.2% C. 2% D. 200% Answer: A
377. The exact percentage difference for a dummy coefficient β in a log-linear model is: A. 100β B.
100(eβ − 1) C. eβ D. β 2 Answer: B
378. An ordinal variable (e.g., credit rating 1 to 5) should typically be entered as: A. A single
continuous variable. B. A set of dummy variables (e.g., rate2, rate3...). C. A log variable. D. It
cannot be used. Answer: B
379. An interaction term between a dummy D and continuous variable x (D ⋅ x) allows for: A.
Different intercepts. B. Different slopes for the groups. C. Constant slopes. D. Reduced variance.
Answer: B
380. The Chow Test is used to test: A. Heteroskedasticity. B. Whether regression coefficients differ
across two groups (structural break). C. Normality. D. Serial correlation. Answer: B
381. The Chow Test is essentially an: A. F-test. B. t-test. C. LM test. D. Z-test. Answer: A
382. In the Linear Probability Model (LPM), the dependent variable is: A. Continuous. B. Binary (0 or
1). C. Logged. D. Always positive. Answer: B
383. In an LPM, y^ is interpreted as: A. The predicted probability of success (P (y = 1∣x)). B. The log
difference in wages. B. The percentage difference in wages. C. The slope of wages. D. The
intercept for men. Answer: B
401. If we have 4 regions (North, South, East, West), how many dummies do we include in the
regression (with intercept)? A. 4 B. 3 C. 2 D. 1 Answer: B
402. If we use "South" as the base group, the coefficient on "North" represents: A. The average for
North. B. The difference between North and South. C. The difference between North and the
national average. D. The sum of North and South. Answer: B
403. In the LPM, the variance of the error term depends on: A. y . B. The values of the independent
variables (probabilities). C. It is constant. D. The sample size. Answer: B
404. To fix heteroskedasticity in LPM, we often use: A. Robust standard errors. B. OLS. C. Larger
sample. D. Fewer variables. Answer: A
405. Which model ensures predicted probabilities are between 0 and 1? A. LPM B. Logit or Probit. C.
OLS. D. WLS. Answer: B
406. The Chow test follows which distribution? A. t B. F C. Chi-square D. Normal Answer: B
407. If the interaction term D ⋅ x is insignificant: A. The slopes are different. B. The slopes are not
statistically different. C. The intercepts are different. D. The model is invalid. Answer: B
408. Using a dummy dependent variable with OLS is called: A. Probit. B. Logit. C. Linear Probability
Model. D. Tobit. Answer: C
409. If a dummy variable is 1 for 50% of the sample, its variance is: A. 0.25 B. 0.50 C. 1.00 D. 0.00
Answer: A
410. In a study of discrimination, if the dummy for "minority" is negative and significant after
controlling for qualifications: A. There is evidence of potential discrimination. B. There is no
discrimination. C. The model is wrong. D. Qualifications don't matter. Answer: A
411. Which is a valid reason to use LPM over Logit/Probit? A. It is easier to interpret coefficients
directly. B. It always gives better predictions. C. It has no heteroskedasticity. D. It fits the data
perfectly. Answer: A
412. If we want to test if a program had ANY effect (intercept or slope), we test joint significance of:
A. The dummy and all its interactions. B. Just the dummy. C. Just the interactions. D. The intercept.
Answer: A
413. Intercept shifts allow for: A. Different slopes. B. Parallel regression lines with different heights. C.
Curved lines. D. Zero error. Answer: B
414. Can we use R2 to compare an LPM and a Logit model? A. Yes directly. B. No, they are estimated
differently. C. Only if N is large. D. Always. Answer: B
415. With panel data, a dummy variable for "Year" captures: A. Differences across units. B. Time-
specific effects common to all units. C. Measurement error. D. Slope changes. Answer: B
416. Policy analysis often uses "Difference-in-Differences" which relies on: A. Dummy variables for
time and treatment group. B. Only continuous variables. C. Only one period. D. Random guessing.
Answer: A
417. The estimated propensity score is often used to: A. Replace OLS. B. Match control and
treatment units. C. Calculate R2 . D. Test for normality. Answer: B
418. If we reject the null in a Chow test, it implies: A. There is a structural break (groups are different).
B. Groups are identical. C. There is heteroskedasticity. D. We need more data. Answer: A
419. Usually, the group with Dummy = 0 is called the: A. Reference category. B. Treatment group.
C. Outlier. D. Variable. Answer: A
420. If t stat for a dummy is 0.5, the difference between groups is: A. Statistically Significant. B. Not
Statistically Significant. C. Huge. D. Negative. Answer: B
Mixed Review
421. Bias vanishes as n → ∞ for: A. Unbiased estimators. B. Consistent estimators. C. All estimators.
D. Biased estimators. Answer: B
422. Which assumption allows us to separate the effect of x from the error u? A. Zero conditional
mean (E(u∣x) = 0). B. Homoskedasticity. C. Normality. D. Linear parameters. Answer: A
The OLS estimator is considered random because it depends on the random sample drawn from the population, affecting the variability and reliability of statistical inferences made from the estimates .
A positive slope with a binary independent variable indicates the difference in means of the dependent variable between the two groups defined by the binary variable .
"Partialling out" refers to isolating the effect of one independent variable on the dependent variable, controlling for the influence of other variables in the model, essentially capturing the relationship with only the uncorrelated part of the variable .
Increasing the sample size generally decreases the variance of the estimated regression coefficients, leading to more precise estimates .
Under the Gauss-Markov assumptions, the OLS estimators are the Best Linear Unbiased Estimators (BLUE), meaning they have the smallest variance among all linear unbiased estimators .
A very large F-statistic typically implies that we reject the joint null hypothesis, indicating that at least one of the variables significantly contributes to the model .
Heteroskedasticity violates the OLS assumption of constant variance among the errors, which can lead to inefficient and biased estimates, inflating Type I error rates .
Attenuation bias occurs when there is measurement error in the independent variable, leading to an underestimation of the absolute value of the slope coefficient, biasing it towards zero .
Converting units of the dependent variable, like changing from Fahrenheit to Celsius, results in a linear transformation of the coefficients without changing the fundamental relationships, such as t-statistics .
Perfect collinearity means that the independent variables are exact linear combinations of each other, leading to indeterminacy in OLS estimates, as they cannot be computed .