ECONOMETRICS
CHAPTER 3
Regression analysis: Further details
Multivariate Case of CLRM
• Multiple regression models, that is, models in which
the dependent variable, or regressand, Y depends on
two or more explanatory variables, or regressors.
• Has two or more independent/explanatory variables.
• The simplest possible multiple regression model is
three-variable regression, with one dependent
variable and two explanatory variables
• Generalizing the two-variable population regression
function (PRF), we may write the three-variable PRF
as
Yi = β1 + β2X2i + β3X3i + ui
Cont’d…
where Y is the dependent variable, X2 and X3 the
explanatory variables (or regressors), u the stochastic
disturbance term, and i the ith observation; in case the
data are time series, the subscript t will denote the tth
observation.β1 is the intercept term.
As usual, it gives the mean or average effect on Y of
all the variables excluded from the model, although
its mechanical interpretation is the average value of Y
when X2 and X3 are set equal to zero.
The coefficients β2 and β3 are called the partial
regression coefficients, and their meaning will be
explained shortly.
Cont’d…
The assumptions are similar with simple linear
regression.
multiple regression model is linear in the
parameters, that the values of the regressors are
fixed in repeated sampling, and that there is sufficient
variability in the values of the regressors.
no exact linear relationship between X2 and X3,
technically known as the assumption of no
collinearity or no multi-collinearity if more than one
exact linear relationship is involved, is new and needs
some explanation
Cont’d…
• in the two-variable case, multiple regression analysis
is regression analysis conditional upon the fixed
values of the regressors, and what we obtain is the
average or mean value of Y or the mean response of
Y for the given values of the regressors.
Global hypothesis test (F and r2)
• In multiple regression models we will undertake two
tests of significance.
One is significance of individual parameters of the model.
This test of significance is the same as the tests discussed in
simple regression model.
The second test is overall significance of the model.
Cont’d…
Test of Overall Significance(F test)
It is joint test of the relevance of all the included
explanatory variables. Now consider the following:
This null hypothesis is a joint hypothesis that β1,
β2’…,βk , are jointly or simultaneously equal to zero. A
test of such a hypothesis is called a test of overall
significance of the observed or estimated regression line,
that is, whether Y is linearly related to X1, X2,….., Xk.
Cont’d…
• If the calculated F- value is greater than the critical F-
value with df(k-1, n-k) from F-distribution
Cont’d…
The coefficient of determination ( R2 )
In the simple regression model, we introduced R2 as
a measure of the proportion of variation in the
dependent variable that is explained by variation in
the explanatory variable.
In multiple regression model the same measure is
relevant, and the same formulas are valid but now we
talk of the proportion of variation in the dependent
variable explained by all explanatory variables
included in the model.
The coefficient of determination is:
Cont’d…
• As in simple regression, R2 is also viewed as a
measure of the prediction ability of the model over
the sample period, or as a measure of how well the
estimated regression fits the data.
• The value of R2 is also equal to the squared sample
correlation coefficient between Y &Yt estimated.
• Since the sample correlation coefficient measures the
linear association between two variables, if R2 is
high, that means there is a close association between
the values of Yt and the values of predicted by the
model, Yt estimated.
Cont’d…
• In this case, the model is said to “fit” the data well. If
R2 is low, there is no association between the values
of Yt and the values predicted by the model, Yt
estimated and the model does not fit the data well.
Omission of relevant variables and
inclusion of irrelevant variables
• The first assumption related to regression model is that all
relevant variables should be included in the model.
• Some important questions that arise in the specification of
models include what variables should be included in the
model, what are the probabilistic assumptions made about
the (dependent variable), (independent variable and
random error term).
• The specification of a linear regression model consists of
a formulation of the regression relationships and of
statements or assumptions concerning the explanatory
variables and disturbances. If any of these is violated,
e.g., incorrect functional form, incorrect introduction of
disturbance term in the model etc., then specification
error occurs.
Cont’d…
The classical assumption that the error term is independent
of the explanatory variables is violated by exclusion of a
relevant variable. This error term can be seen as a collection
of everything that is not accounted for by observable
variables included in the model
Misspecification are the errors associated with the
specification of the model, which can take many forms such
as omission of relevant variable, inclusion of unnecessary
variables, choosing a wrong functional form, errors of
measurement etc.
Omitting relevant variables from the model as a
specification error has been particularly well studied
relative to multiple regression analysis and the most serious
consequence of this type of error is likely the biased
estimates of the regression coefficients
Cont’d…
Omitted variable bias (OVB) is one of the most
common and vexing problems in ordinary least squares
regression.
It occurs when a variable that is correlated with both
the dependent and one or more included independent
variables is omitted from a regression equation
Omission of a Relevant Variable: In the classical
linear regression model, omission of a variable
specified by the truth introduces bias and decreases the
variance in all the least squares estimates.
That is, the “omission of relevant variables” in the
analysis generates inconsistency and bias in estimating
the effects of variables, though a reduction in the
variance of the estimator.
Cont’d…
Possible reduction of omitted variable bias with
the inclusion of some of the omitted variables;
Inclusion of an Irrelevant Variable: In the
classical linear regression model, inclusion of an
irrelevant variable does not introduce bias but
increases the variance in the least squares
estimates
The estimates are still inconsistent and unbiased,
and the only inconvenience is an increase of the
residual variance and hence of the estimated
standard deviation of the residual increased.
Cont’d…
In conclusion, it was found that inclusion of
irrelevant variable is a safer bias than omission of
relevant variable in model selection of a mis-
specified linear regression model.
It is clear from the result that including a collinear
variable, regardless of whether it is relevant, leads
to error inflation and an increase in VIFs, which
makes it more difficult for the researcher to
identify relevant relationships.
Inclusion of irrelevant variables is not as severe as
the consequences of omitting relevant variables
DUMMY VARIABLE
In regression analysis the dependent variable, or
regressand, is frequently influenced not only by ratio
scale variables (e.g., income, output, prices, costs,
height, temperature) but also by variables that are
essentially qualitative, or nominal scale, in nature,
such as sex, race, color, religion, nationality,
geographical region etc.
such variables usually indicate the presence or
absence of a “quality” or an attribute, such as male or
female, black or white, Catholic or non-Catholic,
Democrat or Republican, they are essentially nominal
scale variables.
Cont’d…
ANOVA models are used to assess the statistical
significance of the relationship between a quantitative
regressand and qualitative or dummy regressors.
They are often used to compare the differences in the
mean values of two or more groups or categories, and are
therefore more general than the t test which can be used
to compare the means of two groups or categories only.
Yi = β1 + β2D2i + β3iD3i + ui
where Yi = salary of public school teacher in state I, D2i = 1
if the state is in the Northeast or North Central = 0 otherwise
(i.e., in other regions of the country) D3i = 1 if the state is in
the South = 0 otherwise (i.e., in other regions of the country)
Cont’d…
If a qualitative variable has m categories, introduce only (m
− 1) dummy variables.
In our example, since the qualitative variable “region” has
three categories, we introduced only two dummies.
If you do not follow this rule, you will fall into what is
called the dummy variable trap, that is, the situation of
perfect collinearity or perfect multicollinearity, if there is
more than one exact relationship among the variables.
This rule also applies if we have more than one qualitative
variable in the model, an example of which is presented
later. Thus we should restate the preceding rule as: For each
qualitative regressor the number of dummy variables
introduced must be one less than the categories of that
variable.
Cont’d…
The category for which no dummy variable is
assigned is known as the base, benchmark, control,
comparison, reference, or omitted category. And all
comparisons are made in relation to the benchmark
category
The intercept value (β1) represents the mean value of
the benchmark category
Relaxing the CLRM basic assumptions
MULTICOLLINEARITY
Assumption 10 of the classical linear regression model
(CLRM) is that there is no multi-collinearity among the
regressors included in the regression model.
Multi-collinearity is due to Ragnar Frisch.3 Originally
it meant the existence of a “perfect,” or exact, linear
relationship among some or all explanatory variables of
a regression model.4 For the k-variable regression
involving explanatory variable X1, X2, ... , Xk, an
exact linear relationship is said to exist if the following
condition is satisfied:
Cont’d…
Why does the classical linear regression model
assume that there is no multi-collinearity among
the X’s? The reasoning is this:
If multi-collinearity is perfect in the sense of, the
regression coefficients of the X variables are
indeterminate and their standard errors are
infinite.
If multi-collinearity is less than perfect, as in, the
regression coefficients, although determinate,
possess large standard errors, which means the
coefficients cannot be estimated with great
precision or accuracy.
Cont’d…
sources of multi-collinearity
Multi-collinearity may be due to the following factors:
The data collection method employed, for example, sampling
over a limited range of the values taken by the regressors in
the population.
Constraints on the model or in the population being sampled.
For example, in the regression of electricity consumption on
income (X2) and house size (X3) there is a physical constraint
in the population in that families with higher incomes
generally have larger homes than families with lower incomes.
Model specification, for example, adding polynomial terms to
a regression model, especially when the range of the X
variable is small.
Cont’d…
An overdetermined model. This happens when
the model has more explanatory variables than
the number of observations. This could happen
in medical research where there may be a
small number of patients about whom
information is collected on a large number of
variables.
Cont’d…
In cases of near or high multi-collinearity, one is likely to
encounter the following consequences:
Although BLUE, the OLS estimators have large variances and
covariances, making precise estimation difficult.
Because of consequence 1, the confidence intervals tend to be much
wider, leading to the acceptance of the “zero null hypothesis” (i.e.,
the true population coefficient is zero) more readily. Type II error
exist(the significant variable will be insignificant).
Also because of consequence 1, the t ratio of one or more
coefficients tends to be statistically insignificant.
Although the t ratio of one or more coefficients is statistically
insignificant, R2, the overall measure of goodness of fit, can be very
high.
The OLS estimators and their standard errors can be sensitive to
small changes in the data.
Variances and standard error will be very high
Cont’d…
REMEDIAL MEASURES
Rule-of-Thumb Procedures
One can try the following rules of thumb to address the problem of
multi-collinearity, the success depending on the severity of the
collinearity problem.
o Combining cross-sectional and time series data, pooling the data
o Dropping a variable(s) and specification bias
o Transformation of variables:(first difference form/ratio transformation)
o Additional or new data
o Do Nothing:
Multi-colinearity is bad when we test hypothesis b/c it affect t-value
and confidence interval, while, it is not bad if we estimate parameter.
Cont’d…
Heteroscedasticity
If the variance of the error term is not constant.
var(ui) = E(ui/Xi) = δ2 f(Xi)
Source of Heteroscedasticity
Following the error-learning models: as people learn,
their errors of behavior become smaller over time. In
this case, σ2 i is expected to decrease.
Outliers: observation with very high or very low
values compere to the majority of observation.
As data collecting techniques improve, σ 2 i is likely
to decrease
Cont’d…
Specification error: the regression model is correctly
specified.
incorrect data transformation (e.g., ratio or first
difference transformations) and incorrect functional
form.
Note that the problem of heteroscedasticity is likely to
be more common in cross-sectional than in time series
data.
In cross-sectional data, one usually deals with members
of a population at a given point in time, such as
individual consumers or their families, firms,
industries, or geographical subdivisions such as state,
country, city, etc.
Cont’d…
Consequences of heteroscedasticity
Estimators(β0, β1, β3) are unbiased and
consistent
Variance of estimators are biased and
inconsistent.
T-value decrease , confidence interval will be
incorrect.
Testing will be inefficient
R2 will be uneffective
Cont’d…
Detection of heteroscedasticity
Nature of the data: if the data is cross-sectional
and heterogeneous
Park test: assume the var(ui) depends on one of
the regresors.
Spear man’s rank correlation test:
Breush-pagan-godfrey test
Goldfeld-Quandt Test
Cont’d…
REMEDIAL MEASURES
Make heteroscedasticity robust estimation
command: regress dependent variable, independent
variable ,vce(robust)
Transform the model
Cont’d…
Auto-correlation( common for time series data)
There is the correlation between error terms
the classical linear regression model assumes that
such autocorrelation does not exist in the disturbances
ui. Symbolically,
E(uiuj) = 0
Cont’d…
MODEL SPECIFICATION ERROR
regression model used in the analysis is “correctly” specified:
If the model is not “correctly” specified, we encounter the
problem of model specification error or model specification
bias.
According to Hendry and Richard, a model chosen for
empirical analysis should satisfy the following criteria:
Be data admissible; that is, predictions made from the model
must be logically possible.
Be consistent with theory; that is, it must make good economic
sense.
Cont’d…
Have weakly exogenous regressors; that is, the
explanatory variables, or regressors, must be uncorrelated
with the error term.
Exhibit parameter constancy; that is, the values of the
parameters should be stable. Otherwise, forecasting will
be difficulty.
Exhibit data coherency
Be encompassing; that is, the model should encompass or
include all the rival models in the sense that it is capable
of explaining their results. In short, other models cannot
be an improvement over the chosen model
Cont’d…
To sum up, in developing an empirical model, one is
likely to commit one or more of the following
specification errors:
Omission of a relevant variable(s)
Inclusion of an unnecessary variable(s)
Adopting the wrong functional form
Errors of measurement
Incorrect specification of the stochastic error term
Cont’d…
CONSEQUENCES OF MODEL SPECIFICATION ERRORS
under fitting a model, that is, omitting relevant variables,
The consequences of omitting variable are as follows:
Estimators are (α1 and α2 estimated) are biased as well as
inconsistent.
The disturbance variance σ2 is incorrectly estimated.
In consequence, the usual confidence interval and
hypothesis-testing procedures are likely to give
misleading conclusions about the statistical significance
of the estimated parameters
the forecasts based on the incorrect model and the
forecast (confidence) intervals will be unreliable
Cont’d…
overfitting a model, that is, including unnecessary
variables.
The OLS estimators of the parameters of the
“incorrect” model are all unbiased and consistent,
The error variance σ2 is correctly estimated.
The usual confidence interval and hypothesis-testing
procedures remain valid.
However, the estimated α’s will be generally
inefficient, that is, their variances will be generally
larger than those of the βˆ’s of the true model.
Cont’d…
DETECTION OF SPECIFICATION ERRORS
• Examination of Residuals.
• The Durbin–Watson d Statistic Once Again
• Ramsey’s RESET Test
• Lagrange Multiplier (LM) Test for Adding Variables.
THANK
YOU