Advanced Econometrics
Chapter: 7
MULTICOLLINEARITY
7.1: Collinearity
In a multiple regression model with two independent
variables, if there is linear relationship between independent
variables, we say that there is collinearity.
7.2: Multicollinearity
If there are more than two independent variables and they
are linearly related, this linear relationship is called
multicollinearity.
Multicollinearity arises from the presence of
interdependence among the regressors in a multivariable equation
system. The departure of orthognality in the set of regressors in a
measure of multicollinearity. It means the existence of a perfect or
exact linear relationship among some or all explanatory variables.
When the explanatory variables are perfectly correlated, the
method of least squares breaks down.
7.3: Sources of Multicollinearity
The data collection method employed for example,
sampling over a limited range of the values taken by the
regressors in the population.
Constraints on the model or in the population being
sampled. In the regression of electricity consumption (Y)
on income ( ) and house size ( ) there is a physical
constraints in the population in that families with higher
income generally larger homes than families with lower
income.
Model specification: For example adding polynomial
terms to a regression model, especially when the range of
the variable is small.
71
Advanced Econometrics
An Over determined Model: This happens when the
model has more explanatory variables than the number of
observations. This could happen in medical research,
where there may be a small number of patients about
whom information is collected on a large number of
variables.
An additional reason for multicollinearity, especially in
time series data may be that the regressors included in the
model share a common trend, that is they all increase or
decrease over time. Thus in the regression of consumption
expenditure on income, wealth and population, the
regressors income, wealth and population may all be
growing over time at more or less the same rate leading to
collinearity among these variables.
7.4: TYPES OF MULTICOLLINEARITY
There are two types of multicollinearity.
Perfect Multicollinearity
Relates to the situation where explanatory variables are
perfectly linearly related with each other. Simply when
correlation between two explanatory variables is exactly one i.e.
). This situation is called perfect multicollinearity.
Imperfect Multicollinearity
If the correlation coefficient between two explanatory
variables is not equal to one but close to one approximately 0.9,
it is called high multicollinearity. If approximately 0.5,
it is called moderate and if it is called low
multicollinearity. Both are troublesome because it cannot be
easily detected.
72
Advanced Econometrics
7.5: ESTIMATION IN THE PRESENCE OF PERFECT
MULTICOLLINEARITY
The three variable regression model using deviation form as
̂ ̂
( )
̂ And
( )( )
( )
̂
( )( )
Assume that , where λ is non-zero constant. Then
( )
̂
( )( )
( ) ( )
̂
( )( ) ( )
[ ( ) ( )]
̂
[( ) ( ) ]
[ ]
̂
[ ]
̂ .
Similarly,
( )
̂
( )( )
( ) ( )
̂
( )( ) ( )
[ ( ) ( )]
̂
[( ) ( ) ]
73
Advanced Econometrics
[ ]
̂
[ ]
(̂ )
( )( ) ( )
(̂ )
( )( ) ( )
(̂ )
( )( ) ( )
(̂ )
[( ) ( ) ]
(̂ )
(̂ )
(̂ ) .
Similarly,
(̂ )
( )( ) ( )
(̂ )
( )( ) ( )
(̂ )
( )( ) ( )
(̂ )
[( ) ( ) ]
(̂ )
74
Advanced Econometrics
(̂ )
(̂ ) .
̂ ̂
Put
̂ ̂
̂ ̂
Where ̂ ̂ ̂
Regression in y on x is:
Therefore, although we can estimate ̂ uniquely, but there is
no way to estimate ̂ ̂ uniquely. Hence in the case of perfect
multicollinearity the variance and standard error of ̂ ̂
individually are infinite.
7.6: CONSEQUENCES OF MULTICOLLINEARITY
1) The estimate of the coefficient of statistical unbiased,
even multicollinearity is strong. The sample property
of unbiased of the estimate does not require that the
X‟s be uncorrelated. On the other hand sample with
multicollinear X‟s may rounder the values of the
estimate seriously imprecise.
2) If the intercorrelation between the explanatory is
perfect. Then the estimates of the coefficient are
indeterminate.
75
Advanced Econometrics
Proof: The three variable regression model using
deviation form as
̂ ̂
( )
̂ And
( )( )
( )
̂
( )( )
Assume that , where λ is non-zero constant. Then
( )
̂
( )( )
( ) ( )
̂
( )( ) ( )
[ ( ) ( )]
̂
[( ) ( ) ]
[ ]
̂
[ ]
̂ .
Similarly,
( )
̂
( )( )
( ) ( )
̂
( )( ) ( )
[ ( ) ( )]
̂
[( ) ( ) ]
[ ]
̂
[ ]
76
Advanced Econometrics
3) If the intercorrelation of the explanatory is perfectly
one. Then the standard error of these estimate become
infinitely large.
Proof:
If , the standard error the estimate become
infinitely large in the two variable model:
0 1
* ⁄
+
* ⁄
+
[ ]
⁄√
* +
Putting
* +
* +
Infinitely large.
Similarly:
77
Advanced Econometrics
* +
* ⁄
+
* ⁄
+
[ ]
⁄√
* +
Putting
* +
* +
Infinitely large
4) In case of strong multicollinearity regression
coefficients are determinate but their standard errors
are large.
Proof:
* +
Put
* +
78
Advanced Econometrics
* +
[ ]
In case of
If
* +
* +
* +
5) In case of multicollinearity the confidence interval
becomes wider.
6) In the presence of multicollinearity the t-test will be
misleading.
7) In the presence of multicollinearity prediction is not
accurate.
7.7: DETECTION OF MULTICOLLINEARITY
1. The Farrar and Glauber Test of Multicollinearity
A statistical test for multicollinearity has been developed by
Farrar and Glauber. It is really a set of three tests.
a) The first test is a 𝟐 test for the detection of the
existence and the severity of multicollinearity in a function
including several explanatory variables.
Procedure:
i.
.
ii. Choose level of significance at
iii. Test statistic to be used
* +
79
Advanced Econometrics
iv. Computations: where is the value of the
standardized correlation determinant. K is
number of explanatory variables.
v. Critical Region:
vi. Conclusion: Reject if our calculated value is
greater than table value. Otherwise accept.
b) The second test is an F-test for locating which
variables are multicollinear.
Procedure:
i.
ii. Choose level of significance at
iii. Test statistic to be used
with d.f
iv. Computations:
Compute the multiple correlation coefficients
among the explanatory variables.
v. Critical Region: F
vi. Conclusion:
Reject if our calculated value is greater than
table value. Otherwise accept.
c) The third test is a t-test for finding out the pattern
of multicillinearity that is for determining which variables are
responsible for the appearance of the multicollinear variable.
Procedure:
i.
ii. Choose level of significance at
iii. Test statistic to be used
√
with
80
Advanced Econometrics
iv. Computations:
Computed the partial correlation
coefficients.
v. Critical Region: | |
vi. Conclusion:
Reject if our calculated value
is greater than table value. Otherwise
accept.
2. High Pair Wise Correlation among Regressors
Multicollinearity exists if the pair wise or zero order
coefficients between the two regressors are very high.
3. Eigen Value and Condition Number
A condition number K is defined as
If K is between 100 and 1000, There is moderate to
strong multicollinearity and if exceeds 1000 there is severe
multicollinearity.
The condition index defined as
If is the condition effect lie between 10 and 30 then
there is moderate to strong multicollinearity and if it exceed
30 there is severe multicollinearity.
4. Tolerance and Variance Inflation Factor
As the coefficient of determination in the regression
of regressors on the remaining regressor in the model
increases towards that is as the collinearity with the other
81
Advanced Econometrics
regressor increases VIF all the increases and the limit it can
be infinite.
VIF
( )
Tolerance can also be used to detect the
multicollinearity. That is
Tolerance ( )
5. High 𝑹𝟐 but Few Significant t-Ratios
If is high the F-test in most cases will reject the
hypothesis that the partial correlation coefficients are
simultaneously equal to zero, but the individual t-test will
show that non are very few of the partial slope of coefficients
are statistically different from zero. This is the symptom of
multicollinearity.
6. Some Other Multivariate Methods
Like Principal Component Analysis (PCA), Factor
Analysis (FA) and Ridge Regression can also be used for
detection of multicollinearity.
7.8. REMEDIAL MEASURES OF MULTICOLLINEARITY
i. A Prior Information
Suppose we consider the model
Where Y = Consumption,
Income and wealth variable tends to be highly collinear.
Suppose that is the rate of change of
consumption with respect to wealth one tended the
corresponding rate with respect to income. We can then run
the regression
82
Advanced Econometrics
Where
Once we obtain we can estimate from the
postulated relationship between and .
ii. Combining Cross-sectional and Time Series Data
A variant of the extraneous are a priori information
technique is the combination of cross-sectional and time
series data known as pooling the data. The combination of
cross-sectional and time series data may be a situation of
reduction of multicollinearity.
iii. Dropping a Variable or Variables
When faced with severe multicollinearity one of the
simplest things to do is to drop one of the collinear variables.
In dropping a variable from the model we may be
committing a specification bias or specification error.
iv. Transformation of Variables
One way of minimizing this dependence is to proceed as
follows:
If the above relation holds at time “t” it must also hold at
time “t-1” because the origin of the time is arbitrary, therefore
we have
…eq(2)
is known as first difference form.
83
Advanced Econometrics
The first difference regression model often reduces the
severity of multicollinearity.
v. Additional or New Data
Since multicollinearity is a sample feature, it is possible
that in another sample involving the same variables.
Multicollinearity may not be as serious as in the first sample.
Sometimes simply increasing the size of slope may reduce the
multicollinearity problem.
vi. Other Methods
Multivariate statistical technique such as factor analysis
and principal components or other techniques such as ridge
regression are often implied to solve the problem of
multicollinearity.
84