Econometrics Module Part II 2018
Econometrics Module Part II 2018
MODULE PART II
March, 2026
Debre Tabor, Ethiopia
Debre Tabor University, CAES, Department of Agricultural Economics
Table of Contents
Contents -------------------------------------------------------------------------------------------------------Page
CHAPTER 4: MULTIPLE REGRESSION ANALYSIS ................................................................................... 1
Introduction............................................................................................................................................................ 1
Review Question................................................................................................................................................... 41
APPENDIX ........................................................................................................................................................... 94
Introduction
The two-variable model studied extensively in the previous chapters is often inadequate in
practice. In our consumption–income example, for instance, it was assumed implicitly that
only income X affects consumption Y. But economic theory is seldom so simple for, besides
income, a number of other variables are also likely to affect consumption expenditure. An
obvious example is wealth of the consumer. As another example, the demand for a commodity
is likely to depend not only on its own price but also on the prices of other competing or
complementary goods, income of the consumer, social status, etc. Therefore, we need to
extend our simple two-variable regression model to cover models involving more than two
variables. Adding more variables leads us to the discussion of multiple regression models, that
is, models in which the dependent variable, or Regressand, Y depends on two or
more explanatory variables, or regressors. The simplest possible multiple regression model is
three-variable regression, with one dependent variable and two explanatory variables. In this
and the next chapter we shall study this model. Throughout, we are concerned with multiple
linear regression models, that is, models linear in the parameters; they may or may not be
linear in the variables.
4.1. Model with two Explanatory Variables
Generalizing the two-variable population regression function (PRF) (4.1.1), we may write the three-
variable PRF as:
𝑌 = 𝛽 + 𝛽 𝑋 + 𝛽 𝑋 + 𝑢 ---------------------------------------------------------------------------------4.1.1
Where 𝛽 the intercept term; the coefficients are 𝛽 𝑎𝑛𝑑𝛽 are called the partial
regression coefficients and Y is the dependent variable, 𝑋 𝑎𝑛𝑑𝑋 the explanatory
variables (or regressors), u the stochastic disturbance term, and ―i‖ the 𝑖 observation; in
case the data are time series, the subscript t will denote the 𝑡 observation.
model, the least-squares estimators, in the class of unbiased linear estimators, have
minimum variance, that is, they are Best Linear Unbiased Estimator (BLUE).
The Gaussian, standard, or classical linear regression model (CLRM),
which is the cornerstone of most econometric theory, makes 10 assumptions.
We continue to operate within the framework of the classical linear regression model
(CLRM) first introduced in Chapter 3.
Assumptions
Assumption 1: Linear regression model. The regression model is linear in the parameters.
As shown in equation (4.1.1). The model should be linear in parameter but not in explanatory
variables. Keep in mind that the Regressand Y and the regressor X themselves may be
nonlinear. Of the two interpretations of linearity, linearity in the parameters is relevant for the development of
the regression theory to be presented shortly. Therefore, from now on the term “linear” regression
will always mean a regression that is linear in the parameters; the β’s (that is, the parameters
are raised to the first power only). It may or may not be linear in the explanatory variables,
the X’s.
= 𝑢 𝑥 𝑢 𝑥
=
Where i and j are two different observations and where cov means covariance.
In words, (3.2.5) postulates that the disturbances ui and uj are uncorrelated. Technically, this is
the assumption of no serial correlation, or no autocorrelation. This means that, given Xi, the
deviations of any two Y values from their mean value do not exhibit patterns such as those
shown in Figure 3.6a and b. In Figure 4.2a, we see that the u‘s are positively correlated, a
positive u followed by a positive u or a negative u followed by a negative u. In Figure 4.2b,
the u‘s are negatively correlated, a positive u followed by a negative u and vice versa
FIGURE 4.2: Patterns of correlation among the disturbances. (a) Positive serial correlation;
(b) negative serial correlation; (c) zero correlation.
Assumption 6: zero covariance between or = .formally,
𝑐𝑜𝑣 𝑢 𝑥 = 𝑢 𝑢 ] 𝑥 𝑥 ]
= 𝑢 𝑥 = 𝑢 𝑥 𝑥 ] since 𝑢 =
= 𝑢 𝑥 𝑢 𝑥 Since 𝑥 non-stochastic
𝑢 𝑥 = 𝑠𝑖𝑛𝑐𝑒 𝑢 =
=
Assumption 7: The number of observations n must be greater than the number of
parameters to be estimated. Alternatively, the number of observations n must be greater than
the number of explanatory variables.
Assumption 8: Variability in X values. The X values in a given sample must not all be the
same. Technically, var (X ) must be a finite positive number.
To find the OLS estimators, let us first write the sample regression function (SRF) corresponding to the
PRF of (4.1.1) as follows:
𝑌̂ = 𝛽̂ + 𝛽̂ 𝑋 + 𝛽̂ 𝑋 + 𝑢̂ …………………………………………………4.3.1
where 𝑢̂ i is the residual term, the sample counterpart of the stochastic disturbance term 𝑢
As noted in Chapter 3, the OLS procedure consists in so choosing the values of the unknown
parameters that the residual sum of squares (RSS) ∑ 𝑢̂ is as small as possible. Symbolically,
𝑚𝑖𝑛 ∑ 𝑢̂ =∑ 𝑌̂ 𝛽̂ 𝛽̂ 𝑋 𝛽̂ 𝑋 … … … … … … … … … … … … … … … … … … … 4.3.2
Where the expression for the RSS is obtained by simple algebraic manipulations of (4.3.1).
The most straightforward procedure to obtain the estimators that will minimize (4.3.1) is to differentiate
it with respect to the unknowns, set the resulting expressions to zero, and solve them simultaneously.
This procedure gives the following normal equations
• Given in the equation 4.3.3 and 4.3.5: differentiate the equation 4.3.2 partially with respect
to the three unknowns and setting the resulting equations to zero, we obtain:
∑ 𝑢̂
= 2∑ 𝑌 𝛽̂ 𝛽̂ 𝑋 𝛽̂ 𝑋 =
𝛽
∑ 𝑢̂
= 2∑ 𝑌 𝛽̂ 𝛽̂ 𝑋 𝛽̂ 𝑋 𝑋 =
𝛽
∑ 𝑢̂
= 2∑ 𝑌 𝛽̂ 𝛽̂ 𝑋 𝛽̂ 𝑋 𝑋 =
𝛽
• Simplifying these, we obtain Equations. (4.3.3) to (4.3.5).
𝑌̅ = 𝛽̂ + 𝛽̂ 𝑋̅ + 𝛽̂ 𝑋̅ …………………….…………………………………….4.3.3
𝛽̂ = 𝑌̅ 𝛽̂ 𝑋̅ + 𝛽̂ 𝑋̅ )
𝛽̂ = 𝑌̅ 𝛽̂ 𝑋̅ 𝛽̂ 𝑋̅ -----------------------------------------------------------------------4.3.3a
∑ 𝑌 𝑋 = 𝛽̂ ∑ 𝑋 + 𝛽̂ ∑ 𝑋 + 𝛽̂ ∑ 𝑋 𝑋 ………………………………………….…4.3.4
∑ 𝑌 𝑋 = 𝛽̂ ∑ 𝑋 + 𝛽̂ ∑ 𝑋 + 𝛽̂ ∑ 𝑋 𝑋 ………………………………………….…4.3.5
Thus, equation known as normal equation of OLS
Differentiating this expression partially with respect to each of the k unknowns, setting the
resulting equations equal to zero, and rearranging, we obtain the following k normal equations in
the k unknowns:
∑ 𝑌 = 𝑛𝛽̂ + 𝛽̂ ∑ 𝑋 + 𝛽̂ ∑ 𝑋 + + 𝛽̂ ∑ 𝑋 …………………………………..4.3.6
∑ 𝑌 𝑋 = 𝛽̂ ∑ 𝑋 + 𝛽̂ ∑ 𝑋 + 𝛽̂ ∑ 𝑋 𝑋 …………………………………………….4.3.7
∑ 𝑌 𝑋 = 𝛽̂ ∑ 𝑋 + 𝛽̂ ∑ 𝑋 + 𝛽̂ ∑ 𝑋 𝑋 …………………………………………….4.3.8
---------------------------------------------------------------------------------
∑ 𝑌 𝑋 = 𝛽̂ ∑ 𝑋 + 𝛽̂ ∑ 𝑋 𝑋 + 𝛽̂ ∑ 𝑋 𝑋 + 𝛽̂ ∑ 𝑋 +……………..……………4.3.9
Or, switching to small letters, these equations can be expressed as
∑ 𝑦 𝑥 = 𝛽̂ ∑ 𝑥 + 𝛽̂ ∑ 𝑥 + 𝛽̂ ∑ 𝑥 𝑥 ……………………………………………...4.3.10
∑ 𝑦 𝑥 = 𝛽̂ ∑ 𝑥 + 𝛽̂ ∑ 𝑥 + 𝛽̂ ∑ 𝑥 𝑥 ……………………………………………4.3.11
---------------------------------------------------------------------------------
∑𝑦 𝑥 = 𝛽̂ ∑ 𝑥 + 𝛽̂ ∑ 𝑥 𝑥 + 𝛽̂ ∑ 𝑥 𝑥 + 𝛽̂ ∑ 𝑥 +…………………………..4.3.12
It should further be noted that the k-variable model also satisfies these equations:
∑ 𝑢̂ = 4.3. 2𝑎
∑ 𝑢̂ 𝑥 = ∑ 𝑢̂ 𝑥 ∑ 𝑢̂ 𝑥 = 4.3. 2𝑏
In the case of three variable regression model we substitute equation (4.3.3a) in equation
(4.3.4)and (4.3.5) finally we get the formula for 𝛽̂ and 𝛽̂
By Appling Cramer's rule of matrix from the normal equation of deviation
∑ 𝑦 𝑥 = 𝛽̂ ∑ 𝑥 + 𝛽̂ ∑ 𝑥 𝑥 ……………………………………………………..4.3.13
∑ 𝑦 𝑥 = 𝛽̂ ∑ 𝑥 𝑥 + 𝛽̂ ∑ 𝑋 ….……………………………………………….…4.3.14
∑ ∑ ∑ ∑
𝛽̂ = ∑ ∑ ∑
------------------------------------------------------------------4.3.15
∑ ∑ ∑ ∑
𝛽̂ = ∑ ∑ ∑
------------------------------------------------------------------4.3.16
Which give the OLS estimators of the population partial regression coefficients β2 and β3,
respectively.
In passing, note the following: (1) Equations (4.3.15) and (4.3.16) are symmetrical in nature
because one can be obtained from the other by interchanging the roles of X1and X2; (2) the
denominators of these two equations are identical; and (3) the three-variable case is a natural
extension of the two-variable case.
Having obtained the OLS estimators of the partial regression coefficients, we can derive the
variances and standard errors of these estimators. As in the two-variable case, we need the
standard errors for two main purposes: to establish confidence intervals and to test statistical
hypotheses. The relevant formulas are as follows:
̅ ∑ ̅ ∑ ̅ ̅ ∑
• 𝑣𝑎𝑟(𝛽̂ ) = [ + ∑ ∑ ∑
] 4.4.
𝛽̂ = +√𝑣𝑎𝑟 𝛽̂ 4.4.2
∑
• (𝛽̂ ) = ∑ ∑ ∑
4.4.3
• 𝛽̂ = +√𝑣𝑎𝑟 𝛽̂ 4.4.5
∑
• (𝛽̂ ) = ∑ ∑ ∑
4.4.6
• (𝛽̂ ) = ∑ 4.4.8
• 𝛽̂ = +√𝑣𝑎𝑟 𝛽̂ 4.4.9
But 𝑢̂ = 𝑌 𝛽̂ 𝛽̂ 𝑋 𝛽̂ 𝑋 . .
We get 𝑢̂ = 𝑌 𝑌̅ 𝛽̂ 𝑋̅ 𝛽̂ 𝑋̅ 𝛽̂ 𝑋 𝛽̂ 𝑋
=𝑌 𝑌̅ + 𝛽̂ 𝑋̅ + 𝛽̂ 𝑋̅ 𝛽̂ 𝑋 𝛽̂ 𝑋
=𝑦 𝛽̂ 𝑋 𝑋̅ 𝛽̂ 𝑋 𝑋̅
𝑋 𝑋̅ 𝑎𝑛𝑑 𝑋 𝑋̅ 𝑎𝑟𝑒 𝑑𝑒𝑣𝑖𝑎𝑡𝑖𝑜𝑛𝑠 𝑓𝑟𝑜𝑚 𝑟𝑒𝑠𝑝𝑒𝑐𝑡𝑖𝑣𝑒 𝑚𝑒𝑎𝑛.
𝑡𝑒𝑟𝑒 𝑓𝑜𝑟𝑒 𝑡𝑒 𝑒𝑞𝑢𝑎𝑡𝑖𝑜𝑛 𝑏𝑒𝑐𝑜𝑚𝑒𝑠
𝑢̂ = 𝑦 𝛽̂ 𝑥 𝛽̂ 𝑥 4.4.8𝑐
Now ∑ 𝑢̂ = ∑ 𝑢̂ 𝑢̂
= ∑ 𝑢̂ 𝑦 𝛽̂ 𝑥 𝛽̂ 𝑥
= ∑ 𝑢̂ 𝑦
∑ 𝑢̂ = ∑𝑦 𝛽̂ 𝑥 𝑦 𝛽̂ 𝑥 𝑦 4.4.8𝑑
The estimator ̂ can be computed from (4.4.8) once the residuals are available, but it can also be
obtained more readily by using the following relation
∑ 𝑢̂ = ∑𝑦 𝛽̂ ∑ 𝑦 𝑥 𝛽̂ ∑ 𝑦 𝑥 4.4.9
2. The mean value of the estimated 𝑌 ( =𝑌̂ ) is equal to the mean value of the actual Yi,
3. . ∑ 𝑢̂ = 𝑢̅̂ = , which can be verified from
SRF which can be expressed in the deviation forms as
𝑦 = 𝑦̂ + 𝑢̂ = 𝛽̂ 𝑥 + 𝛽̂ 𝑥 + 𝑢̂ Because 𝑦̂ = 𝛽̂ 𝑥 + 𝛽̂ 𝑥
4. The residuals 𝑢̂ are uncorrelated with 𝑋 and 𝑋 , that is
∑ 𝑢̂ 𝑋 = ∑ 𝑢̂ 𝑋 =
5. The residuals 𝑢̂ are uncorrelated with 𝑌 ; that is ∑ 𝑢̂ 𝑌 =
6. From (4.4.4) and (4.4.9) it is evident that as 𝑟 , the correlation coefficient between 𝑋 and
𝑋 , increases toward 1, the variances of 𝛽̂ 𝑎𝑛𝑑 𝛽̂ increase for given values of and
∑𝑥 or ∑ 𝑥 .
• 𝐼𝑛 𝑡𝑒 𝑙𝑖𝑚𝑖𝑡, 𝑤𝑒𝑛 𝑟 = (𝑖.𝑒., 𝑝𝑒𝑟𝑓𝑒𝑐𝑡 𝑐𝑜𝑙𝑙𝑖𝑛𝑒𝑎𝑟𝑖𝑡𝑦), 𝑡𝑒𝑠𝑒 𝑣𝑎𝑟𝑖𝑎𝑛𝑐𝑒𝑠
𝑏𝑒𝑐𝑜𝑚𝑒 𝑖𝑛𝑓𝑖𝑛𝑖𝑡𝑒.
7. It is also clear from (4.4.4) and (4.4.9) that for given values of 𝑟 and ∑ 𝑥 or ∑ 𝑥 , the
variances of the OLS estimators are directly proportional to ; that is, they increase as
increases.
Similarly, for given values of and 𝑟 , the variance of β1 is inversely proportional to ∑ 𝑥 ;
that is, the greater the variation in the sample values of X1, the smaller the variance of 𝛽̂ and
therefore 𝛽̂ can be estimated more precisely. A similar statement can be made about the
variance of 𝛽̂ .
8. Given the assumptions of the classical linear regression model, one can prove that the OLS
estimators of the partial regression coefficients not only are linear and unbiased but also have
minimum variance in the class of all linear unbiased estimators. In short, they are BLUE:
Put differently, they satisfy the Gauss-Markov theorem.
The Multiple Coefficients of Determination and the Multiple Coefficients of Correlation R
In the two-variable case we saw that as defined as measures the goodness of fit of the
regression equation; that is, it gives the proportion or percentage of the total variation in the
dependent variable Y explained by the (single) explanatory variable X. This notation of can
be easily extended to regression models containing more than two variables. Thus, in the three
variable model we would like to know the proportion of the variation in Y explained by the
variables X1 and X2 jointly. The quantity that gives this information is known as the multiple
coefficient of determination and is denoted by ; conceptually it is akin to
To derive we may follow the derivation of Recall that:
𝑌 = 𝛽̂ + 𝛽̂ 𝑋 + 𝛽̂ 𝑋 + 𝑢̂
--------------------4.4.12
𝑌 = 𝑌̂ + 𝑢̂
̂ is the estimated value of Yi from the fitted regression line and is an estimator of true
Where 𝑌
𝑌𝑖 𝑋 𝑖 𝑋2𝑖 . Upon shifting to lowercase letters to indicate deviations from the mean values,
Eq. (4.4.12) may be written as
𝑦 = 𝛽̂ 𝑥 + 𝛽̂ 𝑥 + 𝑢̂
𝑦 = 𝑦̂ + 𝑢̂ --------------------4.4.13
Squaring (4.4.13) on both sides and summing over the sample values, we obtain
∑𝑦 =∑ 𝑦̂ +∑ 𝑢̂ + 2∑𝑦
̂𝑖 𝑢
̂𝑖 4.4. 4
= ∑ 𝑦̂ +∑ 𝑢̂ 𝑤𝑦? 4.4. 4
Verbally, Eq. (4.4. 4) states that the total sum of squares (TSS) equals the explained sum of
squares (ESS) + the residual sum of squares (RSS). Now substituting for ∑ 𝑢̂ from (4.4.9),
we obtain
∑𝑦 = ∑ 𝑦̂ + ∑ 𝑦 𝛽̂ ∑ 𝑦 𝑥 𝛽̂ ∑ 𝑦 𝑥
= ∑ 𝑦̂ = 𝛽̂ ∑ 𝑦 𝑥 + 𝛽̂ ∑ 𝑦 𝑥 4.4. 5
∑̂ ̂ ∑ ̂ ∑
By definition = =∑ = ∑
4.4. 6
2
Since the quantities entering (4.4.16) are generally computed routinely, can be computed easily.
2
Note that , like 𝑟2 , lies between 0 and 1. If it is 1, the fitted regression line explains 100 percent of
the variation in Y. On the other hand, if it is 0, the model does not explain any of the variation in Y.
2
Typically, however, lies between these extreme values. The fit of the model is said to be “better’’
the closer is to 1.
Recall that in the two-variable case we defined the quantity r as the coefficient of correlation and
indicated that it measures the degree of (linear) association between two variables. The three-or-more-
variable analogue of r is the coefficient of multiple correlations, denoted by R, and it is a measure of
the degree of association between Y and all the explanatory variables jointly. Although r can be
positive or negative, R is always taken to be positive. In practice, however, R is of little importance.
2
The more meaningful quantity is . Before proceeding further, let us note the following relationship
2
between and the variance of a partial regression coefficient in the k-variable multiple regression
model given in (.4.4.10):
𝑣𝑎𝑟(𝛽̂ ) = [ ] 4.4. 7
∑𝑥
Where 𝛽̂ is the partial regression coefficient of regressor 𝑋 and is the R2 in the regression of 𝑋
on the remaining (k − 2) regressors. [Note: There are (k − 1) regressors in the k-variable regression
model.]
Partial Correlation Coefficients: Explanation of Simple and Partial Correlation Coefficients
In Chapter 2 we introduced the coefficient of correlation r as a measure of the degree of linear
association between two variables. For the three-variable regression model we can compute three
correlation coefficients: 𝑟 (correlation between Y and X1), 𝑟 (correlation coefficient between Y and
X2), and 𝑟 (correlation coefficient between X1 and X2); notice that we are letting the subscript 1
represents Y for notational convenience. These correlation coefficients are called gross or simple
correlation coefficients, or correlation coefficients of zero order. These coefficients can be
computed by the definition of correlation coefficient given in chapter 3.
But now consider this question: Does, say, r1 2 in fact measure the ―true‖ degree of (linear) association
between Y and X1 when a third variable X2 may be associated with both of them? This question is
analogous to the following question: Suppose the true regression model is (4.1.1) but we omit from the
model the variable X2 and simply regress Y on X1, obtaining the slope coefficient of, say, b12. Will this
coefficient be equal to the true coefficient β1 if the model (4.1.1) were estimated to begin with? In
general, r12 is not likely to reflect the true degree of association between Y and X1 in the presence of
X2. As a matter of fact, it is likely to give a false impression of the nature of association between Y and
X2, as will be shown shortly. Therefore, what we need is a correlation coefficient that is independent of
the influence, if any, of X2 on X1 and Y. Such a correlation coefficient can be obtained and is known
appropriately as the partial correlation coefficient. Conceptually, it is similar to the partial regression
coefficient. We define
r12.3 = partial correlation coefficient between Y and X1, holding X2 constant
r13.2 = partial correlation coefficient between Y and X2, holding X1 constant
r23.1 = partial correlation coefficient between X1 and X2, holding Y constant
These partial correlations can be easily obtained from the simple or zero order, correlation coefficients
as follows.
𝑟 𝑟 𝑟
𝑟 . = 4.4. 8
√ 𝑟 𝑟
𝑟 𝑟 𝑟
𝑟 . = 4.4. 9
√ 𝑟 𝑟
𝑟 𝑟 𝑟
𝑟 . = 4.4.2
√ 𝑟 𝑟
The partial correlations given in Equations (4.4.18) to (4.4.20) are called first order correlation
coefficients. By order we mean the number of secondary subscripts. Thus r1 2.3 4 would be the
correlation coefficient of order two, r1 2.3 4 5 would be the correlation coefficient of order three, and so
on. As noted previously, r12, r13, and so on are called simple or zero-order correlations. The
interpretation of, say, r 1 2.3 4 is that it gives the coefficient of correlation between Y and X1, holding X2
and X3 constant.
and the Adjusted ̅
̅ = 𝑛 𝑘 = ∑ 𝑢̂ 𝑛 𝑘
4.4.27
∑𝑦 𝑛
𝑛
Where k = the number of parameters in the model including the intercept term. (In the three-variable
regression, k = 3. Why?) The R2 thus defined is known as the adjusted , denoted by ̅ . The term
adjusted means adjusted for the degree of freedom (df) associated with the sums of squares entering
into (4.4.26): ∑ 𝑢̂ has n − k df in a model involving k parameters, which include the intercept term,
and ∑ 𝑦 has n − 1 df. (Why?) For the three-variable case, we know that ∑ 𝑢̂ has n − 3 df. Equation
(4.4.27) can also be written as:
̂
̅ = 4.4.28
Where ̂ is the residual variance, an unbiased estimator of true , and is the sample variance of
Y. It is easy to see that ̅ and are related because, substituting (4.4.26) into (4.4.27), we obtain
𝑛
̅ = 4.4.29
𝑛 𝑘
It is immediately apparent from Eq. (4.4.29) that:
1. For 𝑘 , ̅ which implies that as the number of X variables increases, the adjusted R2
increases less than the unadjusted ; and
2. ̅ can be negative, although R2 is necessarily non-negative. In case ̅ turns out to be negative in
an application, its value is taken as zero.
Which should one use in practice? As Theil notes:
“It is good practice to use ̅ rather than because tends to give an overly optimistic picture
of the fit of the regression, particularly when the number of explanatory variables is not very small
compared with the number of observations”.
But Thiel’s view is not uniformly shared, for he has offered no general theoretical justification for the
“superiority’’ of ̅ . For example, Goldberger argues that the following , call it modified , will
do just as well
𝑘
𝑜𝑑𝑖𝑓𝑖𝑒𝑑 =( ) 4.4.3
𝑛
Note, However, that if = 1, ̅ = = 1. When = 0, ̅ = (1 − k)/ (n− k), in which case
̅ can be negative if k > 1
Besides and adjusted as goodness of fit measures, other criteria are often used to judge the
adequacy of a regression model.
Numerical Example: Child Mortality In Relation To Per Capita GNP and Female Literacy Rate: based
∑𝑥 𝑦 ∑𝑥 == 4.76797 + 3
∑𝑥 𝑦 ∑𝑥 𝑥 = 6.379 9 + 2
= 4.76797 + 3 6.379 9 + 2 = 4. 3 6 + 3
4. 3 6 + 3
𝛽̂ = = 2.23 585732
.85 73 + 3
∑𝑥 𝑦 ∑𝑥 ∑𝑥 𝑦 ∑𝑥 𝑥
𝛽̂ =
∑𝑥 ∑𝑥 ∑𝑥 𝑥
∑𝑥 𝑦 ∑𝑥 ∑𝑥 𝑦 ∑𝑥 𝑥
2.2667 + .22 67 + = . 45 3 +
∑𝑥 ∑𝑥 ∑𝑥 𝑥 = .85 73 + 3
. 45 3 +
𝛽̂ = = . 5646595
.85 73 + 3
Observation Yi X1 X2 X3 Observation Yi X1 X2 X3
CM FLFP PGNP TFR CM FLFP PGNP TFR
1 128 37 1870 6.66 33 142 50 8640 7.17
2 204 22 130 6.15 34 104 62 350 6.6
3 202 16 310 7 35 287 31 230 7
4 197 65 570 6.25 36 41 66 1620 3.91
5 96 76 2050 3.81 37 312 11 190 6.7
6 209 26 200 6.44 38 77 88 2090 4.2
7 170 45 670 6.19 39 142 22 900 5.43
8 240 29 300 5.89 40 262 22 230 6.5
9 241 11 120 5.89 41 215 12 140 6.25
10 55 55 290 2.36 42 246 9 330 7.1
11 75 87 1180 3.93 43 191 31 1010 7.1
12 129 55 900 5.99 44 182 19 300 7
13 24 93 1730 3.5 45 37 88 1730 3.46
14 165 31 1150 7.41 46 103 35 780 5.66
15 94 77 1160 4.21 47 67 85 1300 4.82
16 96 80 1270 5 48 143 78 930 5
17 148 30 580 5.27 49 83 85 690 4.74
18 98 69 660 5.21 50 223 33 200 8.49
19 161 43 420 6.5 51 240 19 450 6.5
20 118 47 1080 6.12 52 312 21 280 6.5
21 269 17 290 6.19 53 12 79 4430 1.69
22 189 35 270 5.05 54 52 83 270 3.25
23 126 58 560 6.16 55 79 43 1340 7.17
24 12 81 4240 1.8 56 61 88 670 3.52
25 167 29 240 4.75 57 168 28 410 6.09
26 135 65 430 4.1 58 28 95 4370 2.86
27 107 87 3020 6.66 59 121 41 1310 4.88
28 72 63 1420 7.28 60 115 62 1470 3.89
29 128 49 420 8.12 61 186 45 300 6.9
30 27 63 19830 5.23 62 47 85 3630 4.1
31 152 84 420 5.79 63 178 45 220 6.09
32 224 23 530 6.5 64 142 67 560 7.2
Note: CM=Child mortality ; the number of death of children under the age of 5 in a year per 1000 live
births FLFP=Female Literacy Rate Percent PGNP= Per capita GNP in 1980
TFR= Total Fertility Rate; the average number of children born to a woman using age specific fertility
rate for a given year
Table 4.2: Estimation Procedure of the Parameters /Coefficients of the Child Mortality Regression Model
𝑦 𝑥 𝑥 𝑥 𝑥 x1y x2y
𝑌 𝑌̅ 𝑋 𝑋̅ 𝑋 𝑋̅ 𝑥 𝑥
-13.5 -14.1875 468.75 201.2852 219726.5625 -6650.390625 191.53125 -6328.125
62.5 -29.1875 -1271.25 851.9102 1616076.563 37104.60938 -1824.21875 -79453.125
60.5 -35.1875 -1091.25 1238.16 1190826.563 38398.35938 -2128.84375 -66020.625
55.5 13.8125 -831.25 190.7852 690976.5625 -11481.64063 766.59375 -46134.375
-45.5 24.8125 648.75 615.6602 420876.5625 16097.10938 -1128.96875 -29518.125
67.5 -25.1875 -1201.25 634.4102 1443001.563 30256.48438 -1700.15625 -81084.375
28.5 -6.1875 -731.25 38.28516 534726.5625 4524.609375 -176.34375 -20840.625
98.5 -22.1875 -1101.25 492.2852 1212751.563 24433.98438 -2185.46875 -108473.125
99.5 -40.1875 -1281.25 1615.035 1641601.563 51490.23438 -3998.65625 -127484.375
-86.5 3.8125 -1111.25 14.53516 1234876.563 -4236.640625 -329.78125 96123.125
-66.5 35.8125 -221.25 1282.535 48951.5625 -7923.515625 -2381.53125 14713.125
-12.5 3.8125 -501.25 14.53516 251251.5625 -1911.015625 -47.65625 6265.625
-117.5 41.8125 328.75 1748.285 108076.5625 13745.85938 -4912.96875 -38628.125
23.5 -20.1875 -251.25 407.5352 63126.5625 5072.109375 -474.40625 -5904.375
-47.5 25.8125 -241.25 666.2852 58201.5625 -6227.265625 -1226.09375 11459.375
-45.5 28.8125 -131.25 830.1602 17226.5625 -3781.640625 -1310.96875 5971.875
6.5 -21.1875 -821.25 448.9102 674451.5625 17400.23438 -137.71875 -5338.125
-43.5 17.8125 -741.25 317.2852 549451.5625 -13203.51563 -774.84375 32244.375
19.5 -8.1875 -981.25 67.03516 962851.5625 8033.984375 -159.65625 -19134.375
-23.5 -4.1875 -321.25 17.53516 103201.5625 1345.234375 98.40625 7549.375
𝛽̂ = 𝑌̅ 𝛽̂ 𝑋̅ 𝛽̂ 𝑋̅
𝛽̂ = 263.64 5856
Now let us bring in female literacy as measured by the female literacy rate (FLR). A priori, we
expect that FLR too will have a negative impact on CM. Now when we introduce both the variables
in our model, we need to net out the influence of each of the regressors. That is, we need to estimate
the (partial) regression coefficients of each regressor. Thus our model is:
̂ =𝛽 + 𝛽 𝑋 + 𝛽 𝑋 + 𝑢
𝑠𝑒 = .5932 . 9 .2 99 = .7 77 ̅ = .698
Where figures in parentheses are the estimated standard errors. Before we interpret this regression,
observe the partial slope coefficient of PGNP, namely, −0.0056.
Let us now interpret these regression coefficients: The coefficient −2.2316 tells us that holding the
influence of PGNP constant, on average, the number of deaths of children under 5 goes down by
about 2.23 per thousand live births as the female literacy rate increases by one percentage point. The
intercept value of about 263, mechanically interpreted, means that if the values of FLR and PGNP
rate were fixed at zero, the mean child mortality would be about 263 deaths per thousand live births.
−0.0056 is the partial regression coefficient of PGNP and tells us that with the influence of FLR held
constant, as PGNP increases, say, by a dollar, on average, child mortality goes down by 0.0056
units. To make it more economically interpretable, if the per capita GNP goes up by a thousand
dollars, on average, the number of deaths of children under age 5 goes down by about 5.6 per
thousand live births. Of course, such an interpretation should be taken with a grain of salt. All one
could infer is that if the two regressors were fixed at zero, child mortality will be quite high, which
makes practical sense. The R2 value of about 0.71 means that about 71 percent of the variation in
child mortality is explained by FLR and PGNP, a fairly high value considering that the maximum
value of R2 can at most be 1. All told, the regression results make sense.
We know by now that if our sole objective is point estimation of the parameters of the regression
models, the method of ordinary least squares (OLS), which does not make any assumption about the
probability distribution of the disturbances ui, will suffice. But if our objective is estimation as well
as inference, then, as argued in chapter 3, we need to assume that the ui follow some probability
distribution. For reasons already clearly spelled out, we assumed that the ui follow the
normal distribution with zero mean and constant variance σ2. We continue to make the same
assumption for multiple regression models. With the normality assumption and following the
discussion of Chapters 3, we find that the OLS estimators of the partial regression coefficients,
estimators, are best linear unbiased estimators (BLUE). Moreover, the estimators, 𝛽̂ 𝛽̂ , and 𝛽̂ are
themselves normally distributed with means equal to true β2, β1, and β0 and the variances given in
this Chapter.
As a result and following, one can show that, upon replacing ζ 2 by its unbiased estimator ̂ 2 in the
computation of the standard errors, each of the following variables
We can derive the t-value of the OLS estimates follows the t distribution with n− 3 df.
̂
𝑡̂ = ̂
-------------------------4.5.1
̂
𝑡̂ = ̂
Where: SE = is standard error and k = number of parameters in the model including the
intercept.
𝑠𝑒 = .5932 .2 99 . 9
-------4.5.2
𝑡 = 22.74 .6293 2.8 87
𝑝 𝑣𝑎𝑙𝑢𝑒 = . . . 65
= .7 77 ̅ = .698
What about the statistical significance of the observed results? Consider, for example, the of
PGNP of −0.0056. Is this coefficient statistically significant, that is, statistically different
from zero? Likewise, is the coefficient of FLR of −2.2316 statistically significant? Are both
coefficients statistically significant? To answer this and related questions, let us first consider
the kinds of hypothesis testing that one may encounter in the context of a multiple regression
model.
Hypothesis Testing In Multiple Regressions:
General Comments
Once we go beyond the simple world of the two-variable linear regression model, hypothesis
testing assumes several interesting forms, such as the following:
1. Testing hypotheses about an individual partial regression coefficient
2. Testing the overall significance of the estimated multiple regression model, that is, finding
out if all the partial slope coefficients are simultaneously equal to zero
3. Testing that two or more coefficients are equal to one another
4. Testing that the partial regression coefficients satisfy certain restrictions
5. Testing the stability of the estimated regression model over time or in different cross-
sectional units
6. Testing the functional form of regression models. Since testing of one or more of these
types occurs so commonly in empirical analysis, we devote a section to each type.
Hypothesis Testing About Individual Regression Coefficients
If we invoke the assumption that ui ∼ N (0, ζ 2), then, as noted first, we can use the t test to
21 Bish200821@[Link] | Prepared by: Bishaw A. (MSc) : Econometrics Module
Debre Tabor University, CAES, Department of Agricultural Economics
test a hypothesis about any individual partial regression coefficient. To illustrate the
mechanics, consider the child mortality regression(4.5.2): Let us postulate that
Example: 𝛽 =
𝛽
The null hypothesis states that, with X1 (female literacy rate) held constant, X2 (PGNP) has
no (linear) influence on Y (child mortality). To test the null hypothesis, we use the t test
given in (4.5.1). Following Chapter 3 if the computed t value exceeds the critical t value at
the chosen level of significance, we may reject the null hypothesis; otherwise, we may not
reject it. For our illustrative example, using (4.5.1) and noting that β2 = 0 under the null
hypothesis, we obtain
̂ .
𝑡̂ = ̂ = = 2.8 87 = 𝑡
.
.
FIGURE 4.3: The 95% confidence interval for t (60 df)
which in the present case is 0.0065. The interpretation of this p value (i.e., the exact level of
significance) is that if the null hypothesis were true, the probability of obtaining a t value of as much
as 2.8187 or greater (in absolute terms) is only 0.0065 or 0.65 percent, which is indeed a small
probability, much smaller than the artificially adopted value of α = 5%.
This example provides us an opportunity to decide whether we want to use a one-tail or a two-tail t
test. Since a priori child mortality and per capita GNP are expected to be negatively related (why?),
we should use the one-tail test. That is, our null and alternative hypothesis should be:
𝛽
𝛽
As the reader knows by now, we can reject the null hypothesis on the basis of the one-tail t test in the
present instance.
In Chapter 3 we saw the intimate connection between hypothesis testing and confidence interval
estimation. For our example, the 95% confidence interval for β2 is:
𝛽̂ 𝑡 𝑠𝑒 (𝛽̂ ) 𝛽2 𝛽̂ + 𝑡 𝑠𝑒 (𝛽̂ ) 4.5.3
. 56 2. . 2 𝛽2 . 56 + 2. . 2
or
𝛽̂ (𝛽̂ )𝑡 = . 56 .85 . 2 2. = . 56 .79
. 96 𝛽 . 6 4.5.4
That is, the interval, -0.0096 to -0.0016 includes the true β2 coefficient with 95% confidence
coefficient. Thus, if 100 samples of size 64 are selected and 100 confidence intervals like (4.5.4) are
constructed, we expect 95 of them to contain the true population parameter β2. Since the interval
(4.5.4) does not include the null-hypothesized value of zero, we can reject the null hypothesis that
the true β2 is zero with 95% confidence.
Testing the Overall Significance of the Sample Regression
Throughout the previous section we were concerned with testing the significance of the estimated
partial regression coefficients individually, that is, under the separate hypothesis that each true
population partial regression coefficient was zero. But now consider the following hypothesis:
𝛽 =𝛽 = 4.5.5
This null hypothesis is a joint hypothesis that β1 and β2 are jointly or simultaneously equal to zero. A
test of such a hypothesis is called a test of the overall significance of the observed or estimated
regression line, that is, whether Y is linearly related to both X1 and X2.
Can the joint hypothesis in (4.5.5) be tested by testing the significance of 𝛽̂ and 𝛽̂ individually as
in previous? The answer is no, and the reasoning is as follows.
testing a series of single [individual] hypotheses is not equivalent to testing those same hypotheses
jointly. The intuitive reason for this is that in a joint test of several hypotheses any single hypothesis
is ―affected‘‘ by the information in the other hypotheses.
The upshot of the preceding argument is that for a given example (sample) only one confidence
interval or only one test of significance can be obtained. How, then, does one test the simultaneous
null hypothesis that β1 = β2 = 0? The answer follows.
The Analysis of Variance Approach to Testing the Overall Significance of an Observed
Multiple Regression: The F Test
For reasons just explained, we cannot use the usual t test to test the joint hypothesis that the true
partial slope coefficients are zero simultaneously. However, this joint hypothesis can be tested by the
analysis of variance (ANOVA) technique, which can be demonstrated as follows.
Model Sum of square df Mean square F-statistic
Regression/ ∑ 𝑌̂ 𝑌̅ 𝑘
= =
model/SSE
Due to residual/SSR ∑ 𝑌 𝑌̂ 𝑛 𝑘
= = ̂
𝑛 𝑘
Total + 𝑛
Recall identity
∑𝑦 = ∑ 𝑦̂ + ∑ 𝑢̂
∑𝑦 = 𝛽̂ ∑ 𝑦 𝑥 + 𝛽̂ ∑ 𝑦 𝑥 + ∑ 𝑢̂
------------4.5.6
= +
Now it can be shown that, under the assumption of normal distribution for ui and the null hypothesis
β1 = β2 = 0,
.
We obtain F –statistic: = = 73.8325
.
The p value of obtaining an F value of as much as 73.8325 or greater is almost zero, leading to the
rejection of the hypothesis that together PGNP and FLR have no effect on child mortality. If you
were to use the conventional 5 percent level-of-significance value, the critical F value for 2 df in the
numerator and 60 df in the denominator (the actual df, however, are 61) is about 3.15 or about 4.98
if you were to use the 1 percent level of significance. Obviously, the observed F of about 74 far
exceeds any of these critical F values.
𝑛 𝑘 𝑛 𝑘 𝑘
= = = 4.5.9
𝑘 𝑘 𝑛 𝑘
Thus the F test, which is a measure of the overall significance of the estimated regression, is also a
test of significance of R2. In other words, testing the null hypothesis (4.5.5) is equivalent to testing
the null hypothesis that (the population) R2 is zero.
Table 4.5: ANOVA table in terms of R2
with (n− k) df, where k is the total number of parameters estimated, including the constant term. The
se (𝛽̂ 𝛽̂ ) is obtained from the following well-known formula
If we substitute the null hypothesis and the expression for the 𝑠𝑒(𝛽̂ 𝛽̂ ) into (4.5.12), our test
statistic becomes
(𝛽̂ 𝛽̂ )
𝑡= 4.5. 4
√𝑣𝑎𝑟(𝛽̂ + 𝑣𝑎𝑟 𝛽̂ ) 2 𝑐𝑜𝑣(𝛽̂ 𝛽̂ )
3.9
= = 3.3 3
. 442
The reader can verify that for 6 df (why?) the observed t value exceeds the critical t value even at the
0.002 (or 0.2 percent) level of significance (two-tail test); the p value is extremely small, 0.000006.
Hence we can reject the hypothesis that the coefficients of X2and X3 in the cubic cost function are
identical.
There are occasions where economic theory may suggest that the coefficients in a regression model
satisfy some linear equality restrictions. For instance, consider the Cobb–Douglas production
function:
𝑌 =𝛽 𝑋 𝑋 𝑒 4.5. 6
Where Y = output, X1 = labor input, and X2 = capital input. Written in log form, the equation
becomes
𝑌𝑖 = 𝛽 + 𝛽 𝑋 + 𝛽 𝑋 + 𝑢 4.5.17
Now if there are constant returns to scale (equiproportional change in output for an equiproportional
change in the inputs), economic theory would suggest that
H0:β1 + β2 = 1----------------------------------------------------4.5.18
estimated β1 and β2 (say, by OLS method), a test of the hypothesis or restriction (4.5.18) can be
conducted by the t test of (4.5.12), namely
(𝛽̂ + 𝛽̂ ) 𝛽 +𝛽
𝑡= 4.5. 9
𝑠𝑒(𝛽̂ 𝛽̂ )
(𝛽̂ + 𝛽̂ )
𝑡=
√𝑣𝑎𝑟(𝛽̂ + 𝑣𝑎𝑟 𝛽̂ ) 2 𝑐𝑜𝑣(𝛽̂ 𝛽̂ )
Where (𝛽̂ + 𝛽̂ ) = under the null hypothesis and where the denominator is the standard error of
(𝛽̂ + 𝛽̂ ). Then following if the t value computed from (4.5.19) exceeds the critical t value at the
chosen level of significance, we reject the hypothesis of constant returns to scale; otherwise we do
not reject it.
The preceding t test is a kind of postmortem examination because we try to find out whether the
linear restriction is satisfied after estimating the ―unrestricted‘‘ regression. A direct approach would
be to incorporate the restriction (4.5.18) into the estimating procedure at the outset. In the present
example, this procedure can be done easily. From (4.5.18) we see that
𝛽 = 𝛽 4.5.2 𝑎
𝑜𝑟 𝛽 = 𝛽 4.5.2 𝑏
Therefore, using either of these equalities, we can eliminate one of the β coefficients in (4.5.17) and
estimate the resulting equation. Thus, if we use (4.5.20a), we can write the Cobb–Douglas
production function as:
𝑌𝑖 = 𝛽 + 𝛽 𝑋 + 𝛽 𝑋 + 𝑢 4.5.21
𝑌𝑖 = 𝛽 + 𝑋 𝛽 𝑋 + 𝛽 𝑋 +𝑢
= 𝛽 + 𝑋 𝛽 𝑋 + 𝑋 +𝑢
= 𝛽 + 𝑋 + 𝛽 𝑋 𝑋 +𝑢
𝑌𝑖 𝑋 = 𝛽 + 𝛽 𝑋 𝑋 + 𝑢 𝑜𝑟
4.5.22𝑎
𝑋
𝑌𝑖 𝑋 = 𝛽 + 𝛽 𝑙𝑛 ( ) + 𝑢 4.5.22𝑏
𝑋
Where RSSUR =∑ 𝑢̂ RSS of the unrestricted regression; ∑ 𝑢̂ 𝑜𝑟 RSSR = RSS of the restricted
regression m= number of linear restrictions (1 in the present example) k = number of parameters in
the unrestricted regression n = number of observations. Follows the F distribution with m, (n − k) df.
(Note: UR and R stand for unrestricted and restricted, respectively.) The F test above can also be
expressed in terms of R2 as follows:
𝑚
= 4.5. 24
𝑛 𝑘
where R2 UR and R2R are, respectively, the R2 values obtained from the unrestricted and restricted
regressions, that is, from the regressions (4.5.17) and (4.5.22). It should be noted that
𝑎𝑛𝑑 ∑ 𝑢̂ ∑ 𝑢̂ 4.5.25
Example: The Cobb–Douglas Production Function for the Ethiopian Economy, 1955–1974
By way of illustrating the preceding discussion consider the data given in Table 4.6. Attempting to
fit the Cobb–Douglas production function to these data, yielded the following results:
𝑝 𝑣𝑎𝑙𝑢𝑒 = . 44 . 849 .
= .995 = . 36
32 Bish200821@[Link] | Prepared by: Bishaw A. (MSc) : Econometrics Module
Debre Tabor University, CAES, Department of Agricultural Economics
Where RSSUR is the unrestricted RSS, as we have put no restrictions on estimating (4.5.26). We will
see in nest section how to interpret the coefficients of the Cobb– Douglas production function. As
you can see, the output/labor elasticity is about 0.34 and the output/capital elasticity is about 0.85. If
we add these coefficients, we obtain 1.19, suggesting that perhaps the Ethiopian economy during the
stated time period was experiencing increasing returns to scale. Of course, we do not know if 1.19 is
statistically different from 1.
To see if that is the case, let us impose the restriction of constant returns to scale, which gives the
following regression:
̂ = .4947 + . 53 𝑙𝑛𝑐𝑎𝑝𝑖𝑡𝑎𝑙 𝑙𝑎𝑏𝑜𝑟 + 𝑢
𝑡 = 4. 6 2 28. 56
-------4.5.27
𝑝 𝑣𝑎𝑙𝑢𝑒 = . 7 .
= .9777 = . 66
Where RSSR is the restricted RSS, for we have imposed the restriction that there are constant returns
to scale.
Since the dependent variable in the preceding two regressions is different, we have to use the F test
given in (4.5.24). We have the necessary data to obtain the F value.
. 66 . 36
= == 3.75
. 36 2 3
Note in the present case m = 1, as we have imposed only one restriction and (n − k) is 17, since we
have 20 observations and three parameters in the unrestricted regression. This F value follows the F
distribution with 1 df in the numerator and 17 df in the denominator. The reader can easily check
that this F value is not significant at the 5% level.
The conclusion then is that the Ethiopian economy was probably characterized by constant returns to
scale over the sample period and therefore there may be no harm in using the restricted regression
given in (4.5.27). As this regression shows, if capital/labor ratio increased by 1 percent, on average,
labor productivity went up by about 1 percent.
Testing the Functional Form of Regression: Choosing Between Linear and Log–Linear
Regression Models:
The choice between a linear regression model (the Regressand is a linear function of the regressors)
or a log–linear regression model (the log of the Regressand is a function of the logs of the
regressors) is a perennial question in empirical analysis. We can use a test proposed by MacKinnon,
White, and Davidson, which for brevity we call the MWD test to choose between the two models.
To illustrate this test, assume the following
H0: Linear Model: Y is a linear function of regressors, the X‘s.
H1: Log–Linear Model: ln Y is a linear function of logs of regressors, the logs of X‘s. Where, as
usual, H0 and H1 denote the null and alternative hypotheses.
The MWD test involves the following steps:
Step I: Estimate the linear model and obtain the estimated Y values. Call them Yf (i.e.,𝑌̂).
Step: II: Estimate the log–linear model and obtain the estimated ln Y values; call them ln f (i.e.,
̂ ).
𝑙𝑛𝑌
Step III: Obtain Z1 = (𝑙𝑛 𝑌 𝑓 𝑙𝑛 𝑓 ).
Step IV: Regress Y on X‘s and Z1 obtained in Step III. Reject H0 if the coefficient of Z1 is
statistically significant by the usual t test.
Step V: Obtain Z2 = (antilog of ln f - Y f ).
Step VI: Regress log of Y on the logs of X‘s and Z2. Reject H1 if the coefficient of Z2 is statistically
significant by the usual t test.
Example:
The demand for roses.* Table 4.7 gives quarterly data on these variables:
Y = quantity of roses sold, dozens
X2 = average wholesale price of roses, $/dozen
X3 = average wholesale price of carnations, $/dozen
X4 = average weekly family disposable income, $/week
X5 = the trend variable taking values of 1, 2, and so on, for the period
1971–III to 1975–II in the Detroit metropolitan area
You are asked to consider the following demand functions:
a). Estimate the parameters of the linear model and interpret the results.
b. Estimate the parameters of the log-linear model and interpret the results.
c. β2, β3, and β4 give, respectively, the own-price, cross-price, and income elasticities of demand.
What are their a priori signs? Do the results concur with the a priori expectations?
For illustrative purposes, we will consider the demand for roses as a function only of the prices of
roses and carnations, leaving out the income variable for the time being. Now we consider the
following models:
Linear model: 𝑌𝑡 = + 𝑋 + 𝑋 + 𝑢𝑡 4.5.28
– 𝑙𝑛𝑌𝑡 = 𝛽 + 𝛽 𝑙𝑛𝑋 + 𝛽 𝑙𝑛𝑋 + 𝑢𝑡 4.5.29
Where Y is the quantity of roses in dozens, X2 is the average wholesale price of roses ($/dozen), and
X3 is the average wholesale price of carnations ($/dozen). A priori, α2 and β2 are expected to be
negative (why?), and α3 and β3 are expected to be positive (why?). As we know, the slope
coefficients in the log–linear model are elasticity coefficients.
The regression results are as follows:
In this section we demonstrate transformations of nonlinear relationships into linear ones so that we can work
within the framework of the classical linear regression model by taking up the multivariable extension of the
two-variable log–linear model; The specific example we discuss is the celebrated Cobb–Douglas production
function of production theory .The Cobb–Douglas production function, in its stochastic form, may be
The properties of the Cobb–Douglas production function are quite well known:
1. β1 is the (partial) elasticity of output with respect to the labor input, that is, it measures the percentage
change in output for, say, a 1 percent change in the labor input, holding the capital input constant.
2. Likewise, β2 is the (partial) elasticity of output with respect to the capital input, holding the labor input
constant.
3. The sum (β1 + β2) gives information about the returns to scale, that is, the response of output to a
proportionate change in the inputs. If this sum is 1, then there are constant returns to scale, that is, doubling
the inputs will double the output, tripling the inputs will triple the output, and so on. If the sum is less than 1,
there are decreasing returns to scale doubling the inputs will less than double the output. Finally, if the sum is
greater than 1, there are increasing returns to scale doubling the inputs will more than double the output.
A Numerical Example: we obtained the data shown in Table 4.8; these data are for the agricultural
sector of Ethiopia for 1958–1972. Estimate and interpret the coefficient of X1 and X2 if the Cobb–
Douglas production function model, in its stochastic form, is given as follows:
𝑌 =𝛽 𝑋 𝑋 𝑒 4.6.2𝑎
where Y = output X1 = labor input X2 = capital input u = stochastic disturbance term
e = base of natural logarithm or
𝑌𝑖 = 𝛽 + 𝛽 𝑋 + 𝛽 𝑋 + 𝑢 4.6.2𝑏
𝑌𝑖 = + 𝛽 𝑙𝑛 𝑋 + 𝛽 𝑙𝑛 𝑋 + 𝑢 4.6.2𝑐
where = 𝑙𝑛𝛽 . Thus written, the model is linear in the parameters β0, β2, and β3 and is
therefore a linear regression model. Notice, though, it is nonlinear in the variables Y and X
but linear in the logs of these variables. In short, (4.6.2b/c) is a log-log, double-log, or log-
linear model, the multiple regression counterpart of the two-variable log-linear model.
Table 4.8: Real Gross Product, Labor Days, and Real Capital Input in the Agricultural Sector of Taiwan, 1958–1972
𝑠𝑒 = 2.4495 .5398 . 2
From Equation (4.6.3) we see that in the Ethiopian agricultural sector for the period 1958–1972 the
output elasticities of labor and capital were 1.4988 and 0.4899, respectively. In other words, over the
period of study, holding the capital input constant, a 1% increase in the labor input led on the
average to about a 1.5 % increase in the output. Similarly, holding the labor input constant, a 1%
increase in the capital input led on the average to about a 0.5% increase in the output. Adding the
two output elasticities, we obtain 1.9887, which gives the value of the returns to scale parameter. As
is evident, over the period of the study, the Taiwanese agricultural sector was characterized by
increasing returns to scale.
From a purely statistical viewpoint, the estimated regression line fits the data quite well. The R 2
value of 0.8890 means that about 89% of the variation in the (log of) output is explained by the (logs
of) labor and capital. We were seen in previous section how the estimated standard errors can be
used to test hypotheses about the ―true‖ values of the parameters of the Cobb– Douglas production
function for Ethiopian economy.
We now consider a class of multiple regression models, the polynomial regression models that
have found extensive use in econometric research relating to cost and production functions. In
introducing these models, we further extend the range of models to which the classical linear
regression model can easily be applied.
Consider the cost function given by = 𝛽 + 𝛽 𝑋 + 𝛽2𝑋 4.6.4𝑎
which is called a quadratic function, or more generally, a second-degree polynomial in the variable
X—the highest power of X represents the degree of the polynomial.
The stochastic version of (4.6.3) may be written as
= 𝛽 + 𝛽 𝑋 + 𝛽2𝑋 + 𝑢̂ 4.6.4𝑏
The general kth degree polynomial regression may be written as
= 𝛽 + 𝛽 𝑋 + 𝛽2𝑋 + +𝛽𝑘𝑋 + 𝑢̂ 4.6.5
Notice that in these types of polynomial regressions there is only one explanatory variable on the
right-hand side but it appears with various powers, thus making them multiple regression models.
Incidentally, note that if Xi is assumed to be fixed or nonstochastic; the powered terms of Xi also
become fixed or nonstochastic.
Do these models present any special estimation problems? Since the second-degree polynomial
(4.6.4b) or the kth degree polynomial (4.6.5) is linear in the parameters, the β‘s, they can be
estimated by the usual OLS or ML methodology. But what about the collinearity problem? Aren‘t
the various X‘s highly correlated since they are all powers of X? Yes, but remember that terms like
39 Bish200821@[Link] | Prepared by: Bishaw A. (MSc) : Econometrics Module
Debre Tabor University, CAES, Department of Agricultural Economics
X 2, X 3, X 4, etc., are all nonlinear functions of X and hence, strictly speaking, do not violate the no
Multicollinearity assumption. In short, polynomial regressions models can be estimated by the
techniques presented in this chapter and present no new estimation problems.
Example: Estimating the Total Cost Function: As an example of the polynomial regression, consider the
data on output and total cost of production of a commodity in the short run given in Table 4.9. What type of
regression model will fit these data?
Table 4.9: Total Cost (Y) and Output(X)
This S shape of the total cost curve can be captured by the following cubic or third-degree polynomial:
𝑌 = 𝛽 + 𝛽 𝑋 + 𝛽2𝑋 + 𝛽3𝑋 + 𝑢̂ 4.6.6
Where Y = total cost and X = output
When the third-degree polynomial regression was fitted to the data of Table 4.3, we obtained the following
results:
𝑌̂ = 4 .7667 + 63.4776𝑋 + 2.96 5𝑋 + .9396𝑋 4.6.7
𝑠𝑒 = (6.3753 4.7786 .9857 . 59 = .9983
Review Question
1. based on the data given in the Table 4.7 gives quarterly data on each variables and You are asked
to consider the following demand functions:
Linear model: 𝑌𝑡 = + 𝑋 + 𝑋 + 𝑋 + 𝑋 + 𝑢𝑡
– 𝑙𝑛𝑌𝑡 = 𝛽 + 𝛽 𝑙𝑛𝑋 + 𝛽 𝑙𝑛𝑋 + 𝛽 𝑙𝑛𝑋 + 𝛽 𝑙𝑛𝑋 + 𝑢𝑡
Y = quantity of roses sold, dozens
X2 = average wholesale price of roses, $/dozen
X3 = average wholesale price of carnations, $/dozen
X4 = average weekly family disposable income, $/week
X5 = the trend variable taking values of 1, 2, and so on, for the period 1971–III to 1975–II in the
Detroit metropolitan area
a) Estimate the parameters of the linear model and interpret the results.
b) Estimate the parameters of the log-linear model and interpret the results.
c) β2, β3, and β4 give, respectively, the own-price, cross-price, and income elasticities of demand.
What are their a priori signs? Do the results concur with the a priori expectations?
d) How would you compute the own-price, cross-price, and income elasticities for the linear
model?
e) On the basis of your analysis, which model, if either, would you choose and why?
2. Is it possible to obtain the following from a set of data?
a) 𝑟 = .9 𝑟 = .2 𝑟 = .8
b) 𝑟 = .6 𝑟 = .9 𝑟 = .5
c) 𝑟 = . 𝑟 = .66 𝑟 = .7
3. From the following data estimate the partial regression coefficients, their standard errors, and the
adjusted and unadjusted R2 values:
∑ 𝑋 𝑋̅ = 84855. 96
∑ 𝑋 𝑋̅ = 28 .
∑ 𝑋 𝑋̅ 𝑋 𝑋̅ = 4796.
∑ 𝑋 𝑋̅ 𝑌 𝑌̅ = 74778.346
∑ 𝑋 𝑋̅ 𝑌 𝑌̅ = 425 .9
𝑌 = + 𝑋 +𝑢
𝑌 =𝛽 +𝛽 𝑋 +𝑢
𝑌 = + 𝑋 + 𝑋 +𝑢
Note: Estimate only the coefficients and not the standard errors.
a) is = ?? Why or why not?
b) Is = 𝛽 ? ? Why or why not?
c) What important conclusion do you draw from this exercise?
( )
5. Show that 𝑟 . = 𝑎𝑛𝑑 𝑖𝑛𝑡𝑒𝑟𝑝𝑟𝑢𝑡 𝑡𝑒 𝑒𝑞𝑢𝑎𝑡𝑖𝑜𝑛
6. If the relation 𝑋 + 𝑋 + 𝑋 = for all values of X1, X2, and X3, find the values of the
three partial correlation coefficient?
7. Consider the following regression results 𝑌̂ = 6899 2978.5𝑋 𝑟 = .6 49
𝑡 = 8.5 52 4.728
̂𝑌 = 9734.2 3782.2𝑋 + 28 5𝑋
𝑡 = 3.37 5 6.6 7 2.97 2 = .77 6
8. Based on our discussion of individual and joint tests of hypothesis based, respectively, on the t
and F tests, which of the following situations are likely?
a) Reject the joint null on the basis of the F statistic, but do not reject each separate null on the
̂
𝑠𝑎𝑙𝑎𝑟𝑦 = 4.32 + .28 𝑠𝑎𝑙𝑒𝑠 + . 74𝑟𝑜𝑒 + . 24𝑟𝑜𝑠
𝑠𝑒 = .32 . 35 . 4 . 54 = .283
Where salary = salary of CEO
sales = annual firm sales
roe = return on equity in percent
ros = return on firm‘s stock and where figures in the parentheses are the estimated standard
errors.
a. Interpret the preceding regression taking into account any prior expectations that you may have
about the signs of the various coefficients.
b. Which of the coefficients are individually statistically significant at the 5 percent level?
c. What is the overall significance of the regression? Which test do you use? And why?
d. Can you interpret the coefficients of ―roe‖ and ―ros‖ as elasticity coefficients? Why or why not?
10. Return to the child mortality example that we have discussed several times. In regression (4.7.3)
we regressed child mortality (CM) on per capita GNP (PGNP) and female literacy rate (FLR).
Now we extend this model by including total fertility rate (TFR). The data on all these variables
are already given in Table 4.5. We reproduce regression (4.7.3) and give results of the extended
regression model below:
Model I: ̂ = 263.64 6 . 56 2.23 6
43 Bish200821@[Link] | Prepared by: Bishaw A. (MSc) : Econometrics Module
Debre Tabor University, CAES, Department of Agricultural Economics
4.7.3𝑎
𝑠𝑒 = .5932 . 9 .2 99 = .7 77
Model II: ̂ = 68.3 67 . . 55 .768 + 2.8686
4.7.3𝑏
𝑠𝑒 = 32.89 6 . 8 .248 = .7474
a. How would you interpret the coefficient of TFR? A priori, would you expect a positive or
negative relationship between CM and TFR? Justify your answer.
b. Have the coefficient values of PGNP and FR changed between the two equations? If so, what
may be the reason(s) for such a change? Is the observed difference statistically significant?
Which test do you use and why?
c. How would you choose between models 1 and 2? Which statistical test would you use to
answer this question? Show the necessary calculations.
d. We have not given the standard error of the coefficient of TFR. Can you find it out? (Hint:
Recall the relationship between the t and F distributions.)
11. Given the table below which gave data on advertising impressions retained and advertising
expenditure for a sample of 21 firms. Decide on an appropriate model about the relationship
between impressions and advertising expenditure. Letting Y represent impressions retained and
X the advertising expenditure, the following regressions were obtained:
Model I: 𝑌̂ = 22. 63 + .363𝑋
𝑠𝑒 = 7. 89 . 97 𝑟 = .424=
Model II: 𝑌̂ = 7. 59 + . 847𝑋 + . 4 𝑋
𝑠𝑒 = 9.986 .3699 . 9 = .53
a. Interpret both models.
b. Which is a better model? Why?
c. Which statistical test(s) would you use to choose between the two models?
d. Are there ―diminishing returns‖ to advertising expenditure, that is, after a certain level of
advertising expenditure (the saturation level) it does not pay to advertise? Can you find out
what that level of expenditure might be? Show the necessary calculations.
Table 4.9: impact of advertising expenditure
2 99.6 74.1
3 11.7 19.3
4 21.9 22.9
5 60.8 82.4
6 78.6 40.1
7 92.4 185.9
8 50.7 26.9
9 21.4 20.4
10 40.1 166.2
11 40.8 27.0
12 10.4 45.6
13 88.9 154.9
14 12.0 5.0
15 29.2 49.7
16 38.0 26.9
17 10.0 5.7
18 12.3 7.6
19 23.4 9.2
20 71.1 32.4
21 4.4 6.1
12. Suppose the demand for commodity; its price and the price of related goods at different market
outlet recoded in the table 1 as follows:
Table 4.10: The Demand for Commodity Y Depend on its own price (X1) and the Price of Related Goods(X2)
Y X1 X2 Observation
64 57 8 1
71 59 10 2
53 49 6 3
67 62 11 4
55 51 8 5
58 50 7 6
77 55 10 7
57 48 9 8
a) a). Estimate the coefficient of the parameter in the econometric demand model of commodity Y given
as: 𝑌̂ = 𝛽̂ + 𝛽̂ 𝑋 + 𝛽̂ 𝑋 + 𝑢̂
b) Estimate the variance and standard error of the model as well as the parameters?
c) Compute TSS, RSS and ESS?
d) Compute R2 and F- statistic for the model?
e) Test whether the coefficient for price of related goods becomes zero at 5% significance level?
f) Test whether the coefficient for both X1and X2 are becomes zero?
g) Test whether the coefficient for both X1and X2 are equal?
13. suppose the demand for tea and its own price as well as price of coffee is shown is given as
follows in different coffee house owners
Table 4.11: demand for tea and its price
Y X1 X2
-3.7 3 8
3.5 4 5
2.5 5 7
11.5 6 3
5.7 2 1
a) Estimate the coefficient of the parameter in the econometric demand model of commodity Y given
as: 𝑌̂ = 𝛽̂ + 𝛽̂ 𝑋 + 𝛽̂ 𝑋 + 𝑢̂
b) Estimate the variance and standard error of the model as well as the parameters?
c) Compute TSS, RSS and ESS?
d) Compute R2 and F- statistic for the model?
e) Test whether the coefficient for price of related goods becomes zero at 5% significance level?
f) Test whether the coefficient for both X1and X2 are becomes zero?
g) Test whether the coefficient for both X1and X2 are equal?
5.1 Non-Normality
The normality assumption for multiple regressions is one of the most misunderstood in all of
statistics. In multiple regressions, the assumption requiring a normal distribution applies only to the
residuals, not to the independent variables as is often believed. Perhaps the confusion about this
assumption derives from difficulty understanding what the residuals are simply put, the residuals are
the error in the relationship between the independent variables and the dependent variable in a
regression model. Each case in the sample has a residual value that represents the difference in the
observed and predicted values produced by a regression equation. It is the distribution of the
residuals or noise for all cases in the sample that should be normally distributed.
Classical normal linear regression (CNLR) assumes that each is distributed normally
𝑤𝑖𝑡
Mean = E (Ui) = 0
Variance = E (Ui2) =
Cov (Ui, Uj) = E (Ui, Uj) = 0 (i # j)
Note: For two normally distributed variables, the zero covariance or correlation means independence
of them, so Ui and Uj are not only uncorrelated but also independently distributed. Therefore
NID (0, ) is Normal and independently distributed(Note: ID stands for Normally and
Independently Distributed).
distribution of their sum tends to a normal distribution as the number of such variables increase
indefinitely.1 It is the CLT that provides a theoretical justification for the assumption of normality of
𝑢.
2. A variant of the CLT states that, even if the number of variables is not very large or if these
variables are not strictly independent, their sum may still be normally distributed.
3. With the normality assumption, the probability distributions of OLS estimators can be easily
derived because, one property of the normal distribution is that any linear function of normally
distributed variables is itself normally distributed. As we discussed earlier, OLS estimator‘s 𝛽̂ and
𝛽̂ are linear functions of ui. Therefore, if ui are normally distributed, so are 𝛽̂ and 𝛽̂ , which makes
our task of hypothesis testing very straightforward.
4. The normal distribution is a comparatively simple distribution involving only two parameters
(mean and variance); it is very well known and its theoretical properties have been extensively
studied in mathematical statistics. Besides, many phenomena seem to follow the normal distribution.
5. Finally, if we are dealing with a small, or finite, sample size, say data of less than 100
observations, the normality assumption assumes a critical role. It not only helps us to derive the
exact probability distributions of OLS estimators but also enables us to use the t, F, and χ2 statistical
tests for regression models.. As we will show subsequently, if the sample size is reasonably large, we
may be able to relax the normality assumption. A cautionary note: Since we are ―imposing‖ the
normality assumption, it behooves us to find out in practical applications involving small sample
size data whether the normality assumption is appropriate. Later, we will develop some tests to do
just that. Also, later we will come across situations where the normality assumption may be
inappropriate. But until then we will continue with the normality assumption for the reasons
discussed previously.
Violations of normality create problems for determining whether model coefficients are significantly
different from zero and for calculating confidence intervals for forecasts. Sometimes the error
distribution is "skewed" by the presence of a few large outliers. Since parameter estimation is based
on the minimization of squared error, a few extreme observations can exert a disproportionate
influence on parameter estimates. Calculation of confidence intervals and various significance tests
for coefficients are all based on the assumptions of normally distributed errors. If the error
distribution is significantly non-normal, confidence intervals may be too wide or too narrow.
Technically, the normal distribution assumption is not necessary if you are willing to assume the
model equation is correct and your only goal is to estimate its coefficients and generate predictions
in such a way as to minimize mean squared error. The formulas for estimating coefficients require
no more than that, and some references on regression analysis do not list normally distributed errors
among the key assumptions. But generally we are interested in making inferences about the model
and/or estimating the probability that a given forecast error will exceed some threshold in a
particular direction, in which case distributional assumptions are important. Also, a significant
violation of the normal distribution assumption is often a "red flag" indicating that there is some
other problem with the model assumptions and/or that there are a few unusual data points that should
be studied closely and/or that a better model is still waiting out there somewhere.
There are four main sources of non –normality problem in our data.
Extreme values or outliers: examples the income data in developing countries may have
extreme values or have outliers. That is some observations have minimum value but the other
have maximum value. In this case the data have large range.
Two or more process overlapping: happens when the distribution of data which is the
combination of two or more normal distributions. Therefore we have bimodal or multimodal
distribution.
Insufficient data: the more data you have the more likely to follow normal distribution. You
should have the data at least from 30 observations. Example, if we increase observation to
50; the data the more likely to follow normal distribution.
Subset of main sample: if the data is the subset of main sample and you have not taken the
indirect data. This is happen because of sampling error.
Data follows same other distribution: no reason every data should be follow normal
distributions. In real world the data may follow other distributions like beta distribution,
logistic distribution, exponential distribution, Poisson distribution, gamma distribution etc.
There are few consequences associated with a violation of the normality assumption, as it does not
contribute to bias or inefficiency in regression models. It is only important for the calculation of p
49 Bish200821@[Link] | Prepared by: Bishaw A. (MSc) : Econometrics Module
Debre Tabor University, CAES, Department of Agricultural Economics
values for significance testing, but this is only a consideration when the sample size is very small.
When the sample size is sufficiently large 2 ), the normality assumption is not needed at all as
the Central Limit Theorem ensures that the distribution of residuals will approximate normality.
When dealing with very small samples, it is important to check for a possible violation of the
normality assumption. This can be accomplished through an inspection of the residuals from the
regression model (some programs will perform this automatically while others require that you save
the residuals as a new variable and examine them using summary statistics and histograms). There
are several statistics available to examine the normality of variables, including skewness and
kurtosis, as well as numerous graphical depictions, such as the normal probability plot.
Unfortunately, the statistics to assess it are unstable in small samples, so their results should be
interpreted with caution. When the distribution of the residuals is found to deviate from normality,
Quantile plot: the quantile plot for residuals also used to detect the problem of normality in the
data.
Normality plot p plot: such as histogram probability plot for residuals also shows whether there is
the problem or not.
What do you if your data are not normal? : Possible solutions include transforming the data,
removing outliers, or conducting an alternative analysis that does not require normality (e.g., a
nonparametric regression).
1. Remove outliers: If the observation is large enough we can remove outlier individual data‘s if
they are few in number.
2. Data transformation: use the uniform log transformation or other mathematical change used
across the board to increase normality. This is ok and legitimate to minimize the gap between
outliers and make the distribution normal. For example log transformations are good when the
data are positively skewed. However keep in mind interpretability of transformation.
3. Fit another distribution:
4. Use non-parametric test: if you do not want to deal with transformations you can use non-
parametric tests that do not require data to be normally distributed.
50 Bish200821@[Link] | Prepared by: Bishaw A. (MSc) : Econometrics Module
Debre Tabor University, CAES, Department of Agricultural Economics
5.2 MULTICOLLINEARITY
One of the assumptions of the classical linear regression model (CLRM) is that there is no
perfect Multicollinearity among the regressors included in the regression model. Although the
assumption is said to be violated only in the case of exact Multicollinearity (i.e. an exact linear
relationship among some of the regressors), the presence of Multicollinearity (an approximate
linear relationship among some of the regressors) leads to estimation problems.
Multicollinearity does not depend on any theoretical or actual linear relationship among any of the
regressors; it depends on the existence of an approximate linear relationship in the data set at hand.
Unlike most other estimation problems, this problem is caused by the particular sample available.
The existence of Multicollinearity will affect seriously the parameter estimates. Intuitively, when any
two explanatory variables are changing in nearly the same way, it becomes extremely difficult to
establish the influence of each regressor on the dependent variable separately.
Consider the consumption–income model
𝑜𝑛𝑠𝑢𝑚𝑝𝑡𝑖𝑜𝑛 = 𝛽 + 𝛽 + 𝛽 +𝑢
It may happen that when we obtain data on income and wealth, the two variables may be highly, if
not perfectly, correlated: Wealthier people generally tend to have higher incomes. Thus, although in
theory income and wealth are logical candidates to explain the behavior of consumption expenditure,
in practice (i.e., in the sample) it may be difficult to disentangle the separate influences of income
and wealth on consumption expenditure.
Ideally, to assess the individual effects of wealth and income on consumption expenditure we need a
sufficient number of sample observations of wealthy individuals with low income, and high-income
individuals with low wealth. Although this may be possible in cross sectional studies (by increasing
the sample size), it is very difficult to achieve in aggregate time series work.
In general, the problem of Multicollinearity arises when individual effects of explanatory variables
cannot be isolated and the corresponding parameter magnitudes cannot be determined with the
desired degree of precision. Though it is quite frequent in cross section data as well,
Multicollinearity tends to be more common and more serious problem in time series data.
Building a linear regression model is only half of the work. In order to actually be usable in practice,
the model should conform to the assumptions of linear regression.
51 Bish200821@[Link] | Prepared by: Bishaw A. (MSc) : Econometrics Module
Debre Tabor University, CAES, Department of Agricultural Economics
𝑦 = 𝛽̂ 𝑥 + 𝛽̂ 𝑥 + 𝑢̂
∑𝑥 𝑦 ∑𝑥 ∑𝑥 𝑦 ∑𝑥 𝑥
𝛽̂ =
∑𝑥 ∑𝑥 ∑𝑥 𝑥
Suppose
X 2i = kX1i where k is non-zero constant.
Then
∑𝑥 𝑦 𝑘 ∑𝑥 𝑘∑𝑥 𝑦 𝑘∑𝑥
𝛽̂ = =
∑𝑥 𝑘 ∑𝑥 𝑘 ∑𝑥 𝑥
This is an indeterminate expression. It can also be shown that the expression for 𝛽̂ is
indeterminate.
Recall that 𝛽̂ gives the rate of change in the average value of Y as X1 changes by a unit, holding X2
constant. But if X1 and X2 are perfectly collinear, there is no way X2 can be kept constant: As X1
changes, so does X2 by the factor k. What it means, then, is that there is no way of disentangling the
separate influences of X1 and X2 from the given sample.
Moreover, for a three variable model:
∑𝑥
𝛽̂ =
∑𝑥 ∑𝑥 ∑𝑥 𝑥
∑𝑥
𝛽̂ =
∑𝑥 ∑𝑥 ∑𝑥 𝑥
Substituting 𝑥 = 𝑘𝑥
𝑘 ∑𝑥 ∑𝑥
𝛽̂ = = =
𝑘 ∑𝑥 ∑𝑥 𝑘 ∑𝑥 𝑥
𝑣𝑎𝑟 𝛽̂ =
∑𝑥 𝑟
It is apparent from the above formula that as 𝑟 (which is the coefficient of correlation between X1
and X2) tends towards 1, that is, as collinearity increases, the Variance of the estimator increases.
The same holds for 𝑣𝑎𝑟 𝛽̂ and the 𝛽̂ 𝛽̂ )
Because of consequence (1), the confidence intervals tend to be much wider, leading to the
acceptance of the ―Zero null hypothesis‖ (i.e., the true population coefficient is zero). Because of
consequence (1), the t-ratio of one or more coefficients tends to bestatistically insignificant.
Although the t-ratio of one or more coefficients is statistically insignificant, R2, the overall
measure of goodness of fit, can be very high. The OLS estimators and their standard errors can
be sensitive to small changes in the data.
Note that Multicollinearity is a question of degree and not of a kind. It should also be noted that
since Multicollinearity refers to the condition of the explanatory variables that are assumed to be
nonstochastic, it is a feature of the sample and not of the population. Therefore, we do not ―test for
Multicollinearity‖ but can, if we wish, measure its degree in any particular sample. The following
are some rules of thumb and formal rules to detection of Multicollinearity.
a) High R2 but few significant t-ratios. If R2 is high, say in excess of 0.8, the F-test in most cases
will reject the hypothesis that the partial slope coefficients are simultaneously equal to zero, but
the individual t tests will show that none or very few of the partial slope coefficients are
statistically different from zero.
b) High pair-wise correlation among regressors. If the pair-wise correlation coefficient among
two regressors is high, say in excess of 0.8, then Multicollinearity is a serious problem.
c) Auxiliary Regression. Since Multicollinearity arises because one or more of the regressors are
exact or approximately linear combinations of the other regressors, one way of finding out which
X variable is related to other X variables is to regress each Xi on the remaining X variables and
compute the corresponding R2. As a rule of thumb, Multicollinearity may be a troublesome
problem only if the R2 obtained from an auxiliary regression is greater than the overall R2 (that
obtained from the regression of Y on all the regressors).
d) Tolerance (TOL) and variance inflation factor (VIF): For a linear regression with two
explanatory variables (X1 and X2) =
VIF shows how the variance of an estimator is inflated by the presence of Multicollinearity.
12
As 𝑟 approaches 1, the VIF approaches infinity. That is, as the extent of collinearity increases,
the variance of an estimator increases, and in the limit it can become infinite. If there is no
collinearity between X1 and X2, VIF will be 1.
𝑣𝑎𝑟(𝛽̂ ) = = 𝐼
∑𝑥 𝑟 ∑𝑥
𝑣𝑎𝑟(𝛽̂ ) = = 𝐼
∑𝑥 𝑟 ∑𝑥
Similarly for k-variables model
𝑣𝑎𝑟(𝛽̂ ) = = 𝐼
∑𝑥 ( ) ∑𝑥
It is also possible to use TOL as the measure of Multicollinearity in the view of its intimate
connection with VIF
= =
𝐼
The larger the value of the VIFi, the more ―troublesome‖ or collinear the variable Xi. As a rule of
Thumb, if the VIF of a variable exceeds 10, which will happen if R 2 exceeds 0.90, that variable is
i
said to be highly collinear. In other words, the closer is TOLi to zero, the greater the degree of
collinearity of that variable with the other regressors. On the other hand, the closer TOLi is to 1, the
greater the evidence that Xi is not collinear with the other regressors.
The existence of Multicollinearity in a data set does not necessarily mean that the coefficient
estimators in which the researcher is interested have unacceptably high variance. Because
Multicollinearity is essentially a sample problem there are no infallible guides. However one can
try the following rules of thumb, the success of which depends on the severity of the collinearity
problem.
a) Obtain more data: - Because the Multicollinearity is essentially a data problem, additional data
that do not contain the Multicollinearity feature could solve the problem. For example, in the
2
three variable models we saw that now as the sample size increases, x will generally
increases. Thus, for a given r12 the variance of ̂ 1 decreases (thus, the standard error
decreases), which will enable us to estimate ̂ 1 more precisely.
b) Transformation of variables: - In time series analysis, one reason for high Multicollinearity
between two variables is that over time both variables tend to move in the same direction. One
way of minimizing then dependence is to transform the variables.
Suppose𝑌 = 𝛽 + 𝛽 𝑋 + 𝛽𝑋 +
This relation must also hold at time t-1 because the origin of time is arbitrary anyway. Therefore
we have 𝑌 = 𝛽 + 𝛽𝑋 + 𝛽𝑋 +
Subtracting this from the above gives
𝑌 𝑌 = +𝛽 𝑋 𝑋 + 𝛽 𝑋 𝑋 +
This is known as the first difference form because we run the regression, not on the original
variables, but on the difference of successive values of the variables. The first difference regression
model often reduces the severity of Multicollinearity. Although the levels of X1 and X2 may be
highly correlated, there is no a priori reason to believe that their difference will also be highly
correlated.
Another commonly used transformation in practice is the ratio transformation.
Consider the model:
𝑌 = 𝛽 + 𝛽𝑋 + 𝛽𝑋
where Y is consumption expenditure in real dollars, X1 is GDP, and X2 is total population. Since GDP
and population grow over time, they are likely to be correlated. One ―solution‖ to this problem is to
56 Bish200821@[Link] | Prepared by: Bishaw A. (MSc) : Econometrics Module
Debre Tabor University, CAES, Department of Agricultural Economics
5.3 HETEROSCEDASTICITY
The assumption of homoscedasticity (or constant variance) about the random variable u is that its
probability distribution remains the same over all observations of X, and in particular that the
variance of each Ui is the same for all values of the explanatory variable. That is, the variation of
each ui around its zero mean does not depend on the value of X.
Symbolically we have Var (ui) = E {(ui – E (ui)}2 = E (u 2) =
If the above condition is not satisfied in any particular case, we say that the ui‟s are heterosckedastic.
That is, Var (ui) =
The problem of heteroskedasticity is more serious in cross section data rather than time series data.
Suppose we have a cross-section sample of family budget from which we want to measure the
savings function. That means Saving = f(income). In this case, the assumption of constant variance
of the ui‟s is not appropriate, because high-income families show a much greater variability in their
saving behavior than do low income families. Families with high income tend to stick to a certain
57 Bish200821@[Link] | Prepared by: Bishaw A. (MSc) : Econometrics Module
Debre Tabor University, CAES, Department of Agricultural Economics
standard of living and when their income falls they cut down their savings rather than their
consumption expenditure. But this is not the case in low income families. Hence, the variance of
ui‘s increases as income increases.
If the assumption of homoscedastic disturbance is not fulfilled we have the following consequences:
Heteroskedasticity does not destroy the unbiasedness and consistency properties of OLS
estimators.
The OLS estimates do not have the minimum variance property in the class of unbiased
As in the case of Multicollinearity, there are no hard-and-fast rules for detecting heteroskedasticity,
only a few rules of thumb.
A) Informal method
Nature of the problem: As a matter of fact, in cross-sectional data involving heterogeneous units,
heteroskedasticity may be the rule rather than the exception. For example, in a cross-sectional
analysis involving the investment expenditure in relation to sales, rate of interest, etc.,
heteroskedasticity is generally expected if small, medium and large-size firms are sampled
together.
Visual Inspection of Residuals / graphical method
This is a postmortem approach when there is no a priori information as the existence of
heteroskedasticity. Hence, this approach examines whether the error term depicts some systematic
pattern or not. To this end, the residuals are plotted against the dependent or independent variable to
which it is suspected the disturbance variance is related. Although are not the same thing as u 2, they
can be used as proxies especially if the sample size is sufficiently large.
B) Formal methods
Park Test: park formalizes the general method by suggesting that is some functions of
explanatory variable 𝑋
The functional form he suggest is 𝑣𝑎𝑟 𝑢 = + 𝑋 𝑒 which could be written as logarithmic
form as 𝑙𝑛 = 𝑙𝑛 + 𝛽𝑙𝑛𝑋 + 𝑣
Where 𝑣 is stochastic disturbance term
Since is generally not known, park suggest using 𝑢̂ as a proxy and running the following
regression
𝑙𝑛𝑢̂ = 𝑙𝑛 + 𝛽𝑙𝑛𝑋 + 𝑣
𝑙𝑛𝑢̂ = + 𝛽𝑙𝑛𝑋 + 𝑣
If 𝛽 turns out to be statistically significant; it would suggest that heteroskedasticity is present in the
59 Bish200821@[Link] | Prepared by: Bishaw A. (MSc) : Econometrics Module
Debre Tabor University, CAES, Department of Agricultural Economics
data. If it turns out to be insignificant, we may accept the assumption of homoscedasticity. The park
test is thus a two-stage procedure. In the first stage we run the OLS regression disregarding the
heteroskedasticity question.
We obtain 𝑢̂ from this regression and then in the second stage we run the regression
𝑙𝑛𝑢̂ = 𝑙𝑛 + 𝛽𝑙𝑛𝑋 + 𝑣
Example: Consider a relationship between Compensation (Y) and Productivity (X). To illustrate the
Park approach, the following regression function is used.
𝑌 =𝛽 +𝛽 𝑋 +𝑢
𝑌 = 992.35 + .23𝑋 + 𝑢
𝑠𝑒 = 936.48 . 99 𝑟 = .44
𝑡 = 2. 3 2.33
Suppose that the residuals obtained from the above regression were regressed on Xi giving the
following results.
𝑙𝑛𝑢̂ = 35.82 2.8 𝑙𝑛𝑋
𝑠𝑒 = 38.2 4.22 𝑟 = .46
𝑡= .93 .67
In the above result, the coefficient of lnXi is not significant. That is, there is no statistically
significant relationship between the two variables. Following the Park test, one may conclude that
there is no heteroskedasticity in the error variance.
Although empirically appealing, the Park test has some problems. For instance, the error term,Vi may
not satisfy the OLS assumptions and may itself be heterosckedastic. Nonetheless, as a strictly
exploratory method, one may use the Park test.
Spearman’s Rank Correlation Test
This test requires calculating rank correlation where its coefficient can be used to detect
heteroskedasticity. The rank correlation coefficient is given by
∑𝑑
𝑟 = 6[ ]
𝑛 𝑛
Where 𝑑 = difference in the ranks assigned to two different characteristics of the 𝑖 individual or
phenomenon and n = number of individuals or phenomena ranked. The steps required in this test are
stated as follows.
Assume
𝑌 =𝛽 +𝛽 𝑋 +𝑢
Step 1: Fit the regression to the data on Y and X and obtain the residuals 𝑢̂
Step 2: Ignoring the sign of 𝑢̂ that is taking their absolute value 𝑢̂ rank both 𝑢̂ and Xi (or 𝑌̂ )
according to an ascending or descending order and compute the Spearman‘s rank correlation
coefficient
Step 3: Assuming that the population rank correlation coefficient ƥs is zero and n > 8, the
significance of the sample rs can be tested by the t test as follows:
𝑟 √𝑛 2
𝑡= 𝑤𝑖𝑡 𝑑𝑒𝑔𝑟𝑒𝑒 𝑜𝑓 𝑓𝑟𝑒𝑒𝑑𝑜𝑚 𝑛 2
√ 𝑟
If the computed t value exceeds the critical t value, we may accept the hypothesis of
heteroskedasticity; otherwise we may reject it. If the regression model involves more than one X
variables, rs can be computed between | û i | and each of the X variable separately and can be
tested for statistical significance by the t-test given.
Example To illustrate the rank correlation test considers the regression
𝑌 = 𝛽 + 𝛽 𝑋 + 𝑢 . Suppose 10 observations are used to this equation. The following table makes
use of the rank correlation approach to test the hypothesis of heteroskedasticity. Notice that column 6
and 7 put rank of |û i | and Xi in an ascending order.
Table 5.1 Rank Correlation Test of Heteroskedasticity
Obs Rankof d (difference between d2
Y X = Xi the two ranking)
1 12.4 12.1 11.37 1.03 9 4 5 25
2 14.4 21.4 15.64 1.24 10 9 1 1
3 14.6 18.4 14.4 0.20 4 7 -3 9
4 16 21.7 15.78 0.22 5 10 -5 25
5 11.3 12.5 11.56 0.26 6 5 1 1
6 10.0 10.4 10.59 0.59 7 2 5 25
7 16.2 20.8 15.37 0.83 8 8 0 0
8 10.4 10.2 10.50 0.10 3 1 2 4
9 13.1 16.0 13.16 0.06 2 6 -4 16
10 11.3 12.0 11.33 0.03 1 3 -2 4
Total 0 110
Then,
∑𝑑
𝑟 = 6[ ]= 6[ ] = .33
𝑛 𝑛
and
.33√ 2
𝑡= = .99
√ .
Note that for 8 (=10-2) df, this t-value is not significant even at the 10% level of significance. Thus,
there is no evidence of systematic relationship between the explanatory variable and the absolute
value of the residuals, which might suggest that there is no heteroskedasticity.
The Goldfield – Quandt Test
This popular method is applicable if one assumes that the heterosckedastic variance, σi2 is positively
related to one of the explanatory variables in the regression model. The test is commonly applicable
to large samples. The observation must be at least twice as many as the parameters to be estimated.
The test assumes normality and serially independent disturbance term, Ui‟s. Consider the
following:
𝑌 = 𝛽 + 𝛽 𝑋 + 𝛽 𝑋 … +𝛽 𝑋 + 𝑢
Note from the F- table in the appendix that the critical F value for 11 numerators and 11
denominator df at the 5% level is 2.82. Since the estimated F* value exceeds the critical value, we
𝑌 = 𝛽 + 𝛽 𝑋 + 𝛽 𝑋 +𝛽 𝑋 + 𝑢 5.2.3
Steps
the null, we conclude that there is homoscedasticity. We can also use an F test of thishypothesis; both
tests have asymptotic justification.
Heteroskedasticity does not destroy the unbiasedness and consistency properties of the OLS
estimators, but they are no longer efficient, not even asymptotically (i.e., large sample size). This
lack of efficiency makes the usual hypothesis testing procedure of dubious value. Therefore, remedial
measures are clearly called for. There are two approaches to remediation: when σ2 is known and when
σi 2 is not known.
𝑌 = 𝛽 + 𝛽 𝑋 +𝑢
65 Bish200821@[Link] | Prepared by: Bishaw A. (MSc) : Econometrics Module
Debre Tabor University, CAES, Department of Agricultural Economics
What is the purpose of transforming the original model? To see this, notice the following feature of
the transformed error term ui*
𝑢
𝑣𝑎𝑟 𝑢 = 𝑢 = ( ) = 𝑢 𝑠𝑖𝑛𝑐𝑒 𝑖𝑠 𝑘𝑛𝑜𝑤𝑛
= 𝑠𝑖𝑛𝑐𝑒 𝑢 =
=
Which is a constant. That is, the variance of the transformed disturbance term ui* is now
homoscedastic. Since we are still retaining the other assumptions of the classical model, ui* is
homoscedastic suggests that if we apply OLS to the transformed model it will produce estimators that
are BLUE. In short, the estimated β0* and β1* are now BLUE and not the OLS estimators
𝛽̂ 𝑎𝑛𝑑 𝛽̂
To obtain GLS estimators, we minimize
∑ 𝑢̂ =∑ 𝑌 𝛽 𝑋 𝛽 𝑋
𝑢̂ 𝑌 𝛽𝑋 𝛽𝑋
∑ =∑ 5.3.
∑ 𝑤 𝑢̂ =∑𝑤 𝑌 𝛽 𝑋 𝛽 𝑋
Where 𝑤 =
Thus, in GLS we minimize a weighted sum of residual squares with wi = 1/σ 2 actingi as the weights,
but in OLS we minimize non-weighted or (what amounts to the same thing) equally weighted RSS.
In GLS, the weight assigned to each observation is inversely proportional to its σi, that is,
observations coming from a population with larger σi will get relatively smaller weight and those
from a population with smaller σi will get proportionately larger weight in minimizing the RSS.
When σi2 Is Not Known: White’s Heteroskedasticity-Consistent Standard Errors
If true σi2 are known, we can use the WLS method to obtain BLUE estimators. Since the true σi2 are
rarely known, there is a way of obtaining consistent (in the statistical sense) estimates of the
variances and covariances of OLS estimators even if there is heteroskedasticity. White has shown that
this estimate can be performed so that asymptotically valid (i.e., large-sample) statistical inferences
can be made about the true parameter values. White‘s heteroskedasticity-corrected standard errors
5.4 AUTOCORRELATION
Var(ϵt) = σ2
Cov (ϵt, ϵt+s) = 0 where subscript „s‟ represent the exact period of lag.
The above specification is of first order because the regression of ut is on itself lagged one period
(where the coefficient ρ is the first order coefficient of autocorrelation). Note that the above
specification postulates that the movement or shift in ut consists of two parts: a part ρut-1, which
accounts for systematic shift, and the other ϵt which is purely random.
Relationships between ut‟s can be shown as:
Cov (ut, ut-1) = E[(ut – E(ut) (ut-1 – E(ut-1)]
= E[ut ut-1]
by substituting ut = ρut-1 + ϵt we obtain:
= E[(ρut-1 + ϵt) ut-1]
= ρE[u2t-1] + E[ϵt ut-1]
Note that E(ϵt) = 0 thus E(ϵt Ut-1) = 0
Since with the assumption of homoscedasticity (i.e., constant variance) Var(ut) = Var (ut-1) = σ2
The result would be Cov (ut, ut-1) = ρ σ2
Now, correlation of ut, ut-1 is given by
𝑐𝑜𝑣 𝑢 𝑢
𝑐𝑜𝑟𝑟 𝑢 𝑢 = = = = 5.4.
√𝑣𝑎𝑟 𝑢 𝑣𝑎𝑟 𝑢 𝑣𝑎𝑟 𝑢
Where
Hence, ρ(rho) is simple correlation of the successive errors of the original model.
Note that when ρ > 0 successive errors are positively correlated and when ρ < 0 successive errors are
negatively correlated. It can be shown that corr (Ut, Ut-s) = ρs (where s represents the exact period of
lag). It implies that the correlation (be it negative or positive) between any two period diminishes as
time goes by; i.e., as s increases.
Inertia: A salient feature of most economic time series is inertia, or sluggishness. As is well known,
time series such as GNP, price indices, production, employment, and unemployment exhibit
(business) cycles. Starting at the bottom of the recession, when economic recovery starts, most of
these series start moving upward. In this upswing, the value of a series at one point in time is greater
than its previous value. Thus, there is a ―momentum‟‟ built into them, and it continues until something
happens (e.g., increase in interest rate or taxes or both) to slow them down. Therefore, in regressions
involving time series data, successive observations are likely to interdependent.
Data manipulation: published data often undergo interpolation or smoothing, procedures that
average true disturbances over successive time periods.
Specification bias
Specification Bias: Excluded Variables Case. In empirical analysis the researcher often starts with
a plausible regression model that may not be the most ―perfect‟‟ one. After the regression analysis,
the researcher does the postmortem to find out whether the results accord with a priori expectations.
For example, suppose we have the following demand model:
Yt = β0 + β1X1t + β2X2t + β3X3t + ut
However, for some reason we run the following regression:
Yt = β0 + β1X1t + β2X2t + vt
Now if the first model is the ―correct‟‟ model, running the second is tantamount to letting vt = β3X3t
+ ut. To the extent that X3 affects Yt, the error term v will reflect a systematic pattern, thus creating
(false) autocorrelation. A simple test of this would be to run both models and see whether
autocorrelation, if any, observed in the second model, disappears when the first model is run.
Specification Bias: Incorrect Functional Form. Suppose the ―true‟‟ or correct model in a cost-
output study is as follows:
Marginal costi = β0 + β1 outputi + β2 output 2 + ui
But we fit the following model:
Marginal costi = α0 + α1 outputi + vi
Because the disturbance term vi is, in fact, equal to output2 + ui , it will catch the systematic effect of
the output2 term on marginal cost. In this case, vi will reflect autocorrelation because of the use of
an incorrect functional form.
Lags. For instance, in a time series regression of consumption expenditure on income, it is not
uncommon to find that the consumption expenditure in the current period depends, among other
things, on the consumption expenditure of the previous period. That is, (Consumption)t= β0 + β1(
income)t + β2 (consumption)t−1 + ut
A regression like this is known as auto regression because one of the explanatory variables is the
lagged value of the dependent variable. The rationale is consumers do not change their consumption
habits readily for psychological, technological, or institutional reasons. Now if we neglect the lagged
consumption in the model, the resulting error term will reflect a systematic pattern due to the
influence of lagged consumption on current consumption.
As in the case of heteroskedasticity, in the presence of autocorrelation the OLS estimators are still
linear unbiased as well as consistent and asymptotically normally distributed, but they are no longer
efficient (i.e., minimum variance). As a consequence, the usual t, F, and χ2 tests cannot be
legitimately applied.
The prediction based on ordinary least squares estimate will be inefficient with auto correlated
errors. This is because of larger variance as compared with predictions based on estimates obtained
from other econometric techniques.
Autocorrelation is potentially a series problem. Hence, it is essential to find out whether Auto-
correlation exists in a given situation. Since the population disturbances Ut, cannot be observed
directly, we use its proxy, the residual Û t which can be obtained from the usual OLS procedure. The
examination of Û t can provide useful information not only about autocorrelation but also about
As the above figure reveals, most of the residuals are bunched in the first and the third quadrants
suggesting very strongly that there is positive correlation in the residuals. However, the graphical
method is essentially subjective or qualitative in nature. There are quantitative tests that can be used
to supplement the purely qualitative approach.
Durbin-Watson d Test
The most celebrated test for detecting serial correlation is the one developed by Durbin and Watson.
It is popularly known as the Durbin-Watson d-Statistic and it is defined as
2
∑ ̂ ̂
= ∑̂
5.4.2
Which is simply the ratio of the sum of squared differences in successive residuals to the residual
sum of squares, RSS. Note that in the numerator of the d statistic the number of observations is n-1
because one observation is lost in taking successive differences.
The proof of d-statistic is as follows
∑ (̂ +̂ + 2̂ ̂ ) ∑ ̂ +∑ ̂ +∑ 2̂ ̂
= = 5.4.3
∑ 𝑢̂ ∑ 𝑢̂
However for large samples ∑ ̂ ∑ ̂ and ∑ 𝑢̂ are approximately equal there fore it
can be written as
2 ∑̂ 2 ∑̂ ̂ 2 ∑̂ ̂
= 2[ ]=2 ̂
∑ 𝑢̂ ∑ 𝑢̂ ∑ 𝑢̂
0 dL dU 2 4-dU 4-dL 4
include them also) and ût -1 , ût -2 , . . . , ût -p , where the latter are the lagged values of
the estimated residuals in step 1. Note that to run this regression we will have only (n −
p) observations.
In short, run the following regression:
𝑢̂ = + 𝑋 + ̂ 𝑢̂ + ̂ 𝑢̂ + ̂ 𝑢̂ 5 .3 .4
+ + ̂ 𝑢̂ +
2
and obtain R from this (auxiliary) regression.
If the sample size is large , Breusch and Godfrey have shown that 𝑛 𝑝
If (n − p) R2 exceeds the critical chi-square value at the chosen level of significance, we reject the
null hypothesis, in which case at least one rho in equation (5.3.4) is statistically significantly
different from zero. That is, there is autocorrelation.
A drawback of the BG test is that the value of p, the length of the lag, cannot be specified apriori.
With time series data, auto correlated residuals are often indications of some error in the way we have
specified the regression equation than genuine autocorrelation in the disturbances. Mostly, positive
autocorrelations in economic data are caused by omission of relevant variables.
Incorrect functional form may also be the cause for auto correlated residuals.
Therefore, we should find out if the autocorrelation is pure autocorrelation and not the result of mis-
specification of the model. If the source of the problem is suspected to be due to omission of
important variables, the remedy is to include those omitted variables. Besides if the source of the
problem is believed to be the result of misspecification of the model, then the solution is to
determine the appropriate mathematical form.
If it is pure autocorrelation, one can use appropriate transformation of the original model so that in
the transformed model we do not have the problem of (pure) autocorrelation. As in the case of
heteroskedasticity, we will use some type of generalized least-square (GLS) method. In large
samples, the Newey–West method can be applied to obtain standard errors of OLS estimators that
are corrected for autocorrelation.
applied to the transformed model that satisfies the classical assumptions. Regression of equation (iv)
is known as the generalized, or quasi, difference equation. It involves regressing Y on X, not in the
original form, but in the difference form, which is obtained by subtracting a proportion (= ρ) of the
value of a variable in the previous time period from its value in the current time period. Note that in
this differencing procedure we lose one observation because the first observation has no
antecedent.
When ρ is not known
Although straight forward to apply, the method of generalized difference is difficult to run because,
ρ, population correlation coefficient is rarely known in practice. Therefore, alternative methods need
to be devised.
The First-Difference Method: Since ρ lies between 0 and ±1, one could start from two extreme
positions. At one extreme, one could assume that ρ = 0, that is, no (first-order) serial correlation, and
at the other extreme we could let ρ = ±1, that is, perfect positive or negative correlation.
As a matter of fact, when a regression is run, one generally assumes that there is no autocorrelation
and then lets the Durbin–Watson or other test show whether this assumption is justified. If, however,
ρ = +1, the generalized difference equation in equation (iv) above reduces to the first-difference
equation:
(Yt - Yt-1) = β1(Xt - Xt-1) + (Ut - ρUt-1)
𝑌𝑡𝑡 == β1 ∆Xt +εt
∆𝑌
The first difference transformation may be appropriate if the coefficient of autocorrelation is very
high, say in excess of 0.8, or the Durbin–Watson d is quite low. Strictly speaking, the first-
difference transformation is valid only if ρ = 1. Maddala has proposed this rough rule of thumb: Use
the first difference form whenever d < R2. An interesting feature of the first-difference model is that
there is no intercept in it. Hence, we have to use the regression through the origin.
Computing ρ from Durbin–Watson d Statistic. If we cannot use the first difference transformation
because ρ is not sufficiently close to unity, we have an easy method of estimating it from the
relationship between d and ρ as follows:
𝑑
2
Thus, in reasonably large samples one can obtain rho and use it to transform the data as shown in the
generalized difference equation. However, the relationship between ρ and d may not hold true in
small samples.
76 Bish200821@[Link] | Prepared by: Bishaw A. (MSc) : Econometrics Module
Debre Tabor University, CAES, Department of Agricultural Economics
Estimating ρ from the residuals. If the AR(1) scheme ut = ρut−1 + εt is valid, a simple way to
estimate rho is to regress the residuals ût on ût -1 , for û t the are consistent estimators of the true
ut. That is, we run the following regression: ût = ρ ût-1 + vt
where, ût are the residuals obtained from the original (level form) regression and where v t are the
error term of this regression. Note that there is no need to introduce the intercept term, because the
OLS residuals sum to zero.
Iterative Methods of Estimating ρ: All the methods of estimating ρ explained above provide us
with only a single estimate of ρ. But there are the so-called iterative methods that estimate ρ
iteratively, that is, by successive approximation, starting with some initial value of ρ. Among these
methods are:
Cochrane–Orcutt iterative procedure,
Cochrane–Orcutt two-step procedure,
Durbin two–step procedure, and
Hildreth–Lu scanning or search procedure.
Of these, the most popular is the Cochrane–Orcutt iterative method. One advantage of this method is
that it can be used to estimate not only an AR(1) scheme, but also higher-order autoregressive
schemes. Having obtained the two rhos, one can easily extend the generalized difference equation.
The Cochrane-Orcutt interactive Procedure
This procedure helps to estimate ρ from the estimated residuals ̂ so that information about the
unknown ρ will be derived.
To explain the method, consider the two-variable model
Yi = β0 + β1Xi + Ui ------------------------------------------------------------------------------------- (a)
and assume that Ut is generated by the AR(1) scheme namely
Ut = ρUt-1 + εt --------------------------------------------------------------------------------------------- (b)
Cochrane and Orcutt then recommended the following steps to estimate ρ:
Step 1: Estimate the two variables model by the standard OLS routine and obtain the residuals
Û t
̂ = ̂̂ +
Step 3: Using ̂ obtained from step 2 regressions, run the generalized difference equation asfollows
Step 4: Since a priori it is not known that the ̂ obtained from the regression in step 2 is the best
estimate of substitute the values of 𝛽̂ * and 𝛽̂ * obtained from the regression in step 3
0 1
into the original regression (a) and obtain the new residuals, say ̂ as
̂ =𝑌 𝛽̂ 𝛽̂ 𝑋
Note that this can be easily computed since Yt , Xt , 𝛽̂ 𝑎𝑛𝑑 𝛽̂ are all known.
Step 5: Now estimate this regression
̂ = ̂̂ ̂ +𝑤
Where ̂̂ is the second round estimate of .
Since we do not know whether this second round estimate ̂̂ is the best estimate of ρ, we can go
into the third estimate, and so on. That is why the Cochrane-Orcutt method is said iterative. But how
long should we go on? The general procedure is to stop carrying out iterations when the successive
estimates of ρ converges. Thus, we select that chosen ρ to transform the model and apply a kind of
GLS estimation that minimizes the problem of autocorrelation.
Note that:
1. Since the OLS estimators are consistent despite autocorrelation, in large samples, it makes little
difference whether we estimate ρ from the Durbin–Watson d, or from the regression of the
residuals in the current period on the residuals in the previous period, or from the Cochrane–
Orcutt iterative procedure because they all provide consistent estimates of the true ρ.
2. The various methods discussed above are basically two-step methods. In step 1 we obtain an
estimate of the unknown ρ and in step 2 we use that estimate to transform the variables to
estimate the generalized difference equation, which is basically GLS. But since we use ̂
instead of the true ρ, all these methods of estimation are known in the literature as feasible
GLS (FGLS) or estimated GLS (EGLS) methods.
3. It is important to note that whenever we use an FGLS or EGLS method to estimate the
parameters of the transformed model, the estimated coefficients will not necessarily have the
usual optimum properties of the classical model, such as BLUE, especially in small samples. In
short, whenever we use an estimator in place of its true value, the estimated OLS coefficients
may have the usual optimum properties asymptotically, that is, in large samples. Also, the
conventional hypothesis testing procedures are, strictly speaking, valid asymptotically. In small
samples, therefore, one has to be careful in interpreting the estimated results.
4. In using EGLS, if we do not include the first observation (as was originally the case with the
Cochrane Orcutt procedure), not only the numerical values but also the efficiency of the
estimators can be adversely affected, especially if the sample size is small and if the regressors
are not strictly speaking nonstochastic. Therefore, in small samples it is important to keep the
first observation à la Prais Winsten.
The Newey–West method of correcting the OLS standard errors
Instead of using the FGLS methods, we can still use OLS but correct the standard errors for
autocorrelation by a procedure developed by Newey and West. This is an extension of White‘s
heteroskedasticity-consistent standard errors. The corrected standard errors are known as HAC
(heteroskedasticity- and autocorrelation-consistent) standard errors or simply as Newey– West
standard errors.
This method is strictly speaking valid in large samples and may not be appropriate in small samples.
Therefore, if a sample is reasonably large, one should use the Newey–West procedure to correct
OLS standard errors not only in situations of autocorrelation only but also in cases of
heteroskedasticity, for the HAC method can handle both, unlike the White method, which was
designed specifically for heteroskedasticity.
Conclusion
Violating Multicollinearity does not impact prediction, but can impact inference. For example, p-
values typically become larger for highly correlated covariates, which can cause statistically
significant variables to lack significance.
Violating linearity can affect prediction and inference. We saw that prediction and precision in
estimating coefficients were only hindered slightly. However, these things will be exacerbated when
stronger levels of non-linearity are unaccounted for.
The no endogeneity assumption was violated in the model due to an omitted variable. This created
biased coefficient estimates, which lead to misleading conclusions. Prediction was also poor since
the omitted variable explained a good deal of variation in housing prices.
This simulation gives a flavor of what can happen when assumptions are violated. Depending on a
multitude of factors (i.e. variance of residuals, number of observations, etc.), the model‘s ability to
predict and infer will vary. Of course, it‘s also possible for a model to violate multiple assumptions.
79 Bish200821@[Link] | Prepared by: Bishaw A. (MSc) : Econometrics Module
Debre Tabor University, CAES, Department of Agricultural Economics
Intrinsically Linear and Intrinsically Nonlinear Regression Models: When we started our
discussion of linear regression models in Chapter 3, we stated that our concern in this book is
basically with models that are linear in the parameters; they may or may not be linear in the
variables. On the other hand, if a model is nonlinear in the parameters it is a nonlinear (in-the-
parameter) regression model whether the variables of such a model are linear or not.
However, one has to be careful here, for some models may look nonlinear in the parameters but are
inherently or intrinsically linear because with suitable transformation they can be made linear-in-
the-parameter regression models. But if such models cannot be linearized in the parameters, they are
called intrinsically nonlinear regression models. From now on when we talk about a nonlinear
regression model, we mean that it is intrinsically nonlinear. For brevity, we will call them NLRM.
Example: Consider now the famous Cobb–Douglas (C–D) production function. Letting Y =
output, X1 = labor input, and X2= capital input, we will write this function in three different ways:
𝑌 =𝛽 𝑋 𝑋 𝑒 6. .
𝑜𝑟 𝑙𝑛𝑌 = 𝑙𝑛𝛽 + 𝛽 𝑙𝑛𝑋 + 𝛽 𝑙𝑛𝑋 + 𝑢 6. .2𝑎
Where = 𝑙𝑛𝛽 . Thus in this format the C–D function is intrinsically linear.
𝑜𝑟 𝑙𝑛𝑌 = + 𝛽 𝑙𝑛𝑋 + 𝛽 𝑙𝑛𝑋 + 𝑢 6. .2𝑏
𝑌 =𝛽 𝑋 𝑋 𝑢 6. .3
𝑜𝑟 𝑙𝑛𝑌 = + 𝛽 𝑙𝑛𝑋 + 𝛽 𝑙𝑛𝑋 + 𝑙𝑛𝑢 6. .4
But now consider the following version of the C–D function:
𝑌 =𝛽 𝑋 𝑋 +𝑢 6. .5
As we just noted, C–D versions (6.1.2a) and (6.1.2b) are intrinsically linear (in the parameter)
regression models, but there is no way to transform (6.1.5) so that the transformed model can be
made linear in the parameters. Therefore, (6.1.4) is intrinsically a nonlinear regression model.
Another well-known but intrinsically nonlinear function is the constant elasticity of substitution
(CES) production function of which the Cobb– Douglas production is a special case.
Example: 𝑌 = + 6. .6
Where Y = output, K = capital input, L = labor input, A = scale parameter, δ = distribution parameter
(0 < δ < 1), and β = substitution parameter (β ≥ −1). No matter in what form you enter the stochastic
error term ui in this production function, there is no way to make it a linear (in parameter) regression
model. It is intrinsically a nonlinear regression model.
Therefore, although we can apply the method of least squares to estimate the parameters of the
nonlinear regression models, we cannot obtain explicit solutions of the unknowns. Incidentally, OLS
applied to a nonlinear regression model is called nonlinear least squares (NLLS). So, what is the
solution?
Estimating Nonlinear Regression Models:
There are several methods of obtaining estimates of NLRMs, such as
The Trial-and-Error Method
Nonlinear least squares (NLLS), and
Linearization through Taylor series expansion.
1. Although linear regression models predominate theory and practice, there are occasions where
nonlinear-in-the-parameter regression models (NLRM) are useful.
2. The mathematics underlying linear regression models is comparatively simple in that one can
obtain explicit, or analytical, solutions of the coefficients of such models. The small-sample and
large-sample theory of inference of such models is well established.
3. In contrast, for intrinsically nonlinear regression models, parameter values cannot be obtained
explicitly. They have to be estimated numerically, that is, by iterative procedures.
4. There are several methods of obtaining estimates of NLRMs, such as (1) trial and error, (2)
nonlinear least squares (NLLS), and (3) linearization through Taylor series expansion.
5. Computer packages now have built-in routines, such as Gauss– Newton, Newton–Raphson, and
Marquard. These are all iterative routines.
6. NLLS estimators do not possess optimal properties in finite samples, but in large samples they do
have such properties. Therefore, the results of NLLS in small samples must be interpreted carefully.
7. Autocorrelation, heteroskedasticity, and model specification problems can plague NLRM, as they
do linear regression models.
8. We illustrated the NLLS with several examples. With the ready availability of user-friendly
software packages, estimation of NLRM should no longer be a mystery. Therefore, the reader should
not shy away from such models whenever theoretical or practical reasons dictate their use.
One objective of analyzing economic data is to predict or forecast the future values of economic
variables. One approach to do this is to build a more or less structural econometric model, describing
the relationship between the variable of interest with other economic quantities, to estimate this
model using a sample of data, and to use it as the basis for forecasting and inference. Although this
approach has the advantage of giving economic content to one‘s predictions, it is not always very
useful. For example, it may be possible to adequately model the contemporaneous relationship
between unemployment and the inflation rate, but as long as we cannot predict future inflation rates
we are also unable to forecast future unemployment.
A time series is a sequence of numerical data in which each item is associated with a particular instant
in time. Time series data is, as its name suggests, ordered by time. One can quote numerous examples:
monthly unemployment, weekly measures of money supply, daily closing prices of stock indices, and
so on.
In general we consider a time series of observations on some variable, e.g. the unemployment rate,
denoted as Y1… YT. These observations will be considered realizations of random variables that can
be described by some stochastic process. It is the properties of this stochastic process that we try to
describe by a relatively simple model. It will be of particular importance how observations
corresponding to different time periods are related, so that we can exploit the dynamic properties of
the series to generate predictions for future periods.
Stationary processes
A stochastic process {yt} is said to be weakly stationary or covariance stationary if it satisfies the
following conditions.
1. E (yt) is constant (constant mean)
2. Var (yt) < ∞ (finite variance)
3. Cov (yt, yt-k) is a function of k, but not of t. where k = 1, 2, 3……
In other words, a stochastic process {yt} is said to be weakly stationary (or covariance stationary) if
its mean, variance and autocovariances are unaffected by changes of time. Note that condition (3) is
the equivalent to saying that the covariance between observations in the series is a function only of
how far apart the observations are in time.
Example: A simple way to model dependence between consecutive observations states that Yt
depends linearly upon its previous value Yt−1.
That is, Yt = δ + θYt−1 + ut |θ| < 1. …………………………………………………6.2.1
where Ut denotes a serially uncorrelated innovation with a mean of zero and a constant variance. The
process in (7.1) is referred to as a first order autoregressive process or AR (1) process. It says that
the current value Yt equals a constant δ plus θ times its previous value plus an unpredictable
component ut. For the moment, we shall assume that |θ| < 1. The process for ut is an important
building block of time series models and is referred to as a white noise process. In this chapter, ut
will always denote such a process that is homoscedastic and exhibits no autocorrelation.
The expected value of Yt can be solved from
E (Yt) = δ + θ E (Yt -1), which, assuming that E (Yt) does not depend upon t, allows us to write
E (Yt) = δ + θ E (Yt), E (Yt ) ……………………………………………6.2.2
1
Defining yt ≡ Yt - μ, we can write (7.1) as
Yt = θyt-1 + εt. ………………………………………………….……………………………6.2.3
Writing time series models in terms of yt rather than Yt is often Notationally more convenient, and
we shall do so frequently in the rest of this chapter. One can allow for nonzero means by adding an
intercept term to the model. While Yt is observable, yt is only observed if the mean of the series is
known. Note that V (yt) = V (Yt).
The model in (5.321) is a parsimonious way of describing a process for the Yt series with certain
properties. That is, the model in (5.3.2) implies restrictions on the time series properties of the
process that generates Yt. In general, the joint distribution of all values of Yt is characterized by the
so-called autocovariances, the covariances between Yt and one of its lags, Yt-k.
For the AR (1) model, the dynamic properties of the Yt series can easily be determined using (9.1) or
(9.3) if we impose that variances and autocovariances do not depend upon the index t. This is a so-
called stationarity assumption and we return to it below. Writing
Var (Yt ) Var (Yt 1 ut ) E Yt 1 ut 2 2 Var (Yt 1 ) Var (ut )
2
Var (Yt ) …………………………………………………….………………..6.2.4
1 2
It is clear from the resulting expression that we can only impose Var (Yt) = Var (Yt-1) if |θ| < 1, as
was assumed before. Furthermore, we can determine that
2
Cov (Yt Yt 1 ) E (Yt Yt 1 ) E Yt 1 Yt 1 Var (Yt 1 ) .......6.2.5
1 2
And, generally (for k = 1, 2, 3, . . .),
2
Cov (Yt Yt k ) K
...................................................................................6.2.6
1 2
As long as θ is nonzero, any two observations on Yt have a nonzero correlation, while this
dependence is smaller (and potentially arbitrary close to zero) if the observations are further apart.
Note that the covariance between Yt and Yt−k depends on k only, not on t. This reflects the
stationarity of the process.
Example: Suppose y follows the random walk process:
Yt = yt-1 + ut ………………………………………………………………………………… 6.2.7
Where the stochastic term ut is white noise. One key trait of random walks is that the most recently
observed value of the variable is the best forecaster of future values. By repeated substitution
(assuming y0 = 0), we can write equation (9.7) as:
t
yt ut
t 1
t
E ( yt ) E (ut ) 0 (cons tan t )
t 1
2
Var ( yt ) E ( yt ) E ut E (ut 2 ) t u 2
2 t t
t 1 t 1
The variance depends on t, meaning that the variance of the series is diverging to infinity with t.
thus, the series is not stationary.
Another simple time series model is the first order moving average process or MA (1) process, given
by
Yt = μ + εt + αεt−1. …………………………………………………………………………..6.2.8
Apart from the mean μ, this says that Y1 is a weighted average of ε1 and ε0, Y2 is a weighted average
of ε2 and ε1, etc. The values of Yt are defined in terms of drawings from the white noise process εt.
The variances and autocovariances in the MA (1) case are given by
Var (Yt ) E t t 1 2 E ( t 2 ) 2 E ( t 2 ) 1 2 2
Cov (Yt , Yt 1) E t t 1 ( t 1 t 2 E ( t 12 ) 2
stationary time series, we examine the correlogram to decide on the appropriate orders of the AR and
MA components. The correlogram of a MA process is zero after a point. That of an AR process
declines geometrically. The correlogram of ARMA processes show different patterns (but all
dampen after a while). Based on these, one arrives at a tentative ARMA model. This step involves
more of a judgmental procedure than the use of any clear-cut rules.
3. The next step is the estimation of the tentative ARMA model identified in step 2.
4. The next step is diagnostic checking to check the adequacy of the tentative model.
5. The final step is forecasting
Forecasting from Box-Jenkins Models
Suppose that we have estimated the model with n observations. We want to forecast X . This is
t k
called a k-period ahead forecast. It is denoted by X . The first subscript gives the time period
n, k
when the forecast is made, and the second subscript denotes the time periods ahead for which the
forecast is made. Let us start with k = 1 so that we need a forecast of X at time period n. We
n 1
have
X n 1 1 X n 2 X n 1 n 1 1 n 2 n 1
We observe Xn and Xn-1. We can replace n and n1 by the predicted residuals.
Predicted residuals: Suppose that we take sample data of n observations and estimate the regression
equation with (n-1) observations at a time by omitting one observation and then use this estimated
equation to predict the y value for the omitted observation. Let us denote the prediction error by
U *
yt y i . The U * are the predicted residuals. By y i we mean a prediction of yt from
a regression equation that is estimated from all observations except the ith observation. This is in
contrast to yt , which is the predicted value of yt from a regression equation that is estimated using
all the observations.
The only unknown is n 1 . Thus we replace by its expected value, zero. Hence
X n ,1
1 X n
2 X n 1 1 n
2 n 1
X n 2 1 X n 1 2 X n n 2 1 n 1 2 n
We replace n 2 and n 1 by zero, their expected value. X n 1 is not known, but we have
the forecast X n ,1
. Thus we get
X n, 2
1 X n,1
2 X n 2 n
We continue like this. The procedure is:
In the previous section we considered models for the stochastic process of a single economic time
series. One reason why it may be more interesting to consider several series simultaneously is that it
may improve forecasts. For example, the history of a second variable, Xt say, may help forecasting
future values of Yt. It is also possible that particular values of Xt are associated with particular
movements in the Yt variable. For example, oil price shocks may be helpful in explaining gasoline
consumption. In addition to the forecasting issue, this also allows us to consider ‗what if‘ questions.
For example, what is the expected future development of gasoline consumption if oil prices are
decreasing by 10% over the next couple of years?
Dynamic Models with Stationary Variables
Considering an economic time series in isolation and applying techniques from the previous
discussion to model it may provide good forecasts in many cases. It does not, however, allow us to
determine what the effects are of, for example, a change in a policy variable. To do so, it is possible
to include additional variables in the model. Let us consider two (stationary) variable Yt and Xt, and
assume that
The interesting element in (7.9) is that it describes the dynamic effects of a change in X t upon
current and future values of Yt.
Taking partial derivatives, we can derive that the immediate response is given by
Yt
0 ...................................................................................................................................... 6.2.10
X t
Sometimes this is referred to as the impact multiplier. An increase in X with one unit has an
immediate impact on Y of 0 units.
In regression analysis involving time series data, if the regression model includes not only current but
also the lagged (past) values of the explanatory variables, then it is called a distributed lag-model.
Thus,
In this sub section, we will discuss the importance of lag in economics. Then we provide the
techniques to estimate a model with lagged values of the explanatory variables. That means it
provides techniques for estimation a distributive lagged model.
s
i
i 0
Lagged values of the variables are important explanatory variables, because most economic
variables are influenced by past patterns of the variable. For example, take the consumption
function. It postulates that the current level of consumption depends on past levels of consumption
and current and past levels of income, i.e.
Ct = f (Ct-1, Yt, Yt-1, X1t, X2t…)
The investment function postulates that it depends on past outputs, on expectation about future
profits, on capital stock and other factors, i.e.
It= f (Qt, Qt-1, Qt-2, …, t, Kt-1,it,…)
Where: Q is the level of output
is profit
K is capital stock
i is interest rate
Very often, dependent variable responds to the explanatory variables with a lapse of time. Such a
lapse of time is called a lag.
Lags are important for decision making especially, for government officials to know how fast, after
how many time periods the economic units will react to changes of various policy variables
(instruments). For example, how fast will producers or consumers react to the imposition of sales tax
and other incentives for investment? How fast will investors react to changes in the interest rate?
Similarly, the lags involved in the demand function following a change in the policy instruments like
price, quantity or advertising of a firm are important for managerial decisions.
Lagged variables are one-way for taking into account the length of time in the adjustment processes
of economic behavior and for handling of expectations about future events. However, economic
89 Bish200821@[Link] | Prepared by: Bishaw A. (MSc) : Econometrics Module
Debre Tabor University, CAES, Department of Agricultural Economics
theory never suggests the precise number of lags that should be included in a function, even if it
recognizes the importance of time lags. The researcher will choose among different lags patterns the
one that gives the most satisfactory fit on the basis of statistical criteria.
Estimation of Distributed lag models
Assume that Y depends on the value of X over s periods. This is called a finite (lag) distributed lag
model because the length of lag is specified.
U N (0, 2)
U
E (UiUj) = 0 for i j
E (UiXj) =0 for j=1, 2, 3…K
Since the explanatory variable Xt is non-stochastic, it is uncorrelated with the disturbance term Ut.
Thus, Xt, X t-1 and so on are non-stochastic and the ordinary least squares can be applied.
This method suggests that to estimate such a model, one may proceed sequentially., i.e. first regress
Yt on Xt, then regress Yt on Xt and X t-1 and so on. This sequential procedure stops when the
regression coefficients of the lagged variables start becoming statistically insignificant and/or the
coefficient of at least one of the variables changes signs from positive to negative or vice versa.
However, the following problems will arise in attempting to apply this approach.
If the number of lags is large and the sample is small (in the case of time series data), we may be
unable to estimate the parameters because there will be no adequate degrees of freedom to carry out
the statistical tests of significance.
There will be a Multicollinearity problem, since there is strong correlation between successive
values of the same variable. With strong collinearity, the values of the estimates will be imprecise
90 Bish200821@[Link] | Prepared by: Bishaw A. (MSc) : Econometrics Module
Debre Tabor University, CAES, Department of Agricultural Economics
and their standard errors will be large so that we may be led to mis-specification of the model by
dropping variables.
To avoid these problems, various methods have been suggested to reduce the number of lagged
variables. This is achieved by imposing restrictions on the 𝛽‘s and constructing new variables from a
linear combination of the lagged variables. The methods differ in the weights which are used in
constructing these new variables.
One of the most popular distributed lag models with endogenous lagged variables is Koyck‘s
Geometric lag scheme. This model assumes that the weights/lag coefficients are declining
continuously following the pattern of geometric progression.
Characteristic equation and unit roots
Define the lag operator Z as: ZkXt= Xt-k k = 1, 2, 3…
e.g. Z1Xt = ZXt = Xt-1 Z2Xt = Xt-2 etc.
Suppose that a stochastic process {yt} follows an autoregressive process of order p AR (p):
yt 1Zyt 2 Z 2 yt ..... p Z p y p ut
1 1Z 2 Z 2 .......... p Z p yt ut
a (Z ) yt ut
3. In particular, if Z = 1, is a root of the characteristic equation (7.15), then the stochastic process is said to
have a unit root.
yt Zyt ut
(1 Z ) yt ut
a (Z ) yt ut
The characteristic equation is given by: 1 Z 0 and the root of the equation is Z 1 .
a) If | | < 1, then Z = 1/p will be larger than one (outside the unit circle), and hence, the process will be
stationary.
b) If = 1 (meaning the process is a random walk), then the root of the characteristic equation is Z = 1
(on the unit circle). Hence, the process has a unit root (is integrated of order one) and is a non-stationary
time series.
Integrated processes and differencing
The difference operator is defined as:
d yt 1 Z d yt d 1,2,3,...
where Z is the lag operator: Zyt yt 1
e.g. First difference: 1 yt yt 1 Z yt yt Zyt yt yt 1
yt - yt-1 = ut y t u t
Since the error term ut is white noise (which is stationary by definition), the first difference yt is also
stationary. The series yt is said to be integrated of order one, denoted by I (1), since taking a first
difference produces a stationary process. A series is said to be integrated of order d, denoted by I (d),
if the series becomes stationary after being differenced d times.
Most macroeconomic time series such as output or unemployment rate are I (1). An I (2) series is
growing at an ever-increasing rate. But series that are I (3) or greater are extremely unusual.
Correcting for error autocorrelation of AR (1) scheme
Consider the model:
t t 1 ut Where ut fulfills all assumptions of the CLRM. Suppose by applying any one of
the above tests you come to the conclusion that the errors are auto correlated. What to do next?
Where ut t t 1.
The above transformation is known as the Cochrane-Orcutt transformation. Since ut fulfills all
assumptions of the CLRM, we can apply OLS to equation (***) to get estimates which are BLUE.
APPENDIX
F Distribution Tables
The F distribution is a right-skewed distribution used most commonly in Analysis of Variance. When referencing the F
distribution, the numerator degrees of freedom are always given first, as switching the order of degrees of freedom
changes the distribution (e.g., F(10,12) does not equal F(12,10) ). For the four F tables below, the rows represent
denominator degrees of freedom and the columns represent numerator degrees of freedom. The right tail area is given in
the name of the table. For example, to determine the .05 critical value for an F distribution with 10 and 12 degrees of
freedom, look in the 10 column (numerator) and 12 row (denominator) of the F Table for alpha=.05. F (.05, 10, 12) =
2.7534