0% found this document useful (0 votes)
4 views26 pages

Classical Regression Analysis Explained

Chapter Two discusses classical regression analysis, focusing on the estimation and prediction of dependent variables based on independent variables. It introduces simple linear regression models, distinguishing between deterministic and stochastic relationships, and outlines key assumptions of the classical linear stochastic regression model. The chapter also covers methods of estimation, particularly the Ordinary Least Squares (OLS) method for estimating parameters in regression models.

Uploaded by

gechabout
Copyright
© All Rights Reserved
We take content rights seriously. If you suspect this is your content, claim it here.
Available Formats
Download as PDF, TXT or read online on Scribd
0% found this document useful (0 votes)
4 views26 pages

Classical Regression Analysis Explained

Chapter Two discusses classical regression analysis, focusing on the estimation and prediction of dependent variables based on independent variables. It introduces simple linear regression models, distinguishing between deterministic and stochastic relationships, and outlines key assumptions of the classical linear stochastic regression model. The chapter also covers methods of estimation, particularly the Ordinary Least Squares (OLS) method for estimating parameters in regression models.

Uploaded by

gechabout
Copyright
© All Rights Reserved
We take content rights seriously. If you suspect this is your content, claim it here.
Available Formats
Download as PDF, TXT or read online on Scribd

CHAPTER TWO

THE CLASSICAL REGRESSION ANALYSIS


Regression is the estimation or prediction of the average value of a dependent variable on the basis
of the fixed values of other variables. In regression analysis Y is dependent variable & X is
independent variable or there are many independent variables. Regression model or equation is
the mathematical equation that helps to predict or forecast the value of the dependent variable
based on the known independent variables. Regression analysis may how cause-&- effect
relationship between X & Y. Correlation analysis shows the existence of a relationship between
two variables X & Y, but not they have cause-&-effect relationship. In correlation analysis both Y
& X are independent variables.

2.1. The simple linear regression model (SLR)

In this topic we consider a simple linear regression (SLR) model, i.e., a relationship between two variables
related in a linear form. We shall first discuss two important forms of relation: stochastic and non-
stochastic, among which we shall be using the former in econometric analysis.
Non-stochastic: A relationship between X and Y, characterized as Y = f(X) is said to be
deterministic or non-stochastic if for each value of the independent variable (X) there is one and
only one corresponding value of dependent variable (Y).
Example; Q = f ( P) =  + P − − − − − − − − − − − − − − − − − − − − − − − − − − − (2.1)

The above relationship between P and Q is such that for a particular value of P, there is only one
corresponding value of Q. This is, therefore, a deterministic (non-stochastic) relationship since
for each price there is always only one corresponding quantity supplied.
In non-stochastic relationship since for each price there is always only one corresponding quantity
supplied. This implies that all the variation in Y is due solely to changes in X, and that there are
no other factors affecting the dependent variable.
If this were true all the points of price-quantity pairs, if plotted on a two-dimensional plane, would
fall on a straight line. However, if we gather observations on the quantity actually supplied in the
market at various prices and we plot them on a diagram we see that they do not fall on a straight
line.

1|Page ZigaleY.
The derivation of the observation from the line may be attributed to several factors.
a. Omission of variables from the function
b. Random behavior of human beings
c. Imperfect specification of the mathematical form of the model
d. Error of aggregation
e. Error of measurement
Stochastic Relationships; a relationship between X and Y is said to be stochastic if for a particular
value of X there is a whole probabilistic distribution of values of Y. In such a case, for any given
value of X, the dependent variable Y assumes some specific value only with some probability.
In order to take into account, the above sources of errors we introduce in econometric functions a
random variable which is usually denoted by the letter ‘u’ or ‘  ’ and is called error term or random
disturbance or stochastic term of the function, so called because u is supposed to ‘disturb’ the exact
linear relationship which is assumed to exist between X and Y. By introducing this random
variable in the function, the model is rendered stochastic of the form:
Yi =  + X + u i …………………………………..…………….………………….(2.2)

Thus, a stochastic model is a model in which the dependent variable is not only determined by the
explanatory variable(s) included in the model but also by others which are not included in the
model.
2.1. Simple Linear Regression model.
The above stochastic relationship (2.2) with one explanatory variable is called simple linear
regression model. The true relationship which connects the variables involved is split into two
parts: a part represented by a line and a part represented by the random term ‘u’.

2|Page ZigaleY.
The scatter of observations represents the true relationship between Y and X. The line represents
the exact part of the relationship and the deviation of the observation from the line represents the
random component of the relationship.
- Were it not for the errors in the model, we would observe all the points on the line Y1' , Y2' ,......,Yn'

corresponding to X 1 , X 2 ,...., X n . However, because of the random disturbance, we observe

Y1 , Y2 ,......,Yn corresponding to X 1 , X 2 ,...., X n . These points diverge from the regression line by

u1 , u 2 ,...., u n .

Yi =  +  xi + ui
 
  
the dependent var iable the regression line random var iable

The first component in the bracket is the part of Y explained by the changes in X and the second
is the part of Y not explained by X, that is to say the change in Y is due to the random influence
of u i .

2.1.1 Assumptions of the Classical Linear Stochastic Regression Model.


The classicals made important assumption in their analysis of regression. The most important of
these assumptions are discussed below.
1. The model is linear in parameters.
The classicals assumed that the model should be linear in the parameters regardless of whether the
explanatory and the dependent variables are linear or not. This is because if the parameters are
non-linear, it is difficult to estimate them since their value is not known but you are given with the
data of the dependent and independent variable.
Example 1. Y =  + x + u is linear in both parameters and the variables, so it
Satisfies the assumption
2. ln Y =  +  ln x + u is linear only in the parameters. Since the

3|Page ZigaleY.
the classicals worry on the parameters, the model satisfies
the assumption.
Dear distance students! Check yourself whether the following models satisfy the above assumption
and give your answer to your tutor.
a. ln Y 2 =  +  ln X 2 + U i

b. Yi =  + X i + U i

2). U i is a random real variable

This means that the value which “u” may assume in any one period depends on chance; it may
be positive, negative or zero. Every value has a certain probability of being assumed by ‘u’ in
any particular instance.
3). The mean value of the random variable(U) in any particular period is zero
This means that for each value of x, the random variable(u) may assume various values, some
greater than zero and some smaller than zero, but if we considered all the possible and negative
values of u, for any given value of X, they would have on average value equal to zero. In other
words, the positive and negative values of “u” cancel each other.
Mathematically, E (U i ) = 0 ………………………………………………...…. (2.3)

4). The variance of the random variable(U) is constant in each period (The assumption of
homoscedasticity)
For all values of X, the u’s will show the same dispersion around their mean. In Fig.2.c this
assumption is denoted by the fact that the values that u can assume lie with in the same limits,
irrespective of the value of X. For X 1 , u can assume any value with in the range AB; for X 2 , u
can assume any value with in the range CD which is equal to AB and so on.
Graphically;

4|Page ZigaleY.
Mathematically;
Var (U i ) = E[U i − E (U i )] 2 = E (U i ) 2 =  2 (Since E (U i ) = 0 ). This constant variance is called
homoscedasticity assumption and the constant variance itself is called homoscedastic variance.
5). The random variable (U) has a normal distribution
This means the values of u (for each x) have a bell-shaped symmetrical distribution about their
zero mean and constant variance  2 , i.e.
U i  N (0,  2 ) …………………………………………..…………..……2.4

6). The random terms of different observations (U i ,U j ) are independent. (The assumption of no

autocorrelation)
This means the value which the random term assumed in one period does not depend on the value
which it assumed in any other period.
Algebraically,
 
Cov(u i u j ) =  [(u i − (u i )][u j − (u j )]

= E (u i u j ) = 0 ………………………….…………..….(2.5)

7). The X i are a set of fixed values in the hypothetical process of repeated sampling which

underlies the linear regression model.


This means that, in taking large number of samples on Y and X, the X i values are the same in all

samples, but the u i values do differ from sample to sample, and so of course do the values of y i .

8). The random variable (U) is independent of the explanatory variables.


This means there is no correlation between the random variable and the explanatory variable. If
two variables are unrelated their covariance is zero.
Hence Cov( X i ,U i ) = 0 ………………………………………..…………..….(2.6)

Proof: - cov( XU ) = [( X i − ( X i )][U i − (U i )]

= [( X i − ( X i )(U i )] given E (U i ) = 0

= ( X iU i ) − ( X i )(U i )

=  ( X iU i )

= X i (U i ) , given that the x i are fixed

=0

5|Page ZigaleY.
8). The explanatory variables are measured without error
- U absorbs the influence of omitted variables and possibly errors of measurement in the y’s.
i.e., we will assume that the regressors are error free, while y values may or may not include
errors of measurement.
Dear students! We can now use the above assumptions to derive the following basic concepts.
A. The dependent variable Yi is normally distributed.

 
i.e  Yi ~ N ( + x i ),  2 …………………………………….…………(2.7)

Proof: Mean: (Y ) = ( + xi + u i )

=  + X i Since (u i ) = 0

Variance: Var(Yi ) = (Yi − (Yi ) )


2

= ( + X i + ui − ( + X i ))
2

= (u i ) 2 =  2 (since (u i ) 2 =  2 )

 var(Yi ) =  2 …………………………………………………….(2.8)

The shape of the distribution of Yi is determined by the shape of the distribution of u i which is

normal by assumption 4. Since  and  , being constant, they don’t affect the distribution of y i .

Furthermore, the values of the explanatory variable, x i , are a set of fixed values by assumption 5

and therefore don’t affect the shape of the distribution of y i .

 Yi ~ N( + x i ,  2 )

B. successive values of the dependent variable are independent, i.e


Cov(Yi , Y j ) = 0

Proof: Cov(Yi , Y j ) = E{[Yi − E (Yi )][Y j − E (Y j )]}

= E{[ + X i + U i − E ( + X i + U i )][ + X j + U j − E ( + X j + U j )}

(Since Yi =  +  X i + U i and Y j =  + X j + U j )

= E[( + X i + Ui −  − X i )( + X j + U j −  − X j )] ,Since (u i ) = 0

= E (U iU j ) = 0 (From equation (2.5))

Therefore, Cov(Yi ,Y j ) = 0 .

6|Page ZigaleY.
2.1.2 Methods of estimation
Specifying the model and stating its underlying assumptions are the first stage of any econometric
application. The next step is the estimation of the numerical values of the parameters of economic
relationships. The parameters of the simple linear regression model can be estimated by various
methods. Three of the most commonly used methods are:
1. Ordinary least square method (OLS)
2. Maximum likelihood method (MLM)
3. Method of moments (MM)
But here we will deal with the OLS methods of estimation.
[Link] The ordinary least square (OLS) method
The model Yi =  +  X i + U i is called the true relationship between Y and X because Y and X

represent their respective population value, and  and  are called the true parameters since they
are estimated from the population value of Y and X But it is difficult to obtain the population value
of Y and X because of technical or economic reasons. So, we are forced to take the sample value
of Y and X. The parameters estimated from the sample value of Y and X are called the estimators
of the true parameters  and  and are symbolized as ˆ and ˆ .

The model Yi = ˆ + ˆX i + ei , is called estimated relationship between Y and X since ˆ and ˆ

are estimated from the sample of Y and X and ei represents the sample counterpart of the

population random disturbance U i

Estimation of  and  by least square method (OLS) or classical least square (CLS) involves

finding values for the estimates ˆ and ˆ which will minimize the sum of square of the squared

residuals (  ei2 ).

From the estimated relationship Yi = ˆ + ˆX i + ei , we obtain:

ei = Yi − (ˆ + ˆX i ) ……………………………(2.6)

e 2
i =  (Yi − ˆ − ˆX i ) 2 ……………………….(2.7)

To find the values of ˆ and ˆ that minimize this sum, we have to partially differentiate e 2
i

with respect to ˆ and ˆ and set the partial derivatives equal to zero.

7|Page ZigaleY.
  ei2
1. = −2 (Yi − ˆ − ˆX i ) = 0.......................................................(2.8)

ˆ

Rearranging this expression, we will get: Y i = n + ˆX i …………………...… (2.9)

If you divide (2.9) by ‘n’ and rearrange, we get


ˆ = Y − ˆX ..........................................................................(2.10)

  ei2
2. = −2 X i (Yi − ˆ − ˆX ) = 0..................................................(2.11)
ˆ

Note: at this point that the term in the parenthesis in equation 2.8and 2.11 is the residual,
e = Yi − ˆ − ˆX i . Hence it is possible to rewrite (2.8) and (2.11) as − 2 ei = 0 and

− 2 X i ei = 0 . It follows that;

e i = 0 and X e i i = 0............................................(2.12)

If we rearrange equation (2.11) we obtain;

Y Xi i = ˆX i + ˆX i2 ……………………………………….(2.13)

Equation (2.9) and (2.13) are called the Normal Equations. Substituting the values of ̂ from
(2.10) to (2.13), we get:

Y X i i = X i (Y − ˆX ) + ˆX i2

= YX i − ˆXX i + ˆX i2

Y X i i − Y X i = ˆ (X i2 − XX i )

XY − nXY = ˆ ( X i2 − nX 2)

XY − nXY
ˆ = ………………….(2.14)
X i2 − nX 2

Equation (2.14) can be rewritten in somewhat differen;2t way as follows;


( X − X )(Y − Y ) = ( XY − XY − XY + XY )

= XY − Y X − XY + nXY


= XY − nY X − nXY + nXY

8|Page ZigaleY.

( X − X )(Y − Y ) = XY − n X Y − − − − − − − − − − − − − −(2.15)

( X − X ) 2 = X 2 − nX 2 − − − − − − − − − − − − − − − − − (2.16)
Substituting (2.15) and (2.16) in (2.14), we get
( X − X )(Y − Y )
ˆ =
( X − X ) 2

Now, denoting ( X i − X ) as x i , and (Yi − Y ) as y i we get;

xi yi
ˆ = ……………………………………… (2.17)
xi2

The expression in (2.17) to estimate the parameter coefficient is termed is the formula in deviation
form.
2.1.3 Estimation of a function with zero intercept
Suppose it is desired to fit the line Yi =  +  X i + U i , subject to the restriction  = 0. To estimate

ˆ , the problem is put in a form of restricted minimization problem and then Lagrange method is
applied.
n
We minimize: ei2 =  (Yi − ˆ − ˆX i ) 2
i =1

Subject to: ˆ = 0
The composite function then becomes
Z =  (Yi − ˆ − ˆX i ) 2 − ˆ , where  is a Lagrange multiplier.

We minimize the function with respect to ˆ , ˆ , and 


Z
= −2(Yi − ˆ − ˆX i ) −  = 0 − − − − − − − −(i)
ˆ
Z
= −2(Yi − ˆ − ˆX i ) ( X i ) = 0 − − − − − − − −(ii)
ˆ

z
= −2 = 0 − − − − − − − − − − − − − − − − − − − (iii)

Substituting (iii) in (ii) and rearranging we obtain:
X i (Yi − ˆX i ) = 0

Yi X i − ˆX i = 0
2

9|Page ZigaleY.
X i Yi
ˆ = ……………………………………..(2.18)
X i2

This formula involves the actual values (observations) of the variables and not their deviation
forms, as in the case of unrestricted value of ˆ .
2.1.4. Statistical Properties of Least Square Estimators
There are various econometric methods with which we may obtain the estimates of the parameters
of economic relationships. We would like to an estimate to be as close as the value of the true
population parameters i.e., to vary within only a small range around the true parameter. How are
we to choose among the different econometric methods, the one that gives ‘good’ estimates? We
need some criteria for judging the ‘goodness’ of an estimate.
‘Closeness’ of the estimate to the popular;8tion parameter is measured by the mean and variance
or standard deviation of the sampling distribution of the estimates of the different econometric
methods. We assume the usual process of repeated sampling i.e., we assume that we get a very
large number of samples each of size ‘n’; we compute the estimates ˆ ’s from each sample, and
for each econometric method and we form their distribution. We next compare the mean (expected
value) and the variances of these distributions and we choose among the alternative estimates the
one whose distribution is concentrated as close as possible around the population parameter.
PROPERTIES OF OLS ESTIMATORS
The ideal or optimum properties that the OLS estimates possess may be summarized by well-
known theorem known as the Gauss-Markov Theorem.
Statement of the theorem: “Given the assumptions of the classical linear regression model, the
OLS estimators, in the class of linear and unbiased estimators, have the minimum variance, i.e.
the OLS estimators are BLUE.
According to this theorem, under the basic assumptions of the classical linear regression model,
the least squares estimators are linear, unbiased and have minimum variance (i.e. are best of all
linear unbiased estimators). Sometimes the theorem referred as the BLUE theorem i.e. Best,
Linear, Unbiased Estimator. An estimator is called BLUE if:
a. Linear: a linear function of a random variable, such as, the dependent variable Y.
b. Unbiased: its average or expected value is equal to the true population parameter.

10 | P a g e ZigaleY.
c. Minimum variance: It has a minimum variance in the class of linear and unbiased
estimators. An unbiased estimator with the least variance is known as an efficient
estimator.
According to the Gauss-Markov theorem, the OLS estimators possess all the BLUE properties.
The detailed proof of these properties is read by yourself:
Example of applying OLS
The marketing manager of Safaricom, the new competitor of Ethio Telecom in the Ethiopian
telecommunication market wants to know the relation between expenses on advertisement and
new clients. To this end, the team collects the following data:

Month January February April May June July August Sept Oct
X: Advertisement
expenses (in 100,000 2.5 3 2.25 4 3.5 10 3.5 2.5 1.2
Birr)
Y: New clients (in
4 5 3 5.5 4 10 5 3 1 Sum
100,000)

-
𝑋𝑖 − 𝑋̄
-1.11 -0.61 -1.36 0.39 0.11 6.39 -0.11 -1.11 -2.41
-
𝑌𝑖 − 𝑌̄
-0.50 0.50 -1.50 1.00 0.50 5.50 0.50 -1.50 -3.50

(𝑋𝑖 − 𝑋̄) ∗ ( 𝑌𝑖 − 𝑌̄) 0.55 -0.30 2.03 0.39 0.05 35.17 -0.05 1.66 8.42 47.93

(𝑋𝑖 − 𝑋̄)2 1.22 0.37 1.84 0.16 0.01 40.89 0.01 1.22 5.79 51.50

Note that the relationship displayed in Figure 3 indicates a positive linear relation between
advertisement expenses and new clients.

11 | P a g e ZigaleY.
Now, let’s use the Least Sum of Squares
approach to estimate 𝛼 and 𝛽, or, in other
words, find 𝛼̂ and 𝛽̂ . Firstly, we need to
find the mean of X (𝑋̄) and the mean of Y
(𝑌̄).

2.5+3+2.25+4+3.5+10+3.5+2.5+1.2
𝑋̄ = =
9

3.606

4 + 5 + 3 + 5.5 + 4 + 10 + 5 + 3 + 1 Figure 1: Scatterplot displaying the relation between advertisement


𝑌̄ = = 4.5 expenditure and new clients.
9

Next, we need to find (𝑋𝑖 − 𝑋̄)(𝑌𝑖 − 𝑌̄) and (𝑋𝑖 − 𝑋̄)2 for each i. These are displayed in the table
above. We find that the sum of (𝑋𝑖 − 𝑋̄)2 is 51.5, and the sum of (𝑋𝑖 − 𝑋̄)(𝑌𝑖 − 𝑌̄) is 47.93. This,
we use to calculate 𝛽̂ .
𝛴(𝑋𝑖 − 𝑋̄)(𝑌𝑖 − 𝑌̄) 47.93
𝛽̂ = = = 0.9305
𝛴(𝑋𝑖 − 𝑋̄)2 51.5

That means that, if advertisement expenses increase by 100,000 Birr, Safaricom gets, on average,
93,050 extra clients. Next, we can calculate 𝛼̂
𝛼̂ = 𝑌̅ − 𝛽̂ 𝑋̅ = 4.5 − 0.9305 ∗ 3.606 = 1.1449

Summarizing, the relation between advertisement expenses (X) and extra clients (Y) for Safaricom
can be expressed by the following equation:
𝑌𝑖 = 1.1449 + 0.9305𝑋𝑖 + 𝑢𝑖

This information is also visible in Figure 4, which shows Stata output of the Simple Linear
Regression. Note that 𝛼̂ is the _cons value in the output, and 𝛽̂ is the coefficient for the variable
expenditure. The other elements of this Stata output will be discussed in the sections that follow.

12 | P a g e ZigaleY.
Figure 2: Stata output for SLR with Y= new clients (in 100,000) and X=expenditure on
advertisement (in 100,000 Birr)
2.2 Covariance, correlation coefficient, coefficient of determination (R2)

Simple linear regression is used to analyze two continuous variables that are linearly related. To see how
strong this relation is, and if it is a positive or a negative relation, it is common to calculate measures of
association: covariance, correlation and the coefficient of determination.

The covariance between X and Y measures if X and Y are positively, negatively or not related. If COV
(X, Y)>0, X and Y are positively related, if COV (X, Y) <0, X and Y are negatively related and if COV
(X, Y) is 0, or close to 0, there is no relation between X and Y. The mathematical expression for
covariance is:

1
𝐶𝑂𝑉(𝑋, 𝑌) = ∑(𝑋𝑖 − 𝑋̄)(𝑌𝑖 − 𝑌̄)
𝑛−1

A disadvantage of covariance is that it is sensitive to the scale of the data and this makes it difficult to
interpret. It is not clear if a COV (X, Y) of 4 is high or low, because it depends on how large X and Y are.

To compensate for this limitation, a more common way to evaluate the relation between two continuous
variables is Pearson’s correlation. Pearson’s correlation is a figure between -1 and 1. If CORR(X,Y)=0,
there is no relation between X and Y, if CORR(X,Y)=1 there is a perfect positive relation between X and
Y and if CORR(X,Y)=-1 the is a perfect negative relation between X and Y. When working with statistical
data, perfect relations are uncommon, and therefore the correlation coefficient usually takes a value between
-1 and 1. Consider, for example, the scatterplots in Figure 5. As a rule of thumb, if CORR(X,Y) is between
0 and 0.3 (or -0.3 and 0), there is no relation, if CORR(X,Y) is between 0.3 and 0.5 (or between -0.5 and -
0.3) the correlation is weak, if CORR(X,Y) is between 0.5 and 0.8 (or between -0.8 and -0.5) the correlation

13 | P a g e ZigaleY.
is moderate and if CORR(X,Y)>0.8 (or <-0.8) the correlation is strong. The mathematical expression for
correlation is:

𝐶𝑂𝑉(𝑋, 𝑌)
𝐶𝑂𝑅𝑅(𝑋, 𝑌) =
√𝑉𝐴𝑅(𝑋) ∗ 𝑉𝐴𝑅(𝑌)

1 1
With COV(X,Y) as defined above, VAR(X)= 𝑛−1 ∑(𝑋𝑖 − 𝑋̄)2 and VAR(Y)= 𝑛−1 ∑(𝑌𝑖 − 𝑌̄)2

Figure 3: Various scatterplots displaying correlations that differ in sign (+ or -) and magnitude/strength.

The correlation is easier to interpret than the covariance, but when discussing Simple Linear Regression, it
is not the best way to present the relation between X and Y. To express the explanatory power that X has
to explain the variation in Y, it is best to calculate the coefficient of determination, the R2 (or R-squared).

R-squared is a measure of how much of the variance in the actual value of dependent variable (Y) can be
explained by its relationship to the independent variable (X). In other words, R-squared is the percentage
of variance in Y explained by the linear regression equation between X and Y. The R-squared can take any
value in the range [-∞, 1], but usually it takes values between 0 (0% of the variation in Y is explained by
the model) and 1 (the model explains 100% of the variation in Y). The closer the value to 1, the better the
explanatory power of the independent variable is. The mathematical expression for the R-squared is:

𝑆𝑆𝑟𝑒𝑠𝑖𝑑𝑢𝑎𝑙
𝑅2 = 1 −
𝑆𝑆𝑡𝑜𝑡𝑎𝑙

Where SS residual is the sums of squares of the residual term from the linear regression
2
(∑(𝑌𝑖 − 𝑌̂𝑖 ) = ∑ 𝑢𝑖 2 ) and SS total is the total sum of squares of the dependent variable Y (∑(𝑌𝑖 − 𝑌̅)2 ).

14 | P a g e ZigaleY.
From the equation we can see that if the error term is close to 0, R-squared will be close to 1. If, on the
other hand, the residual is relatively large, the R-squared is small.

Example covariance, correlation and coefficient of determination


The government included the course entrepreneurship in the freshman year to improve the entrepreneurial
attitude of the students, so they may have a good business after graduation. For 10 graduates with an own
business, data was collected regarding their entrepreneurship grade (X) and the profitability of their own
business (Y). To see if the score and the business profit are related, we will calculate the covariance, the
correlation and the coefficient of determination.
∑Y ∑X
By using the equation to get the mean value, 𝑌̄ = 𝑛 𝑖 and 𝑋̄ = 𝑛 𝑖, we get 𝑋̄ =76.4 and 𝑌̄ = 52.9. To save

space, it is given that 𝛼̂ = −19.55 and 𝛽̂ = 0.95. These values have been found by using the procedure as
explained in section 2.2.
Entrepreneurship 88 76 65 63 88 53 74 70 89 98
grade (X)
Business profit after a 48 50 41 58 50 37 39 37 69 100
year (Y) in 1000 Birr
(𝑋𝑖 − 𝑋̄) 11.6 -0.4 -11.4 -13.4 11.6 -23.4 -2.4 -6.4 12.6 21.6
(𝑌𝑖 − 𝑌̄) -4.9 -2.9 -11.9 5.1 -2.9 -15.9 -13.9 -15.9 16.1 47.1
(𝑋𝑖 − 𝑋̄)2 134.6 0.2 130.0 179.6 134.6 547.6 5.8 41.0 158.8 466.6
(𝑌𝑖 − 𝑌̄)2 24.0 8.4 141.6 26.0 8.4 252.8 193.2 252.8 259.2 2218.4
̂𝑖 = −19.55 + 0.95𝑋𝑖
𝑌 64.1 52.7 42.2 40.3 64.1 30.8 50.8 47.0 65.0 73.6
𝑢𝑖 = 𝑌𝑖 − 𝑌̂𝑖 -16.1 -2.7 -1.2 17.7 -14.1 6.2 -11.8 -10.0 4.0 26.5
𝑢𝑖 2 257.6 7.0 1.4 313.3 197.4 38.4 138.1 99.0 16.0 699.6
(𝑋𝑖 − 𝑋̄)(𝑌𝑖 − 𝑌̄) -56.8 1.2 135.7 -68.3 -33.6 372.1 33.4 101.8 202.9 1017.4

Based on the information in the table, we can calculate COV (X, Y), CORR (X, Y) and the coefficient of
determination:
1 1
• 𝐶𝑂𝑉(𝑋, 𝑌) = ∑(𝑋𝑖 − 𝑋̄)(𝑌𝑖 − 𝑌̄) = ∗ 1705.4 = 189.5, which means that X and Y are
𝑛 10−1

positively related, because COV (X, Y) >0.


1 1 1 1
• VAR(X)= 𝑛 ∑(𝑋𝑖 − 𝑋̄)2 = 10−1 ∗ 1798.4 = 199.8 and VAR(Y)= 𝑛 ∑(𝑌𝑖 − 𝑌̄)2 = 10−1 ∗ 3384.9 =
𝐶𝑂𝑉(𝑋,𝑌) 189.5
376.1, so 𝐶𝑂𝑅𝑅(𝑋, 𝑌) = = = 0.69, which indicates a positive moderate
√𝑉𝐴𝑅(𝑋)∗𝑉𝐴𝑅(𝑌) √199.8∗376.1

relation between X and Y. Positive because the correlation is >0, and moderate because the
correlation is between 0.5 and 0.8.

15 | P a g e ZigaleY.
𝑆𝑆𝑅 ∑𝑢 2 1767.9
• 𝑅 2 = 1 − 𝑆𝑆𝑇 = 1 − ∑(𝑌 −𝑌
𝑖
̅ )2
= 1 − 3384.9 = 0.478, which indicates that 47.8% of the variation in
𝑖

business profit (Y) can be explained by the score of the entrepreneurship course (X).
Note that the information for the coefficient of determination is also displayed in the Stata output, see Figure
6.

Figure 4: upper part of Stata regression output, displaying the R-squared of the model.

2.3 Hypotheses testing


Next to examining the coefficient of determination, it is common to use hypothesis tests to evaluate
the performance of a linear regression model.

The OLS estimates ˆ and ˆ are obtained from a sample of observations on Y and X. Since
sampling errors are inevitable in all estimates, it is necessary to apply test of significance
in order to measure the size of the error and determine the degree of confidence in order to
measure the validity of these estimates. This can be done by using various tests. The most
common ones are:
i) Standard error test ii) Student’s t-test iii) Confidence interval and F test
All of these testing procedures reach on the same conclusion. Let us now see these testing
methods one by one.
i) Standard error test
This test helps us decide whether the estimates ˆ and ˆ are significantly different from
zero, i.e., whether the sample from which they have been estimated might have come from
a population whose true parameters are zero.  = 0 and / or  = 0 .
Formally we test the null hypothesis
H 0 :  i = 0 against the alternative hypothesis H 1 :  i  0

The standard error test may be outlined as follows.


First: Compute standard error of the parameters.

16 | P a g e ZigaleY.
SE ( ˆ ) = var(ˆ )

SE(ˆ ) = var(ˆ )

Second: compare the standard errors with the numerical values of ˆ and ˆ .
Decision rule:
• If SE(ˆi )  1 2 ˆi , accept the null hypothesis and reject the alternative hypothesis.

We conclude that ˆi is statistically insignificant.

• If SE(ˆi )  1 2 ˆi , reject the null hypothesis and accept the alternative hypothesis.

We conclude that ˆi is statistically significant.


The acceptance or rejection of the null hypothesis has definite economic meaning. Namely,
the acceptance of the null hypothesis  = 0 (the slope parameter is zero) implies that the
explanatory variable to which this estimate relates does not in fact influence the dependent
variable Y and should not be included in the function, since the conducted test provided
evidence that changes in X leave Y unaffected. In other words, acceptance of H0 implies
that the relationship between Y and X is in fact Y =  + (0) x =  , i.e. there is no relationship
between X and Y.
Numerical example: Suppose that from a sample of size n=30, we estimate the following
supply function.
Q = 120 + 0.6 p + ei
SE : (1.7) (0.025)

Test the significance of the slope parameter at 5% level of significance using the standard
error test.
SE ( ˆ ) = 0.025

( ˆ ) = 0.6

1
2 ˆ = 0.3

This implies that SE(ˆi )  1 2 ˆi . The implication ˆ is statistically significant at 5% level
of significance.

17 | P a g e ZigaleY.
Note: The standard error test is an approximated test (which is approximated from the z-
test and t-test) and implies a two-tail test conducted at 5% level of significance.
ii). Student’s t-test
If, in linear regression, X has no effect on Y,  is (close to) 0. We can use the t-test (as you learned in the
last course: Statistics II), to test if 𝛽̂ from the sample data significantly deviates from 0.

The steps for conducting a student’s t-test for a regression coefficient are as follows:

Step 1: Compute t
To compute t, we need to know the value of 𝛽̂ and its “uncertainty”, which is measured by its standard
deviation (𝑆𝐸(𝛽̂ )). With 𝛽̂ and 𝑆𝐸(𝛽̂ ), it is possible to compute the t- test statistic, using the following
formula:

̂
𝛽
𝑡𝛽̂ = ̂) with (n-k) degrees of freedom
𝑆𝐸(𝛽

where 𝛽̂ is the coefficient for X based on the least squares estimation (see section 2.2) and 𝑆𝐸(𝛽̂ ) is the

1 ∑(𝑌 −𝑌 ) ̂ 2
standard error of this estimate, which is 𝑆𝐸(𝛽̂ )=√𝑛−2 ∗ ∑(𝑋𝑖 −𝑋̅𝑖)2 . Instead of calculating this by hand, you
𝑖

can read it from the Stata regression output. n is the sample size, and k is the number of parameters in the
model (k= 2 in SLR).

Step 2: Decide which significance level you’ll use


Level of significance is the probability of making ‘wrong’ decision, i.e. the probability of rejecting the
hypothesis when it is actually true or the probability of committing a type I error. It is customary in
econometric research to choose the 5% or the 1% level of significance. This means that in making our
decision we allow (tolerate) five times out of a hundred to be ‘wrong’ i.e. reject the hypothesis when it is
actually true.

Step 3: Decide whether you want to conduct a one-tailed test, or a two-tailed test
Either you want to see if X affects Y (positively or negatively), for which you’ll do a two-tailed test with
H0: =0 and H1: ≠0. This is the most common way to test a coefficient in linear regression.

However, in some cases economic/management theory expects you to see if X affects Y positively only
(not negatively). For example, to test the law of supply with Y is demand and X is price, it makes sense to
test if  is significantly higher than 0 (rather than higher or lower than 0). In that case you’ll do a one-tailed
test, with H0:=0 and H1: >0.

18 | P a g e ZigaleY.
Of course, it is also possible to see if X affects Y negatively. In that case H0:=0 and H1: <0 and the test
is also one-tailed. This should, for example, be used to test the law of demand from economics, which
states that an increased price has a negative effect on the quantity demanded.

Step 4: Obtain critical value of t, tc, at (n-2) degrees of freedom

To see if the t-value (calculated in step 1) exceeds the


critical t-value, it is necessary to know the critical t-
value. The critical t-value, tc, depends on the degrees
of freedom. The degrees of freedom is n-k. n is the
sample size, and k is the number of parameters. In
SLR,  and  are parameters, so there are two
parameters. Therefore, the degrees of freedom is n-2.
For a two-tailed test with confidence level 5%, and a
sample size of 30 (so the degrees of freedom is 30-
2=28), tc=2.048.

Step 5: Compare t (the computed value of t in step


1) with tc (critical value of t)

If the outcome of step one, t > 2.048 (or t< -2.048), H0 is rejected and H1 is accepted. That means that ≠0,
so X has a significant effect on Y. In other words, if H0 is rejected,  is statistically significant. If the
outcome of step 1 is less than 2.048, H0 is not rejected and  is statistically insignificant.

Using Stata output to do coefficient t-tests


Instead of using the table, and computing t manually, econometrists rely on software, like Stata, to conduct
t-tests. Basic regression output in Stata reports for each coefficient the standard error (Std. Err.), the t-value
and P>|t|. The P>|t| calculates the probability to get the data from the sample, if H0 is true. The common
name for P>|t| is the p-value, just ‘p’, or the significance level. If you use a two-tailed test, and your
confidence level is 5%, you must:

- reject H0 if p ≤ 0.05. In that case, the  is significant, and X significantly affects Y.


- not reject H0 if p > 0.05. In that case, the  is insignificant, and X does not significantly affect Y.
Note that the p-value presented in Stata output is for a two-tailed test. If you want to do a one-tailed test,
you must divide the p-value by 2.

19 | P a g e ZigaleY.
Figure 5: SLR output with Y=new clients (in 100,000) and X=advertisement expenditure (in 100,000 Birr)

For the output in Figure 7, which shows the least square estimation results for the effect of advertisement
expenditure on new clients for Safaricom, 𝛽̂ is 0.93 (t: 8.48, p:0.000), so the effect of expenditure on new
clients is significant (because p<0.05). 𝛼̂ is 1.14 (t:2.39, p:0.048), so the intercept of the model also
significantly deviates from 0 (because p<0.05). In the presentation of regression coefficients in reports or
other publications, it is common to report the coefficient, it’s t-value and the significance level (p) of the
coefficient.

ii) Confidence interval


Rejection of the null hypothesis doesn’t mean that our estimate ˆ and ˆ is the correct
estimate of the true population parameter  and  . It simply means that our estimate
comes from a sample drawn from a population whose parameter  is different from zero.
In order to define how close, the estimate to the true parameter, we must construct
confidence interval for the true parameter, in other words we must establish limiting values
around the estimate with in which the true parameter is expected to lie within a certain
“degree of confidence”. In this respect we say that with a given probability the population
parameter will be within the defined confidence interval (confidence limits).
We choose a probability in advance and refer to it as confidence level (interval coefficient).
It is customarily in econometrics to choose the 95% confidence level. This means that in
repeated sampling the confidence limits, computed from the sample, would include the true
population parameter in 95% of the cases. In the other 5% of the cases the population
parameter will fall outside the confidence interval.

20 | P a g e ZigaleY.
In a two-tail test at  level of significance, the probability of obtaining the specific t-value
either –tc or tc is 
2 at n-2 degree of freedom. The probability of obtaining any value of t

ˆ − 
which is equal to at n-2 degree of freedom is 1 − ( 2 +  2 ) i.e. 1 −  .
ˆ
SE (  )

i.e. Pr− t c  t*  t c  = 1 −  …………………………………………(2.57)

ˆ − 
but t* = …………………………………………………….(2.58)
SE ( ˆ )

Substitute (2.58) in (2.57) we obtain the following expression.


 ˆ −  
Pr − t c   t c  = 1 −  ………………………………………..(2.59)
 SE ( ˆ ) 

 
Pr − SE(ˆ )t c  ˆ −   SE(ˆ )t c = 1 −  − − − − − by multiplying SE(ˆ )

Pr− ˆ − SE(ˆ )t  −  −ˆ + SE( ˆ )t = 1 −  − − − − − by subtracting ˆ


c c

Pr+ ˆ + SE(ˆ )    ˆ − SE( ˆ )t = 1 −  − − − − − by multiplying by − 1


c

Prˆ − SE( ˆ )t    ˆ + SE(ˆ )t  = 1 −  − − − − − int erchanging


c c

The limit within which the true  lies at (1 −  )% degree of confidence is:

[ ˆ − SE( ˆ )t c , ˆ + SE( ˆ )t c ] ; where t c is the critical value of t at 


2 confidence interval and
n-2 degree of freedom.
The test procedure is outlined as follows.
H0 :  = 0

H1 :   0
Decision rule: If the hypothesized value of  in the null hypothesis is within the

confidence interval, accept H0 and reject H1. The implication is that ˆ is statistically
insignificant; while if the hypothesized value of  in the null hypothesis is outside the

limit, reject H0 and accept H1. This indicates ˆ is statistically significant.


Numerical Example:
Suppose we have estimated the following regression line from a sample of 20 observations.

21 | P a g e ZigaleY.
Y = 128.5 + 2.88 X + e
(38.2) (0.85)

The values in the bracket are standard errors.


a. Construct 95% confidence interval for the slope of parameter
b. Test the significance of the slope parameter using constructed confidence interval.
Solution:
a. The limit within which the true  lies at 95% confidence interval is:

ˆ  SE( ˆ )t c

ˆ = 2.88

SE ( ˆ ) = 0.85

t c at 0.025 level of significance and 18 degree of freedom is 2.10.

 ˆ  SE( ˆ )t c = 2.88  2.10(0.85) = 2.88  1.79.

The confidence interval is:


(1.09, 4.67)
b. The value of  in the null hypothesis is zero which implies it is outside the
confidence interval. Hence  is statistically significant.
iii) F-test
Next to conducting hypothesis tests for the coefficients (model parameters) separately, it is also common
to test the explanatory power of the whole model at once. The F-test compares the full model with the
restricted model. In SLR, the full model is the one with an intercept and one explanatory variable. The full
model is displayed in the left panel of Figure 8. In the full model, the expected value of Y (E(Y)) depends
on the X. The restricted model is presented in the right panel of Figure 8. It predicts Y based on the mean
value of Y only. If there is no relation between X and Y, the restricted model is as good as the full model,
because X cannot explain the variation in Y. If, however, X is able to linearly predict Y, we expect the full
model errors to be smaller than the errors in the restricted model.
- The H0 of the F-test is that the restricted model is as good in explaining the variation in Y as the full
model. In SLM that means =0.
- The H1 of the F-test is that the full model is better in explaining the variation in Y than the restricted
model. In SLM that means ≠0.

22 | P a g e ZigaleY.
Figure 6: Comparison of the full model and the restricted model

The steps for conducting an F test for a regression coefficient are as follows:

Step 1: Compute F
𝑆𝑆𝑚𝑜𝑑𝑒𝑙/𝑘 − 1 𝑀𝑆 𝑚𝑜𝑑𝑒𝑙
𝐹= =
𝑆𝑆𝑟𝑒𝑠𝑖𝑑𝑢𝑎𝑙/𝑛 − 𝑘 𝑀𝑆 𝑟𝑒𝑠𝑖𝑑𝑢𝑎𝑙

Where: SS model is ∑(𝑌𝑖 − 𝑌̅𝑖 )2 − ∑(𝑌𝑖 − 𝑌̂𝑖 )2, SS residual is ∑(𝑌𝑖 − 𝑌̂𝑖 )2 . When divided by the
degrees of freedom (k-1 for the model and n-k for the residuals), you get the mean sum of squared
(MS). Dividing the MS of the model by the MS residual, gives us the F-value.

Consider this example, based on data about new clients and advertisement expenditure, with 𝑌̅ =
4.5, 𝛼̂ = 1.145 and 𝛽̂ = 0.931, so 𝑌̂𝑖 = 1.145 + 0.931 ∗ 𝑋𝑖 ).

Month January February April May June July August Sept Oct
X: Advertisement expenses (in
2.5 3 2.25 4 3.5 10 3.5 2.5 1.2
100,000 Birr)

Y: New clients (in 100,000) 4 5 3 5.5 4 10 5 3 1

𝑌𝑖 − 𝑌̄ -0.500 0.500 -1.500 1.000 -0.500 5.500 0.500 -1.500 -3.500

(𝑌𝑖 − 𝑌̄)2 0.250 0.250 2.250 1.000 0.250 30.250 0.250 2.250 12.250

𝑌𝑖 − 𝑌̂𝑖 0.528 1.062 -0.240 0.631 -0.404 -0.455 0.597 -0.473 -1.262

(𝑌𝑖 − 𝑌̂𝑖 )2 0.278 1.128 0.057 0.398 0.163 0.207 0.356 0.223 1.593

23 | P a g e ZigaleY.
2
Note that ∑(𝑌𝑖 − 𝑌̅𝑖 )2 = 49 and (𝑌𝑖 − 𝑌̂𝑖 )2 = 4.404. Therefore, SSmodel=∑(𝑌𝑖 − 𝑌̅𝑖 )2 − ∑(𝑌𝑖 − 𝑌̂𝑖 ) =
2 𝑆𝑆𝑚𝑜𝑑𝑒𝑙
49 − 4.404 = 44.596 and SSresidual=∑(𝑌𝑖 − 𝑌̂𝑖 ) = 4.04. k=2, and n=9, so MSmodel= 𝑘−1
=
44.596 𝑆𝑆𝑟𝑒𝑠𝑖𝑑𝑢𝑎𝑙 4.404 𝑀𝑆 𝑚𝑜𝑑𝑒𝑙 44.596
1
= 44.596 and MS residual= 𝑛−𝑘
= 9−2
= 0.629. Therefore, 𝐹 = 𝑀𝑆 𝑟𝑒𝑠𝑖𝑑𝑢𝑎𝑙 = 0.629
=

70.89. Instead of computing this manually, it makes more sense to read Stata output. Identify in Figure….
where SS model, SS residual, the degrees of freedom (df), MS model, MS residual and the F test results are

located. Figure 7: Linear regression output of new clients (in 100,000) vs advertisement expenditure (in 100,000
Birr). Pay attention to the SS, MS and the result of the F-test.
Step 2:
Compare F with the critical F-value, Fc

The critical F-value, Fc, can be found in the table, based on the degrees of freedom of the SS model (k-1)
and the degrees of freedom of SS residual (n-k). In our example, the df of SS model is 1, and the df of SS
residual is 7. Therefore, Fc with a confidence level of 10% is 3.59. Our F is 70.89 (see above, step 1), which
is far larger than Fc, so the H0 that the model is not explaining the variation in Y is strongly rejected. The
model is significant.

Using Stata output to do an F-test


Instead of using the table, and computing F manually,
econometrists rely on software, like Stata, to conduct the
F-test. As visible in Figure …, basic regression output in
Stata reports SS model, SS residual, MS model, MS
residual, F and also Prob>F. Prob>F is the p-value, the
significance or simply ‘p’. It shows the probability to get
the data from the sample, if H0 is true.

- If p is large, it is likely to get the data if X cannot


explain Y and the whole model is insignificant.

24 | P a g e ZigaleY.
- If p is small, normally smaller than 0.05, H0 is rejected and H1 is accepted. That means that the model
is significant, and X can be used to explain the variation in Y.
The result of the F-test from the output in Figure … shows that the model is significant (F(1,7):70.89,
p:0.0001), so the expenditure on advertisement (X) is significantly able to explain the variation in new
clients obtained (Y).

2.4 Reporting the Results of Regression Analysis

The results of the regression analysis derived are reported in conventional formats. It is not
sufficient merely to report the estimates of  ’s. In practice we report regression
coefficients together with their standard errors and the value of R2. It has become
customary to present the estimated equations with standard errors placed in parenthesis
below the estimated parameter values. Sometimes, the estimated coefficients, the
corresponding standard errors, the p-values, and some other indicators are presented in
tabular form.
These results are supplemented by R2 on (to the right side of the regression equation).
Y = 128.5 + 2.88 X
Example: , R2 = 0.93. The numbers in the parenthesis
(38.2) (0.85)

below the parameter estimates are the standard errors. Some econometricians report the t-
values of the estimated coefficients in place of the standard errors.

25 | P a g e ZigaleY.
26 | P a g e ZigaleY.

You might also like