Classical Regression Analysis Explained
Classical Regression Analysis Explained
In this topic we consider a simple linear regression (SLR) model, i.e., a relationship between two variables
related in a linear form. We shall first discuss two important forms of relation: stochastic and non-
stochastic, among which we shall be using the former in econometric analysis.
Non-stochastic: A relationship between X and Y, characterized as Y = f(X) is said to be
deterministic or non-stochastic if for each value of the independent variable (X) there is one and
only one corresponding value of dependent variable (Y).
Example; Q = f ( P) = + P − − − − − − − − − − − − − − − − − − − − − − − − − − − (2.1)
The above relationship between P and Q is such that for a particular value of P, there is only one
corresponding value of Q. This is, therefore, a deterministic (non-stochastic) relationship since
for each price there is always only one corresponding quantity supplied.
In non-stochastic relationship since for each price there is always only one corresponding quantity
supplied. This implies that all the variation in Y is due solely to changes in X, and that there are
no other factors affecting the dependent variable.
If this were true all the points of price-quantity pairs, if plotted on a two-dimensional plane, would
fall on a straight line. However, if we gather observations on the quantity actually supplied in the
market at various prices and we plot them on a diagram we see that they do not fall on a straight
line.
1|Page ZigaleY.
The derivation of the observation from the line may be attributed to several factors.
a. Omission of variables from the function
b. Random behavior of human beings
c. Imperfect specification of the mathematical form of the model
d. Error of aggregation
e. Error of measurement
Stochastic Relationships; a relationship between X and Y is said to be stochastic if for a particular
value of X there is a whole probabilistic distribution of values of Y. In such a case, for any given
value of X, the dependent variable Y assumes some specific value only with some probability.
In order to take into account, the above sources of errors we introduce in econometric functions a
random variable which is usually denoted by the letter ‘u’ or ‘ ’ and is called error term or random
disturbance or stochastic term of the function, so called because u is supposed to ‘disturb’ the exact
linear relationship which is assumed to exist between X and Y. By introducing this random
variable in the function, the model is rendered stochastic of the form:
Yi = + X + u i …………………………………..…………….………………….(2.2)
Thus, a stochastic model is a model in which the dependent variable is not only determined by the
explanatory variable(s) included in the model but also by others which are not included in the
model.
2.1. Simple Linear Regression model.
The above stochastic relationship (2.2) with one explanatory variable is called simple linear
regression model. The true relationship which connects the variables involved is split into two
parts: a part represented by a line and a part represented by the random term ‘u’.
2|Page ZigaleY.
The scatter of observations represents the true relationship between Y and X. The line represents
the exact part of the relationship and the deviation of the observation from the line represents the
random component of the relationship.
- Were it not for the errors in the model, we would observe all the points on the line Y1' , Y2' ,......,Yn'
Y1 , Y2 ,......,Yn corresponding to X 1 , X 2 ,...., X n . These points diverge from the regression line by
u1 , u 2 ,...., u n .
Yi = + xi + ui
the dependent var iable the regression line random var iable
The first component in the bracket is the part of Y explained by the changes in X and the second
is the part of Y not explained by X, that is to say the change in Y is due to the random influence
of u i .
3|Page ZigaleY.
the classicals worry on the parameters, the model satisfies
the assumption.
Dear distance students! Check yourself whether the following models satisfy the above assumption
and give your answer to your tutor.
a. ln Y 2 = + ln X 2 + U i
b. Yi = + X i + U i
This means that the value which “u” may assume in any one period depends on chance; it may
be positive, negative or zero. Every value has a certain probability of being assumed by ‘u’ in
any particular instance.
3). The mean value of the random variable(U) in any particular period is zero
This means that for each value of x, the random variable(u) may assume various values, some
greater than zero and some smaller than zero, but if we considered all the possible and negative
values of u, for any given value of X, they would have on average value equal to zero. In other
words, the positive and negative values of “u” cancel each other.
Mathematically, E (U i ) = 0 ………………………………………………...…. (2.3)
4). The variance of the random variable(U) is constant in each period (The assumption of
homoscedasticity)
For all values of X, the u’s will show the same dispersion around their mean. In Fig.2.c this
assumption is denoted by the fact that the values that u can assume lie with in the same limits,
irrespective of the value of X. For X 1 , u can assume any value with in the range AB; for X 2 , u
can assume any value with in the range CD which is equal to AB and so on.
Graphically;
4|Page ZigaleY.
Mathematically;
Var (U i ) = E[U i − E (U i )] 2 = E (U i ) 2 = 2 (Since E (U i ) = 0 ). This constant variance is called
homoscedasticity assumption and the constant variance itself is called homoscedastic variance.
5). The random variable (U) has a normal distribution
This means the values of u (for each x) have a bell-shaped symmetrical distribution about their
zero mean and constant variance 2 , i.e.
U i N (0, 2 ) …………………………………………..…………..……2.4
6). The random terms of different observations (U i ,U j ) are independent. (The assumption of no
autocorrelation)
This means the value which the random term assumed in one period does not depend on the value
which it assumed in any other period.
Algebraically,
Cov(u i u j ) = [(u i − (u i )][u j − (u j )]
= E (u i u j ) = 0 ………………………….…………..….(2.5)
7). The X i are a set of fixed values in the hypothetical process of repeated sampling which
samples, but the u i values do differ from sample to sample, and so of course do the values of y i .
= ( X iU i ) − ( X i )(U i )
= ( X iU i )
=0
5|Page ZigaleY.
8). The explanatory variables are measured without error
- U absorbs the influence of omitted variables and possibly errors of measurement in the y’s.
i.e., we will assume that the regressors are error free, while y values may or may not include
errors of measurement.
Dear students! We can now use the above assumptions to derive the following basic concepts.
A. The dependent variable Yi is normally distributed.
i.e Yi ~ N ( + x i ), 2 …………………………………….…………(2.7)
= + X i Since (u i ) = 0
= ( + X i + ui − ( + X i ))
2
var(Yi ) = 2 …………………………………………………….(2.8)
The shape of the distribution of Yi is determined by the shape of the distribution of u i which is
normal by assumption 4. Since and , being constant, they don’t affect the distribution of y i .
Furthermore, the values of the explanatory variable, x i , are a set of fixed values by assumption 5
Yi ~ N( + x i , 2 )
= E{[ + X i + U i − E ( + X i + U i )][ + X j + U j − E ( + X j + U j )}
(Since Yi = + X i + U i and Y j = + X j + U j )
Therefore, Cov(Yi ,Y j ) = 0 .
6|Page ZigaleY.
2.1.2 Methods of estimation
Specifying the model and stating its underlying assumptions are the first stage of any econometric
application. The next step is the estimation of the numerical values of the parameters of economic
relationships. The parameters of the simple linear regression model can be estimated by various
methods. Three of the most commonly used methods are:
1. Ordinary least square method (OLS)
2. Maximum likelihood method (MLM)
3. Method of moments (MM)
But here we will deal with the OLS methods of estimation.
[Link] The ordinary least square (OLS) method
The model Yi = + X i + U i is called the true relationship between Y and X because Y and X
represent their respective population value, and and are called the true parameters since they
are estimated from the population value of Y and X But it is difficult to obtain the population value
of Y and X because of technical or economic reasons. So, we are forced to take the sample value
of Y and X. The parameters estimated from the sample value of Y and X are called the estimators
of the true parameters and and are symbolized as ˆ and ˆ .
The model Yi = ˆ + ˆX i + ei , is called estimated relationship between Y and X since ˆ and ˆ
are estimated from the sample of Y and X and ei represents the sample counterpart of the
Estimation of and by least square method (OLS) or classical least square (CLS) involves
finding values for the estimates ˆ and ˆ which will minimize the sum of square of the squared
residuals ( ei2 ).
e 2
i = (Yi − ˆ − ˆX i ) 2 ……………………….(2.7)
To find the values of ˆ and ˆ that minimize this sum, we have to partially differentiate e 2
i
with respect to ˆ and ˆ and set the partial derivatives equal to zero.
7|Page ZigaleY.
ei2
1. = −2 (Yi − ˆ − ˆX i ) = 0.......................................................(2.8)
ˆ
ei2
2. = −2 X i (Yi − ˆ − ˆX ) = 0..................................................(2.11)
ˆ
Note: at this point that the term in the parenthesis in equation 2.8and 2.11 is the residual,
e = Yi − ˆ − ˆX i . Hence it is possible to rewrite (2.8) and (2.11) as − 2 ei = 0 and
− 2 X i ei = 0 . It follows that;
e i = 0 and X e i i = 0............................................(2.12)
Equation (2.9) and (2.13) are called the Normal Equations. Substituting the values of ̂ from
(2.10) to (2.13), we get:
Y X i i = X i (Y − ˆX ) + ˆX i2
Y X i i − Y X i = ˆ (X i2 − XX i )
XY − nXY = ˆ ( X i2 − nX 2)
XY − nXY
ˆ = ………………….(2.14)
X i2 − nX 2
8|Page ZigaleY.
−
( X − X )(Y − Y ) = XY − n X Y − − − − − − − − − − − − − −(2.15)
( X − X ) 2 = X 2 − nX 2 − − − − − − − − − − − − − − − − − (2.16)
Substituting (2.15) and (2.16) in (2.14), we get
( X − X )(Y − Y )
ˆ =
( X − X ) 2
xi yi
ˆ = ……………………………………… (2.17)
xi2
The expression in (2.17) to estimate the parameter coefficient is termed is the formula in deviation
form.
2.1.3 Estimation of a function with zero intercept
Suppose it is desired to fit the line Yi = + X i + U i , subject to the restriction = 0. To estimate
ˆ , the problem is put in a form of restricted minimization problem and then Lagrange method is
applied.
n
We minimize: ei2 = (Yi − ˆ − ˆX i ) 2
i =1
Subject to: ˆ = 0
The composite function then becomes
Z = (Yi − ˆ − ˆX i ) 2 − ˆ , where is a Lagrange multiplier.
Yi X i − ˆX i = 0
2
9|Page ZigaleY.
X i Yi
ˆ = ……………………………………..(2.18)
X i2
This formula involves the actual values (observations) of the variables and not their deviation
forms, as in the case of unrestricted value of ˆ .
2.1.4. Statistical Properties of Least Square Estimators
There are various econometric methods with which we may obtain the estimates of the parameters
of economic relationships. We would like to an estimate to be as close as the value of the true
population parameters i.e., to vary within only a small range around the true parameter. How are
we to choose among the different econometric methods, the one that gives ‘good’ estimates? We
need some criteria for judging the ‘goodness’ of an estimate.
‘Closeness’ of the estimate to the popular;8tion parameter is measured by the mean and variance
or standard deviation of the sampling distribution of the estimates of the different econometric
methods. We assume the usual process of repeated sampling i.e., we assume that we get a very
large number of samples each of size ‘n’; we compute the estimates ˆ ’s from each sample, and
for each econometric method and we form their distribution. We next compare the mean (expected
value) and the variances of these distributions and we choose among the alternative estimates the
one whose distribution is concentrated as close as possible around the population parameter.
PROPERTIES OF OLS ESTIMATORS
The ideal or optimum properties that the OLS estimates possess may be summarized by well-
known theorem known as the Gauss-Markov Theorem.
Statement of the theorem: “Given the assumptions of the classical linear regression model, the
OLS estimators, in the class of linear and unbiased estimators, have the minimum variance, i.e.
the OLS estimators are BLUE.
According to this theorem, under the basic assumptions of the classical linear regression model,
the least squares estimators are linear, unbiased and have minimum variance (i.e. are best of all
linear unbiased estimators). Sometimes the theorem referred as the BLUE theorem i.e. Best,
Linear, Unbiased Estimator. An estimator is called BLUE if:
a. Linear: a linear function of a random variable, such as, the dependent variable Y.
b. Unbiased: its average or expected value is equal to the true population parameter.
10 | P a g e ZigaleY.
c. Minimum variance: It has a minimum variance in the class of linear and unbiased
estimators. An unbiased estimator with the least variance is known as an efficient
estimator.
According to the Gauss-Markov theorem, the OLS estimators possess all the BLUE properties.
The detailed proof of these properties is read by yourself:
Example of applying OLS
The marketing manager of Safaricom, the new competitor of Ethio Telecom in the Ethiopian
telecommunication market wants to know the relation between expenses on advertisement and
new clients. To this end, the team collects the following data:
Month January February April May June July August Sept Oct
X: Advertisement
expenses (in 100,000 2.5 3 2.25 4 3.5 10 3.5 2.5 1.2
Birr)
Y: New clients (in
4 5 3 5.5 4 10 5 3 1 Sum
100,000)
-
𝑋𝑖 − 𝑋̄
-1.11 -0.61 -1.36 0.39 0.11 6.39 -0.11 -1.11 -2.41
-
𝑌𝑖 − 𝑌̄
-0.50 0.50 -1.50 1.00 0.50 5.50 0.50 -1.50 -3.50
(𝑋𝑖 − 𝑋̄) ∗ ( 𝑌𝑖 − 𝑌̄) 0.55 -0.30 2.03 0.39 0.05 35.17 -0.05 1.66 8.42 47.93
(𝑋𝑖 − 𝑋̄)2 1.22 0.37 1.84 0.16 0.01 40.89 0.01 1.22 5.79 51.50
Note that the relationship displayed in Figure 3 indicates a positive linear relation between
advertisement expenses and new clients.
11 | P a g e ZigaleY.
Now, let’s use the Least Sum of Squares
approach to estimate 𝛼 and 𝛽, or, in other
words, find 𝛼̂ and 𝛽̂ . Firstly, we need to
find the mean of X (𝑋̄) and the mean of Y
(𝑌̄).
2.5+3+2.25+4+3.5+10+3.5+2.5+1.2
𝑋̄ = =
9
3.606
Next, we need to find (𝑋𝑖 − 𝑋̄)(𝑌𝑖 − 𝑌̄) and (𝑋𝑖 − 𝑋̄)2 for each i. These are displayed in the table
above. We find that the sum of (𝑋𝑖 − 𝑋̄)2 is 51.5, and the sum of (𝑋𝑖 − 𝑋̄)(𝑌𝑖 − 𝑌̄) is 47.93. This,
we use to calculate 𝛽̂ .
𝛴(𝑋𝑖 − 𝑋̄)(𝑌𝑖 − 𝑌̄) 47.93
𝛽̂ = = = 0.9305
𝛴(𝑋𝑖 − 𝑋̄)2 51.5
That means that, if advertisement expenses increase by 100,000 Birr, Safaricom gets, on average,
93,050 extra clients. Next, we can calculate 𝛼̂
𝛼̂ = 𝑌̅ − 𝛽̂ 𝑋̅ = 4.5 − 0.9305 ∗ 3.606 = 1.1449
Summarizing, the relation between advertisement expenses (X) and extra clients (Y) for Safaricom
can be expressed by the following equation:
𝑌𝑖 = 1.1449 + 0.9305𝑋𝑖 + 𝑢𝑖
This information is also visible in Figure 4, which shows Stata output of the Simple Linear
Regression. Note that 𝛼̂ is the _cons value in the output, and 𝛽̂ is the coefficient for the variable
expenditure. The other elements of this Stata output will be discussed in the sections that follow.
12 | P a g e ZigaleY.
Figure 2: Stata output for SLR with Y= new clients (in 100,000) and X=expenditure on
advertisement (in 100,000 Birr)
2.2 Covariance, correlation coefficient, coefficient of determination (R2)
Simple linear regression is used to analyze two continuous variables that are linearly related. To see how
strong this relation is, and if it is a positive or a negative relation, it is common to calculate measures of
association: covariance, correlation and the coefficient of determination.
The covariance between X and Y measures if X and Y are positively, negatively or not related. If COV
(X, Y)>0, X and Y are positively related, if COV (X, Y) <0, X and Y are negatively related and if COV
(X, Y) is 0, or close to 0, there is no relation between X and Y. The mathematical expression for
covariance is:
1
𝐶𝑂𝑉(𝑋, 𝑌) = ∑(𝑋𝑖 − 𝑋̄)(𝑌𝑖 − 𝑌̄)
𝑛−1
A disadvantage of covariance is that it is sensitive to the scale of the data and this makes it difficult to
interpret. It is not clear if a COV (X, Y) of 4 is high or low, because it depends on how large X and Y are.
To compensate for this limitation, a more common way to evaluate the relation between two continuous
variables is Pearson’s correlation. Pearson’s correlation is a figure between -1 and 1. If CORR(X,Y)=0,
there is no relation between X and Y, if CORR(X,Y)=1 there is a perfect positive relation between X and
Y and if CORR(X,Y)=-1 the is a perfect negative relation between X and Y. When working with statistical
data, perfect relations are uncommon, and therefore the correlation coefficient usually takes a value between
-1 and 1. Consider, for example, the scatterplots in Figure 5. As a rule of thumb, if CORR(X,Y) is between
0 and 0.3 (or -0.3 and 0), there is no relation, if CORR(X,Y) is between 0.3 and 0.5 (or between -0.5 and -
0.3) the correlation is weak, if CORR(X,Y) is between 0.5 and 0.8 (or between -0.8 and -0.5) the correlation
13 | P a g e ZigaleY.
is moderate and if CORR(X,Y)>0.8 (or <-0.8) the correlation is strong. The mathematical expression for
correlation is:
𝐶𝑂𝑉(𝑋, 𝑌)
𝐶𝑂𝑅𝑅(𝑋, 𝑌) =
√𝑉𝐴𝑅(𝑋) ∗ 𝑉𝐴𝑅(𝑌)
1 1
With COV(X,Y) as defined above, VAR(X)= 𝑛−1 ∑(𝑋𝑖 − 𝑋̄)2 and VAR(Y)= 𝑛−1 ∑(𝑌𝑖 − 𝑌̄)2
Figure 3: Various scatterplots displaying correlations that differ in sign (+ or -) and magnitude/strength.
The correlation is easier to interpret than the covariance, but when discussing Simple Linear Regression, it
is not the best way to present the relation between X and Y. To express the explanatory power that X has
to explain the variation in Y, it is best to calculate the coefficient of determination, the R2 (or R-squared).
R-squared is a measure of how much of the variance in the actual value of dependent variable (Y) can be
explained by its relationship to the independent variable (X). In other words, R-squared is the percentage
of variance in Y explained by the linear regression equation between X and Y. The R-squared can take any
value in the range [-∞, 1], but usually it takes values between 0 (0% of the variation in Y is explained by
the model) and 1 (the model explains 100% of the variation in Y). The closer the value to 1, the better the
explanatory power of the independent variable is. The mathematical expression for the R-squared is:
𝑆𝑆𝑟𝑒𝑠𝑖𝑑𝑢𝑎𝑙
𝑅2 = 1 −
𝑆𝑆𝑡𝑜𝑡𝑎𝑙
Where SS residual is the sums of squares of the residual term from the linear regression
2
(∑(𝑌𝑖 − 𝑌̂𝑖 ) = ∑ 𝑢𝑖 2 ) and SS total is the total sum of squares of the dependent variable Y (∑(𝑌𝑖 − 𝑌̅)2 ).
14 | P a g e ZigaleY.
From the equation we can see that if the error term is close to 0, R-squared will be close to 1. If, on the
other hand, the residual is relatively large, the R-squared is small.
space, it is given that 𝛼̂ = −19.55 and 𝛽̂ = 0.95. These values have been found by using the procedure as
explained in section 2.2.
Entrepreneurship 88 76 65 63 88 53 74 70 89 98
grade (X)
Business profit after a 48 50 41 58 50 37 39 37 69 100
year (Y) in 1000 Birr
(𝑋𝑖 − 𝑋̄) 11.6 -0.4 -11.4 -13.4 11.6 -23.4 -2.4 -6.4 12.6 21.6
(𝑌𝑖 − 𝑌̄) -4.9 -2.9 -11.9 5.1 -2.9 -15.9 -13.9 -15.9 16.1 47.1
(𝑋𝑖 − 𝑋̄)2 134.6 0.2 130.0 179.6 134.6 547.6 5.8 41.0 158.8 466.6
(𝑌𝑖 − 𝑌̄)2 24.0 8.4 141.6 26.0 8.4 252.8 193.2 252.8 259.2 2218.4
̂𝑖 = −19.55 + 0.95𝑋𝑖
𝑌 64.1 52.7 42.2 40.3 64.1 30.8 50.8 47.0 65.0 73.6
𝑢𝑖 = 𝑌𝑖 − 𝑌̂𝑖 -16.1 -2.7 -1.2 17.7 -14.1 6.2 -11.8 -10.0 4.0 26.5
𝑢𝑖 2 257.6 7.0 1.4 313.3 197.4 38.4 138.1 99.0 16.0 699.6
(𝑋𝑖 − 𝑋̄)(𝑌𝑖 − 𝑌̄) -56.8 1.2 135.7 -68.3 -33.6 372.1 33.4 101.8 202.9 1017.4
Based on the information in the table, we can calculate COV (X, Y), CORR (X, Y) and the coefficient of
determination:
1 1
• 𝐶𝑂𝑉(𝑋, 𝑌) = ∑(𝑋𝑖 − 𝑋̄)(𝑌𝑖 − 𝑌̄) = ∗ 1705.4 = 189.5, which means that X and Y are
𝑛 10−1
relation between X and Y. Positive because the correlation is >0, and moderate because the
correlation is between 0.5 and 0.8.
15 | P a g e ZigaleY.
𝑆𝑆𝑅 ∑𝑢 2 1767.9
• 𝑅 2 = 1 − 𝑆𝑆𝑇 = 1 − ∑(𝑌 −𝑌
𝑖
̅ )2
= 1 − 3384.9 = 0.478, which indicates that 47.8% of the variation in
𝑖
business profit (Y) can be explained by the score of the entrepreneurship course (X).
Note that the information for the coefficient of determination is also displayed in the Stata output, see Figure
6.
Figure 4: upper part of Stata regression output, displaying the R-squared of the model.
The OLS estimates ˆ and ˆ are obtained from a sample of observations on Y and X. Since
sampling errors are inevitable in all estimates, it is necessary to apply test of significance
in order to measure the size of the error and determine the degree of confidence in order to
measure the validity of these estimates. This can be done by using various tests. The most
common ones are:
i) Standard error test ii) Student’s t-test iii) Confidence interval and F test
All of these testing procedures reach on the same conclusion. Let us now see these testing
methods one by one.
i) Standard error test
This test helps us decide whether the estimates ˆ and ˆ are significantly different from
zero, i.e., whether the sample from which they have been estimated might have come from
a population whose true parameters are zero. = 0 and / or = 0 .
Formally we test the null hypothesis
H 0 : i = 0 against the alternative hypothesis H 1 : i 0
16 | P a g e ZigaleY.
SE ( ˆ ) = var(ˆ )
SE(ˆ ) = var(ˆ )
Second: compare the standard errors with the numerical values of ˆ and ˆ .
Decision rule:
• If SE(ˆi ) 1 2 ˆi , accept the null hypothesis and reject the alternative hypothesis.
• If SE(ˆi ) 1 2 ˆi , reject the null hypothesis and accept the alternative hypothesis.
Test the significance of the slope parameter at 5% level of significance using the standard
error test.
SE ( ˆ ) = 0.025
( ˆ ) = 0.6
1
2 ˆ = 0.3
This implies that SE(ˆi ) 1 2 ˆi . The implication ˆ is statistically significant at 5% level
of significance.
17 | P a g e ZigaleY.
Note: The standard error test is an approximated test (which is approximated from the z-
test and t-test) and implies a two-tail test conducted at 5% level of significance.
ii). Student’s t-test
If, in linear regression, X has no effect on Y, is (close to) 0. We can use the t-test (as you learned in the
last course: Statistics II), to test if 𝛽̂ from the sample data significantly deviates from 0.
The steps for conducting a student’s t-test for a regression coefficient are as follows:
Step 1: Compute t
To compute t, we need to know the value of 𝛽̂ and its “uncertainty”, which is measured by its standard
deviation (𝑆𝐸(𝛽̂ )). With 𝛽̂ and 𝑆𝐸(𝛽̂ ), it is possible to compute the t- test statistic, using the following
formula:
̂
𝛽
𝑡𝛽̂ = ̂) with (n-k) degrees of freedom
𝑆𝐸(𝛽
where 𝛽̂ is the coefficient for X based on the least squares estimation (see section 2.2) and 𝑆𝐸(𝛽̂ ) is the
1 ∑(𝑌 −𝑌 ) ̂ 2
standard error of this estimate, which is 𝑆𝐸(𝛽̂ )=√𝑛−2 ∗ ∑(𝑋𝑖 −𝑋̅𝑖)2 . Instead of calculating this by hand, you
𝑖
can read it from the Stata regression output. n is the sample size, and k is the number of parameters in the
model (k= 2 in SLR).
Step 3: Decide whether you want to conduct a one-tailed test, or a two-tailed test
Either you want to see if X affects Y (positively or negatively), for which you’ll do a two-tailed test with
H0: =0 and H1: ≠0. This is the most common way to test a coefficient in linear regression.
However, in some cases economic/management theory expects you to see if X affects Y positively only
(not negatively). For example, to test the law of supply with Y is demand and X is price, it makes sense to
test if is significantly higher than 0 (rather than higher or lower than 0). In that case you’ll do a one-tailed
test, with H0:=0 and H1: >0.
18 | P a g e ZigaleY.
Of course, it is also possible to see if X affects Y negatively. In that case H0:=0 and H1: <0 and the test
is also one-tailed. This should, for example, be used to test the law of demand from economics, which
states that an increased price has a negative effect on the quantity demanded.
If the outcome of step one, t > 2.048 (or t< -2.048), H0 is rejected and H1 is accepted. That means that ≠0,
so X has a significant effect on Y. In other words, if H0 is rejected, is statistically significant. If the
outcome of step 1 is less than 2.048, H0 is not rejected and is statistically insignificant.
19 | P a g e ZigaleY.
Figure 5: SLR output with Y=new clients (in 100,000) and X=advertisement expenditure (in 100,000 Birr)
For the output in Figure 7, which shows the least square estimation results for the effect of advertisement
expenditure on new clients for Safaricom, 𝛽̂ is 0.93 (t: 8.48, p:0.000), so the effect of expenditure on new
clients is significant (because p<0.05). 𝛼̂ is 1.14 (t:2.39, p:0.048), so the intercept of the model also
significantly deviates from 0 (because p<0.05). In the presentation of regression coefficients in reports or
other publications, it is common to report the coefficient, it’s t-value and the significance level (p) of the
coefficient.
20 | P a g e ZigaleY.
In a two-tail test at level of significance, the probability of obtaining the specific t-value
either –tc or tc is
2 at n-2 degree of freedom. The probability of obtaining any value of t
ˆ −
which is equal to at n-2 degree of freedom is 1 − ( 2 + 2 ) i.e. 1 − .
ˆ
SE ( )
ˆ −
but t* = …………………………………………………….(2.58)
SE ( ˆ )
Pr − SE(ˆ )t c ˆ − SE(ˆ )t c = 1 − − − − − − by multiplying SE(ˆ )
The limit within which the true lies at (1 − )% degree of confidence is:
H1 : 0
Decision rule: If the hypothesized value of in the null hypothesis is within the
confidence interval, accept H0 and reject H1. The implication is that ˆ is statistically
insignificant; while if the hypothesized value of in the null hypothesis is outside the
21 | P a g e ZigaleY.
Y = 128.5 + 2.88 X + e
(38.2) (0.85)
ˆ SE( ˆ )t c
ˆ = 2.88
SE ( ˆ ) = 0.85
22 | P a g e ZigaleY.
Figure 6: Comparison of the full model and the restricted model
The steps for conducting an F test for a regression coefficient are as follows:
Step 1: Compute F
𝑆𝑆𝑚𝑜𝑑𝑒𝑙/𝑘 − 1 𝑀𝑆 𝑚𝑜𝑑𝑒𝑙
𝐹= =
𝑆𝑆𝑟𝑒𝑠𝑖𝑑𝑢𝑎𝑙/𝑛 − 𝑘 𝑀𝑆 𝑟𝑒𝑠𝑖𝑑𝑢𝑎𝑙
Where: SS model is ∑(𝑌𝑖 − 𝑌̅𝑖 )2 − ∑(𝑌𝑖 − 𝑌̂𝑖 )2, SS residual is ∑(𝑌𝑖 − 𝑌̂𝑖 )2 . When divided by the
degrees of freedom (k-1 for the model and n-k for the residuals), you get the mean sum of squared
(MS). Dividing the MS of the model by the MS residual, gives us the F-value.
Consider this example, based on data about new clients and advertisement expenditure, with 𝑌̅ =
4.5, 𝛼̂ = 1.145 and 𝛽̂ = 0.931, so 𝑌̂𝑖 = 1.145 + 0.931 ∗ 𝑋𝑖 ).
Month January February April May June July August Sept Oct
X: Advertisement expenses (in
2.5 3 2.25 4 3.5 10 3.5 2.5 1.2
100,000 Birr)
(𝑌𝑖 − 𝑌̄)2 0.250 0.250 2.250 1.000 0.250 30.250 0.250 2.250 12.250
𝑌𝑖 − 𝑌̂𝑖 0.528 1.062 -0.240 0.631 -0.404 -0.455 0.597 -0.473 -1.262
(𝑌𝑖 − 𝑌̂𝑖 )2 0.278 1.128 0.057 0.398 0.163 0.207 0.356 0.223 1.593
23 | P a g e ZigaleY.
2
Note that ∑(𝑌𝑖 − 𝑌̅𝑖 )2 = 49 and (𝑌𝑖 − 𝑌̂𝑖 )2 = 4.404. Therefore, SSmodel=∑(𝑌𝑖 − 𝑌̅𝑖 )2 − ∑(𝑌𝑖 − 𝑌̂𝑖 ) =
2 𝑆𝑆𝑚𝑜𝑑𝑒𝑙
49 − 4.404 = 44.596 and SSresidual=∑(𝑌𝑖 − 𝑌̂𝑖 ) = 4.04. k=2, and n=9, so MSmodel= 𝑘−1
=
44.596 𝑆𝑆𝑟𝑒𝑠𝑖𝑑𝑢𝑎𝑙 4.404 𝑀𝑆 𝑚𝑜𝑑𝑒𝑙 44.596
1
= 44.596 and MS residual= 𝑛−𝑘
= 9−2
= 0.629. Therefore, 𝐹 = 𝑀𝑆 𝑟𝑒𝑠𝑖𝑑𝑢𝑎𝑙 = 0.629
=
70.89. Instead of computing this manually, it makes more sense to read Stata output. Identify in Figure….
where SS model, SS residual, the degrees of freedom (df), MS model, MS residual and the F test results are
located. Figure 7: Linear regression output of new clients (in 100,000) vs advertisement expenditure (in 100,000
Birr). Pay attention to the SS, MS and the result of the F-test.
Step 2:
Compare F with the critical F-value, Fc
The critical F-value, Fc, can be found in the table, based on the degrees of freedom of the SS model (k-1)
and the degrees of freedom of SS residual (n-k). In our example, the df of SS model is 1, and the df of SS
residual is 7. Therefore, Fc with a confidence level of 10% is 3.59. Our F is 70.89 (see above, step 1), which
is far larger than Fc, so the H0 that the model is not explaining the variation in Y is strongly rejected. The
model is significant.
24 | P a g e ZigaleY.
- If p is small, normally smaller than 0.05, H0 is rejected and H1 is accepted. That means that the model
is significant, and X can be used to explain the variation in Y.
The result of the F-test from the output in Figure … shows that the model is significant (F(1,7):70.89,
p:0.0001), so the expenditure on advertisement (X) is significantly able to explain the variation in new
clients obtained (Y).
The results of the regression analysis derived are reported in conventional formats. It is not
sufficient merely to report the estimates of ’s. In practice we report regression
coefficients together with their standard errors and the value of R2. It has become
customary to present the estimated equations with standard errors placed in parenthesis
below the estimated parameter values. Sometimes, the estimated coefficients, the
corresponding standard errors, the p-values, and some other indicators are presented in
tabular form.
These results are supplemented by R2 on (to the right side of the regression equation).
Y = 128.5 + 2.88 X
Example: , R2 = 0.93. The numbers in the parenthesis
(38.2) (0.85)
below the parameter estimates are the standard errors. Some econometricians report the t-
values of the estimated coefficients in place of the standard errors.
25 | P a g e ZigaleY.
26 | P a g e ZigaleY.