Chapter 2
Chapter 2
CHAPTER TWO
SIMPLE LINEAR REGRESSION
As you know, financial and theories are mainly concerned with the relationships between variables. This chapter
introduces the simplest possible regression analysis, namely, the two variables regression in which the dependent
variable is linearly related to one independent explanatory variable. This case is considered here, because it
presents the fundamental idea of regression analysis as simply as possible. We shall first discuss the basic concept
of regression, and then proceed to the core of simple linear regression analysis.
Regression analysis is concerned with the study of the dependence of one variable (the dependent variable) on
one or more other variables (the explanatory variable(s)). In other words, regression analysis is concerned with
describing and evaluating the relationship between a given variable (the dependent variable) and one or
more other variables (the independent variable(s)). The objective of regression analysis is to estimate and/or
predict the unknown (population) mean value of the dependent variable given known values of the explanatory
variables.
In Galton’s example, this is to estimate/predict the mean height of children (dependent variable), given the height
of their parents (explanatory variables). Galton’s example is not related to finance or economics, but the following
example is. A financial expert may be interested in studying the dependence of household monthly savings on
household monthly disposable income. That is, we want to predict average savings, knowing household monthly
disposable income. Such an analysis is helpful in estimating the marginal propensity to save (MPS), that is,
average change in saving for, say, a unit change in disposable income. To see how this can be done, consider
Figure 2.1.
As discussed in section 1.4, regression analysis deals with statistical dependence among variables, but not with
deterministic dependence among variables. In statistical relationships, we essentially deal with random
(stochastic) variables, i.e., variables that have probability distributions.
Page 1
CHAPTER TWO: SIMPLE LINEAR REGRESSION 2022
Figure 2.1 shows the distribution of household monthly savings in a hypothetical population corresponding to the
given or fixed values of the household monthly disposable income. Notice that corresponding to any given
household income is a range or distribution of the household savings. However, notice that despite the variability
of savings for a given value of household income, the average savings generally increases as the income of
household increases.
The line that passes through the average level of consumption expenditure for each level of household income is
known as the regression line. It shows how the average consumption expenditure increases with the household’s
income.
Regression Line
Figure 2.1: Scatter plot diagram with savings as dependent variable and disposable income as explanatory variable
In econometrics we exclusively deal with stochastic relationships. Although regression analysis deals with the
dependence of one variable on other variables, it does not necessarily imply causation. For example, consider
the regression line in Figure 2.2 which displays the relationship between hot drink sales and number of customers
of a resort. If it is busy in the resort (number of customers approaches 100-150, see the x-axis), the tea sales are
low. But if the number of customers is low, the tea sales are high. Does this mean that the number of customers
negatively affects tea sales? That would not make sense from a theoretical perspective. In this case, there is a third
variable, the weather, which affects both the tea sales and the number of customers. If the weather is cold, less
customers will come to the resort and those who come will likely order tea. If the weather is hot, more customers
will come to the resort, and they will likely order cold drinks (not tea).
Page 2
50
40
30
20
10 CHAPTER TWO: SIMPLE LINEAR REGRESSION 2022
0
0 50 100 150
number of customers
Figure 2.2: The scatter plot diagram with tea sales as dependent variable and number of customers as explanatory variable
Therefore, the determination of the direction of causation should come from outside of statistics, i.e., economic
or financial theory. In other words, statistical relationships by themselves cannot logically imply causation. To
ascribe causality, one must appeal to ‘a priori’ or theoretical considerations.
The variables in a regression relation consist of dependent and explanatory variables. The dependent variable is
the variable whose variation is being explained by the other variable(s). The explanatory variable is the variable
whose variation is used to explain the variation in the dependent variable. Other words for the dependent variable
are explained variable, or regressand, or, simply the letter Y. Other words for the explanatory variable are
independent variable, regressor, or, simply the letter X.
and are true population parameters, but research rarely has sufficient resources to check values for the whole
population. Therefore, sample data is used to estimate/calculate 𝛼̂ (which is the best fitting given the sample
data) and 𝛽̂ (which is the best fitting given the sample data).
Page 3
CHAPTER TWO: SIMPLE LINEAR REGRESSION 2022
Consider this example: Cloth production is upcoming in Ethiopia. Some factories manage to make good profit,
while others make very small or no profits. Hana, a financial expert, expects that some make no profit because
they do not optimally use their workers and equipment (sewing machines). Therefore, she expects that technical
efficiency determines how much profit the cloth producing factory makes. In Hana’s analysis the dependent
variable is profitability (Yi) and the explanatory variable is technical efficiency (Xi). Profitability is measured as
𝑡𝑜𝑡𝑎𝑙 𝑝𝑟𝑜𝑓𝑖𝑡
net profit margin ratio (which is ). Technical efficiency is measured as
𝑡𝑜𝑡𝑎𝑙 𝑟𝑒𝑣𝑒𝑛𝑢𝑒
𝑎𝑐𝑡𝑢𝑎𝑙 𝑝𝑟𝑜𝑑𝑢𝑐𝑡𝑖𝑜𝑛 𝑜𝑢𝑡𝑝𝑢𝑡
. Note that both X and Y are ratios, which Hana collected
𝑚𝑎𝑥 𝑝𝑜𝑠𝑠𝑖𝑏𝑙𝑒 𝑝𝑟𝑜𝑑𝑢𝑐𝑡𝑖𝑜𝑛 𝑜𝑢𝑡𝑝𝑢𝑡 𝑔𝑖𝑣𝑒𝑛 𝑒𝑞𝑢𝑖𝑝𝑚𝑒𝑛𝑡 𝑎𝑛𝑑 𝑙𝑎𝑏𝑜𝑟
Page 4
CHAPTER TWO: SIMPLE LINEAR REGRESSION 2022
Figure 2.3: Three types of relations between technical efficiency and profitability
Assumption 4: The variance of the error term (𝑼𝒊 ) is constant across observations
Another name of this assumption is the assumption of homoscedasticity. This means that, for all values of X, the
U’s will show the same dispersion around their mean. This implies that, given the value of X, the variance of 𝑈𝑖
is the same (constant) for all observations. Figure 2.4 displays a situation in which this assumption is met. Figure
Page 5
CHAPTER TWO: SIMPLE LINEAR REGRESSION 2022
2.5 shows a situation where the spread in U increases for higher values of X. Note: this assumption implies that
the values of Y corresponding to various values of X have constant variance.
Figure 2.4: Heteroskedastic Variance--- spread of errors depends on X Figure 2.5: Homoscedastic variance - the error has constant variance
Page 6
CHAPTER TWO: SIMPLE LINEAR REGRESSION 2022
Figure 2.6: The effect of omitting a relevant variable. Note that the estimated line over-estimates X's true impact on Y
Thus, to have unbiased estimator, the exogeneity assumption is important. If two variables are unrelated their
covariance shall be zero.
𝑪𝒐𝒗(𝑿𝒊 𝑼𝒊 ) = 𝟎 (𝟐. 𝟓)
Assumption 8: Variability in X values
The x values in a given sample must not all be the same. This means that x assumes different values in a given
sample; but it assumes fixed values in hypothetically repeated samples. This assumption is very critical since
without this assumption it would be impossible to estimate the parameters and hence, regression analysis would
fail. For example, if there is little variation in household income, we will not be able to explain much of the
variation in the consumption expenditure of the households.
Assumption 9: The distribution of the dependent variable, Y
Based on the assumptions we discussed so far about the distributions of 𝑿 and 𝑼, we can establish that Y is
normally distributed with Mean: 𝑬(𝒀𝒊 ) = 𝜶 + 𝜷𝑿𝒊 , and variance: 𝑽𝒂𝒓(𝒀𝒊 |𝑿𝒊 ) = 𝑽𝒂𝒓(𝑼𝒊 ) = 𝝈𝟐 .
Page 7
CHAPTER TWO: SIMPLE LINEAR REGRESSION 2022
𝒀𝒊 = 𝜶 ̂ 𝑿𝒊 + 𝒆𝒊
̂ +𝜷 (2.6)
Note that 𝛼 changes into 𝛼̂, 𝛽 changes into 𝛽̂ and 𝑈𝑖 changes into 𝑒𝑖 because we do our analysis with sample data.
OLS is the technique used to estimate a line that will minimize the error (the difference between the predicted
and the actual values of a dependent variable, i.e. the e).
From the estimated relationship, 𝑌𝑖 = 𝛼̂ + 𝛽̂ 𝑋𝑖 + 𝑒𝑖 , we obtain 𝑒𝑖 = 𝑌𝑖 − (𝛼̂ + 𝛽̂ 𝑋𝑖 ). That means that we should
find the values for 𝛼̂ and 𝛽̂ which minimize ∑ 𝑒𝑖2 = ∑(𝑌𝑖 − 𝛼̂ − 𝛽̂ 𝑋𝑖 )2 . Therefore, we must partially differentiate
∑ 𝑒𝑖2 with respect to 𝛼̂ and 𝛽̂ and set the partial derivatives equal to zero.
Partial derivative with respect to 𝛼̂:
𝝏 ∑ 𝒆𝟐𝒊
= −𝟐 ∑(𝒀𝒊 − 𝜶 ̂ 𝑿𝒊 ) = 𝟎
̂−𝜷 (2.7)
̂
𝝏𝜶
Or, rearranged, normal equation 1: ∑ 𝒀𝒊 = 𝒏𝜶 ̂ 𝜮𝑿𝒊 . Note that, if we divide both sides by n and rearrange,
̂+𝜷
normal equation 1 is
̂=𝒀
𝜶 ̂𝑿
̅−𝜷 ̅ (2.8)
Partial derivative with respect to 𝛽̂ :
𝝏 ∑ 𝒆𝟐𝒊
= −𝟐 ∑ 𝑿𝒊 (𝒀𝒊 − 𝜶 ̂ 𝑿𝒊 ) = 𝟎
̂−𝜷 (2.9)
̂
𝝏𝜷
Page 8
CHAPTER TWO: SIMPLE LINEAR REGRESSION 2022
̂ (𝜮𝑿𝟐𝒊 − 𝑿̄𝜮𝑿𝒊 )
∑ 𝒀𝒊 𝑿𝒊 − 𝒀̄𝜮𝑿𝒊 = 𝜷
̂ (𝜮𝑿𝟐𝒊 − 𝒏𝑿̄2)
𝜮𝑿𝒊 𝒀𝒊 − 𝒏𝑿̄𝒀̄ =𝜷
̂ = 𝜮𝑿𝒊 𝒀𝟐 𝒊 −𝒏𝑿̄𝒀̄
𝜷 𝜮𝑿 −𝒏𝑿̄𝟐
𝒊
or, rearranged,
̄ ̄
̂ = 𝜮(𝑿𝒊−𝑿)(𝒀𝒊−𝒀)
𝜷 (2.11)
̄ 𝟐
𝜮(𝑿𝒊 −𝑿)
Given consumption (Y) and income (X) data, both in thousands of Birr, of six households, obtain the OLS
estimators of 𝛼 and 𝛽.
42 54
𝑌̅ = = 7, and 𝑋̅ = = 9
6 6
Observations 𝑌𝑖 𝑋𝑖 𝑌𝑖 𝑋𝑖 𝑋𝑖2 𝑦𝑖 = 𝑌𝑖 − 𝑌̅ 𝑥𝑖 = 𝑋𝑖 − 𝑋̅ (𝑌𝑖 − 𝑌̅) (𝑋𝑖 − 𝑋̅)2
∗
(𝑋𝑖 − 𝑋̅)
1. 4 5 20 25 -3 -4 12 16
2. 4 4 16 16 -3 -5 15 25
3. 7 8 56 64 0 -1 0 1
4. 8 10 80 100 1 1 1 1
5. 9 13 117 169 2 4 8 16
6. 10 14 140 196 3 5 15 25
Sums 42 54 429 570 0 0 51 84
∑(𝑋𝑖 − 𝑋̅)(𝑌𝑖 − 𝑌̅) 51
𝛽̂ = , 𝛽̂ = = 0.607
∑(𝑋𝑖 − 𝑋̅)2 84
𝛼̂ = 𝑌̅ − 𝛽̂ 𝑋̅ ∴ 𝛼̂ = 7 − 0.607(9) = 1.53
Therefore, the fitted regression equation (i.e., OLS regression line) is:
Page 9
CHAPTER TWO: SIMPLE LINEAR REGRESSION 2022
̂𝑖 = 1.53 + 0.607𝑖𝑛𝑐𝑖
𝐶𝑜𝑛𝑠
Example 2: Effect of advertisement on number of customers
The marketing manager of Safaricom, the new competitor of EthioTelecom in the Ethiopian telecommunication
market wants to know the relation between expenses on advertisement and new clients. To this end, the team
collects the following data:
Month January February April May June July August Sept Oct
X: Advertisement
2.5 3 2.25 4 3.5 10 3.5 2.5 1.2
expenses (in 100,000 Birr)
Y: New clients (in
4 5 3 5.5 4 10 5 3 1 Sum
100,000)
𝑋𝑖 − 𝑋̄ -1.11 -0.61 -1.36 0.39 -0.11 6.39 -0.11 -1.11 -2.41
𝑌𝑖 − 𝑌̄ -0.50 0.50 -1.50 1.00 -0.50 5.50 0.50 -1.50 -3.50
(𝑋𝑖 − 𝑋̄) ∗ ( 𝑌𝑖 − 𝑌̄) 0.55 -0.30 2.03 0.39 0.05 35.17 -0.05 1.66 8.42 47.93
(𝑋𝑖 − 𝑋̄)2 1.22 0.37 1.84 0.16 0.01 40.89 0.01 1.22 5.79 51.50
Note that the relationship displayed in Figure 2.7 indicates a positive linear relation between advertisement
expenses and new clients.
2.5+3+2.25+4+3.5+10+3.5+2.5+1.2
𝑋̄ = = 3.606
9
4 + 5 + 3 + 5.5 + 4 + 10 + 5 + 3 + 1
𝑌̄ = = 4.5
9
Next, we need to find (𝑋𝑖 − 𝑋̄)(𝑌𝑖 − 𝑌̄) and (𝑋𝑖 − 𝑋̄)2
for each i. These are displayed in the table above. We
find that the sum of (𝑋𝑖 − 𝑋̄)2 is 51.5, and the sum of
(𝑋𝑖 − 𝑋̄)(𝑌𝑖 − 𝑌̄) is 47.93. This, we use to calculate 𝛽̂ .
𝛴(𝑋𝑖 − 𝑋̄)(𝑌𝑖 − 𝑌̄) 47.93
𝛽̂ = = = 0.9305
𝛴(𝑋𝑖 − 𝑋̄)2 51.5
That means that, if advertisement expenses increase by Figure 2.7: Scatterplot displaying the relation between advertisement
expenditure and new clients.
100,000 Birr, Safaricom gets, on average, 93,050 extra
clients. Next, we can calculate 𝛼̂
𝛼̂ = 𝑌̅ − 𝛽̂ 𝑋̅ = 4.5 − 0.9305 ∗ 3.606 = 1.1449
Summarizing, the relation between advertisement expenses (X) and extra clients (Y) for Safaricom can be
expressed by the following equation:
𝑌𝑖 = 1.1449 + 0.9305𝑋𝑖 + 𝑢𝑖
Usually, datafiles consist of many more observations and OLS is rarely calculated by hand. Most econometric
analyses are conducted by software, like Stata. An example of Stata output for this example is displayed in Figure
2.8. Note that the Coef. shows the parameter estimates for the model. 𝛼̂ is the _cons value in the output, and 𝛽̂ is
Page 10
CHAPTER TWO: SIMPLE LINEAR REGRESSION 2022
the coefficient for the variable expenditure. The other information in the output will be discussed in the sections
that follow.
Figure 2.8: Parameter estimates from Stata output. The Coef. of _cons is 𝛼̂ and the Coef. of expenditure is 𝛽̂
̂ is the change in Y (in the units in which Y is measured) by a unit change of X (in the units in which
Therefore, 𝜷
̂ 𝒊 = 𝟏. 𝟓𝟑 +
X is measured). For example, in our consumption example above the fitted model is 𝑪𝒐𝒏𝒔
̂ = 𝟎. 𝟔𝟎𝟕, can be interpreted as: if income increases by 1
𝟎. 𝟔𝟎𝟕𝒊𝒏𝒄𝒊 , which is a linear model. Therefore, 𝜷
thousands Birr, consumption will increase by 0.607 thousands Birr.
The linearity of this model implies that a one-unit change in X always has the same effect on Y, regardless of the
value of X considered.
2.3.2 Log-Linear Model (or Exponential Model)
Page 11
CHAPTER TWO: SIMPLE LINEAR REGRESSION 2022
Some functional relationships are exponential which take the form = 𝒂𝑿 , where a (a constant) is the “base”, X is
the exponent and Y is “growing exponentially” so long as 𝑿 > 𝟎. The most common base for exponential
functions is the constant e (where, 𝒆 = 𝟐. 𝟕𝟏𝟖𝟑).
Suppose the underlying relationship between two variables (X and Y) is given as: 𝒀𝒊 = 𝒆(𝜶+𝜷𝑿𝒊 +𝑼𝒊). By taking
natural logs on both sides of the above model, we obtain the following log-linear model:
𝒍𝒏(𝒀𝒊 ) = 𝜶 + 𝜷𝑿𝒊 + 𝑼𝒊 (2.14)
The corresponding sample regression function to (𝒄) be the following:
̂𝒊 ) = 𝜶
𝒍𝒏(𝒀 ̂ 𝑿𝒊
̂+𝜷 (2.15)
The slope coefficient in this model measures relative change in Y for a given absolute change in X.
𝐑𝐞𝐥𝐚𝐭𝐢𝐯𝐞 𝐜𝐡𝐚𝐧𝐠𝐞 𝐢𝐧 𝐭𝐡𝐞 𝐝𝐞𝐩𝐞𝐧𝐝𝐞𝐧𝐭 𝐯𝐚𝐫𝐢𝐚𝐛𝐥𝐞 ∆𝒀⁄
̂=
I.e., 𝜷 = 𝒀
𝐀𝐛𝐬𝐨𝐥𝐮𝐭𝐞 𝐜𝐡𝐚𝐧𝐠𝐞 𝐢𝐧 𝐭𝐡𝐞 𝐞𝐱𝐩𝐥𝐚𝐧𝐚𝐭𝐨𝐫𝐲 𝐚𝐫𝐢𝐚𝐛𝐥𝐞 ∆𝑿
If we multiply the relative change in Y by 100, (2.15) will give the percentage change or the growth rate in Y for
̂ gives the growth rate in Y.
an absolute change in the explanatory variable, i.e., 100 times 𝜷
In other words, taking first order differences in (2.15), and then multiplying both sides by 100%, we obtain
̂ 𝒊 )% = 𝜶
𝟏𝟎𝟎∆𝒍𝒏(𝒀 ̂ %∆𝑿𝒊
̂ + 𝟏𝟎𝟎𝜷
̂ %.
̂ will increase by 𝟏𝟎𝟎 ∗ 𝜷
Therefore, if 𝑿 increases by 1 unit, then 𝒀
For example, suppose an econometrician is interested to study the effect of years of education on hourly wage.
He expects that each year of education increases wage by a constant percentage. Therefore, based on a sample
of 100 individuals from a given city, the following model is estimated to explain wages:
̂ 𝒊 ) = 𝟎. 𝟕𝟓 + 𝟎. 𝟏𝟐𝟓𝑬𝑫𝑼𝑪𝒊
𝒍𝒏(𝑾𝑨𝑮𝑬
Where, EDUC (education) is measured in years of schooling and Wage is hourly wage in ETB. Note that the
coefficient on educ will have a percentage interpretation when it is multiplied by 100. For every additional year
of education, wage increases by 12.5%, on average.
Page 12
CHAPTER TWO: SIMPLE LINEAR REGRESSION 2022
𝐀𝐛𝐬𝐨𝐥𝐮𝐭𝐞 𝐜𝐡𝐚𝐧𝐠𝐞 𝐢𝐧 𝐘 ∆𝒀
̂=
I.e., 𝜷 = ∆𝑿 or ̂ ∆𝑿⁄
∆𝒀 = 𝜷
𝐑𝐞𝐥𝐚𝐭𝐢𝐯𝐞 𝐜𝐡𝐚𝐧𝐠𝐞 𝐢𝐧 𝐗 ⁄𝑿 𝑿
Therefore, taking first order differences in (2.17) and then multiplying and dividing the right hand side by 100,
̂
̂ 𝒊 = 𝜷 𝟏𝟎𝟎%∆𝒍𝒏𝑿𝒊. Therefore, if X increases by 1%, then 𝑌̂ will increase by (𝛽̂ /100) units.
we have ∆𝒀 𝟏𝟎𝟎
For example, suppose an estimated model of expenditure on Dairy products (in ETB) as a function of income
(in ETB) is given as:
𝑫𝒂𝒊𝒓𝒚 = −𝟏𝟐 + 𝟕. 𝟓𝒍𝒏(𝒊𝒏𝒄)
If the consumer’s income increases by 1%, on average, the demand of dairy products will increase by 0.075 ETB.
2.3.4 Log-Log Model (Double Log Model)
Sometimes, potential models are postulated in economic theory, such as the well-known Cobb-Douglas
functional form. This form is very popular for estimating production and demand functions. A potential model
with a unique explanatory variable is given by:
𝒀𝒊 = 𝒆𝜶 𝑿𝜷 𝒆𝑼
This model is not linear in the parameters, but it is linearizable by taking natural logarithms, and the following
is obtained:
𝒍𝒏(𝒀𝒊 ) = 𝜶 + 𝜷𝒍𝒏𝑿𝒊 + 𝑼𝒊 (2.18)
The corresponding fitted model to (2.18) is the following:
̂𝒊 ) = 𝜶
𝒍𝒏(𝒀 ̂ 𝒍𝒏𝑿𝒊
̂+𝜷 (2.19)
Taking first order differences in (2.19), we obtain
̂ 𝒊) = 𝜶
∆𝒍𝒏(𝒀 ̂ ∆𝒍𝒏𝑿𝒊
̂+𝜷 (2.20)
The slope coefficient in this model measures relative change in Y for a given relative change in X.
𝐑𝐞𝐥𝐚𝐭𝐢𝐯𝐞 𝐜𝐡𝐚𝐧𝐠𝐞 𝐢𝐧 𝐭𝐡𝐞 𝐝𝐞𝐩𝐞𝐧𝐝𝐞𝐧𝐭 𝐯𝐚𝐫𝐢𝐚𝐛𝐥𝐞 ∆𝒀⁄
̂=
I.e., 𝜷 = 𝒀
𝐑𝐞𝐥𝐚𝐭𝐢𝐯𝐞 𝐜𝐡𝐚𝐧𝐠𝐞 𝐢𝐧 𝐭𝐡𝐞 𝐞𝐱𝐩𝐥𝐚𝐧𝐚𝐭𝐨𝐫𝐲 𝐚𝐫𝐢𝐚𝐛𝐥𝐞 ∆𝑿⁄
𝑿
Therefore, multiplying both sides of (2.20) by 100%, we obtain percentage relationships
̂ 𝒊 )𝟏𝟎𝟎% = 𝜶
∆𝒍𝒏(𝒀 ̂ ∆𝒍𝒏𝑿𝒊 (𝟏𝟎𝟎%)
̂+𝜷
In case you remember the term elasticity: in this case 𝛽̂ represents elasticity of Y for X. If X increases by 1%,
then 𝑌̂ will increase by 𝛽̂ %. It is important to remark that, in this model, 𝛽̂ is the estimated elasticity of Y with
respect to X, for any value of X and Y. Consequently, in this model the elasticity is constant.
For example, suppose that to examine the effect of coffee price on quantity demanded of coffee, an investigator
estimated the following Log-Log Model:
̂ 𝒍𝒏(𝑪𝒐𝒇𝒇𝒑𝒓𝒊𝒄𝒆) + 𝑼𝒊
𝒍𝒏(𝑸𝑫𝑪𝒐𝒇𝒇) = 𝜶 + 𝜷
Page 13
CHAPTER TWO: SIMPLE LINEAR REGRESSION 2022
Page 14
CHAPTER TWO: SIMPLE LINEAR REGRESSION 2022
To calculate the total sum of squares, first the mean of Y should be calculated, which is 6. The total sum of
squares (TSS) is ∑(𝑌𝑖 − 𝑌̅)2, which is 34 (see the fourth column the table). That is the total variation for which
explanation is needed. When using the model 𝑌𝑖 = 𝛼̂ + 𝛽̂ 𝑋𝑖 + 𝑒𝑖 , with Y being ROA and X being initial capital,
you will find that the OLS method leads to 𝛼̂=-6.6 and 𝛽̂ =1.8. This helps us to find 𝑌̂𝑖 for each i. In this way, we
can calculate the model sum of squared (MSS), which is ∑(𝑌̂𝑖 − 𝑌̅)2 , is 32.4 (see the 7th column of the table).
This MSS states that from all the variation (TSS = 34), 32.4 could be explained by the linear regression model.
Lastly, the residual sum of squared (RSS) can be calculated by squaring the error terms for each observation,
which is 1.6 (see the last column of the table). Note that TSS=MSS+RSS, because 34=32.4+1.6. We decomposed
the total variation of 34 into two elements: the part explained by the model (32.4) and the part not explained by
the model (1.6).
This decomposition of variation can be used to make statements about the goodness of fit. The goodness of fit of
a model is measured by a statistical index called coefficient of determination (𝐑𝟐 ). The Coefficient of
Determination (R2 ) is the proportion of total variation of (Y) which is explained by the variation of the explanatory
variable (X) included in the model.
𝑬𝒙𝒑𝒍𝒂𝒊𝒏𝒆𝒅 𝑽𝒂𝒓𝒊𝒂𝒕𝒊𝒐𝒏 𝒊𝒏 𝒀
𝑪𝒐𝒆𝒇𝒇𝒊𝒄𝒊𝒆𝒏𝒕 𝒐𝒇 𝑫𝒆𝒕𝒆𝒓𝒎𝒊𝒏𝒂𝒕𝒊𝒐𝒏 ( 𝑹𝟐 ) =
𝑻𝒐𝒕𝒂𝒍 𝑽𝒂𝒓𝒊𝒂𝒕𝒊𝒐𝒏 𝒊𝒏 𝒀
𝑴𝒐𝒅𝒆𝒍 𝑺𝒖𝒎 𝒐𝒇 𝑺𝒒𝒖𝒂𝒓𝒆𝒔 𝑴𝑺𝑺 ∑(𝒀 ̅ )𝟐
̂ 𝒊 −𝒀
= = ∑(𝒀𝒊 −𝒀̅ )𝟐
(2.24)
𝑻𝒐𝒕𝒂𝒍 𝑺𝒖𝒎 𝒐𝒇 𝑺𝒒𝒖𝒂𝒓𝒆𝒔 𝑻𝑺𝑺
The notion of this index is straightforward in a sense that a model will be good fit if the explanatory variable (X)
included in the model determines large part of the actual values of Y. But if X is irrelevant variable, then model
would predict no part of the actual values of X.
Note that the total variation in the observed Y values about their mean value can be partitioned into two parts, one
attributable to the regression line and the other to random forces because not all actual Y observations lie on the
fitted line. Geometrically, this breakdown is displayed in Figure 2.9.
Page 15
CHAPTER TWO: SIMPLE LINEAR REGRESSION 2022
𝑹𝟐 measures part of total variation in Y which is explained by the model. As a result, it is used as an indicator to
measure the explanatory power or goodness of fit of a model. It shows the proportion of total variation in Y which
is attributable for the variation in X.
For example, in the above example where the ROA of mini-enterprises was predicted based on their start capital,
32.4 34−1.6
the TSS was 34, the MSS was 32.4 and the RSS was 1.6. Therefore, the R 2 for this case is 𝑅 2 = = =
34 34
1.6
1 − 34 ≈ 0.95.
How can this measure of the goodness of fit be interpreted? An 𝑅 2 of 0.95 means that 95% of the variation in Y
(the variation in ROA) can be explained by the model. That means that approximately 5% of the variation in Y is
not explained by the model. Equivalently, it would mean that, 90% of the total variation in the values of Y is
explained or determined (caused) by the variation in the values of X (the start up capital size). Therefore, the
model has a good fit.
Note that the maximum value for R2 is 1, which would indicate that 100% of the variation in Y is explained by
the variation in X. This would be a deterministic non-stochastic model, without any residual term. The lowest
value for R2 is 0. If 0% of the variation in Y is explained by the variation in X it means that the model there is no
difference between the mean value of Y (𝑌̅ ) and the predicted value of Y (𝑌̂𝑖 ) for each i.
Page 16
CHAPTER TWO: SIMPLE LINEAR REGRESSION 2022
̂ and 𝜶
𝜷 ̂
Figure 2.10: Stata output of the SLR between ROA and Initial_Capital
Figure 2.10 displays the calculated parameter estimates and decomposition of the sums of square (SS). Note that,
also, the R2 is included in the output.
If the assumptions as described in section 2.2 are met, OLS estimates are known from being very good. The ideal
of optimum properties that the OLS estimates possess may be summarized by the well-known theorem known as
the Gauss-Markov Theorem, named after Carl Friedrich Gauss (a German Mathematician) and Andrey Markov
(a well-known Russian Mathematician).
The Gauss-Markov Theorem can be stated as: “Given the assumptions of the classical linear regression model,
the OLS estimators, in the class of linear and unbiased estimators, have minimum variance. In other words,
the OLS estimators are BLUE.”
econometric analysis because it is useful to deduce the probability distribution of an estimate from the known
probability distribution of the dependent variable of a model.
b. Unbiased: its average or expected value is equal to the true population parameter.
According to this criterion a good estimator is one that produces an unbiased estimate. An estimate is said to be
unbiased if its bias is zero. The bias of an estimate is defined by the difference between the expected value of the
estimate and the value of the population parameter.
That is,
̂) − 𝜷
𝑩𝒊𝒂𝒔 𝒐𝒇 𝒂𝒏 𝑬𝒔𝒕𝒊𝒎𝒂𝒕𝒆, 𝜷 = 𝑬(𝜷
̂ ) is unbiased if
Thus, an estimate 𝑜𝑓(𝜷
̂) = 𝜷
𝑬(𝜷 (2.26)
Unbiasedness is desirable but not sufficient alone. Unbiasedness is good to be attained along with minimum
variance. This is so because even an unbiased estimate could be far from the true population parameter unless it
has minimum variance. Minimum variance means that the estimate has a minimum variance in the class of linear
and unbiased estimators.
An estimator is best if it has the smallest variance as compared to any other estimators obtained from other
̂ is best if:
econometric method. That is, an estimator 𝜷
̂ ) < 𝑽𝒂𝒓(𝜷
𝑽𝒂𝒓(𝜷 ̃)
Page 18
CHAPTER TWO: SIMPLE LINEAR REGRESSION 2022
Efficiency is a desirable statistical property because from two unbiased estimators of the same population
parameter, we prefer the one that has the smaller variance, i.e., the one that is statistically more precise.
Let β̂ and β̃ be two unbiased estimators of the population parameter 𝛃, such that 𝐄(𝛃
̂ ) = 𝛃 and 𝐄(𝛃
̃) = 𝛃.
̂ is efficient relative to the estimator 𝛃
Then the estimator 𝛃 ̃ if the variance of the finite-sample distribution of
̂ is less than the variance of the finite-sample distribution of 𝛃
𝛃 ̃;
Reading assignment: read, as background material, the mathematical proof for the fact that, if the assumptions of
SLR are met, the OLS estimates are BLUE.
An interval of values constructed (estimated) to predict ranges of values for the unknown population parameters
̂ & the actual sample values of 𝛂
̂&𝛃
𝛂 and 𝛃 based on some clues (i.e., the sampling distribution of 𝛂 ̂
̂&𝛃
obtained from the sample) is known as confidence interval for the population parameters 𝛂 and 𝛃. And the
degree of certainty we may assign to the interval to contain the true value of the population parameter within its
range is known as confidence level or confidence coefficient. The confidence interval can be computed from a
Z-distribution (if the sample is large n>30) or from a t-distribution (if the sample is small, n<30).
Page 19
CHAPTER TWO: SIMPLE LINEAR REGRESSION 2022
𝒁𝟏 = −𝟏. 𝟗𝟔 𝒁𝟐 = 𝟏. 𝟗𝟔
Figure 2.12: Graphical representation of a 95% confidence interval
Step-4: Interpretation – if we are curious if X is affecting Y significantly, we have to see if could be 0. If the
95% confidence interval includes 0, it is possible that is 0, indicating that there is no significant effect of X on
Y. However, if the confidence interval does not include 0, it is likely that X affects Y.
Page 20
CHAPTER TWO: SIMPLE LINEAR REGRESSION 2022
For example, let’s construct a 95% confidence interval for the population slope parameter 𝜷, for a supply
function.
𝑸 = 𝟏𝟎𝟎 + 𝟒𝑷
𝑺𝑬: (𝟏𝟎) (𝟏. 𝟓)
Where, the values in the bracket are standard errors and 𝒏 = 𝟕𝟎𝟎 producers
The 95% confidence interval can be calculated by 𝛽̂ ± 1.96. 𝑆𝐸(𝛽̂ ). The lower limit is 𝛽̂ − 1.96 ∗ 𝑆𝐸(𝛽̂ ) = 4 −
1.96 ∗ 1.5 = 1.06 and the upper limit is 𝛽̂ + 1.96 ∗ 𝑆𝐸(𝛽̂ ) = 4 + 1.96 ∗ 1.5 = 6.94. This interval is a random
interval. If the sampling process and confidence interval calculation will be repeated, it will include the true
population in at least 95% of all cases. Because 0 is not included in the confidence interval, we can say that
is statistically significantly different from 0. Therefore, P does significantly positively affect Q.
Stata regression output also includes 95% confidence intervals. Consider the output in Figure 2.13 where the
number of clients of Safaricom are predicted, based on their expenses on advertisement.
Page 21
CHAPTER TWO: SIMPLE LINEAR REGRESSION 2022
The confidence interval for is 0.669---1.192
The confidence interval for is 0.014---2.276
Both and are likely different from 0, because the confidence interval excludes 0.
Page 22
CHAPTER TWO: SIMPLE LINEAR REGRESSION 2022
Step - 2: State the null hypothesis (𝑯𝟎 ) and the alternative hypothesis (𝑯𝑨 ) of the test.
Table 2.4 provides an overview of the null hypothesis and the alternative hypothesis for various standard t-tests.
Table 2.4: Overview of the hypotheses for three types of tests
For one-tailed test
For two-tailed test Right-tailed Left-tailed
𝐻0 : 𝛼 = 0 𝑜𝑟 𝛽 = 0 𝐻0 : 𝛼 = 0 𝑜𝑟 𝛽 = 0 𝐻0 : 𝛼 = 0 𝑜𝑟 𝛽 = 0
𝐻𝐴 : 𝛼 ≠ 0 𝑜𝑟 𝛽 ≠ 0 𝐻𝐴 : 𝛼 > 0 𝑜𝑟 𝛽 > 0 𝐻𝐴 : 𝛼 < 0 𝑜𝑟 𝛽 < 0
Step – 4: Get tt (given the level of significance and the degrees of freedom) from the table
tt is the critical value for the t-test, which is obtained from Table 2.5. The critical value/s is/are the maximum
and/or minimum standardized value/s beyond which there is only 𝛼% (level of significance) chance for the
estimate to assume values under the null hypothesis. The critical value/s of the test is/are obtained from the t-
table. As a result, the critical values are also known as table or theoretical values. 𝑡𝑡 should be found at 𝛼⁄2 and
with n-2 df for a two tailed test. For a one tailed test, use 𝛼. For example, for a two-tailed test with a sample size
of 28 respondents and a significance level of 5%, tn-2,0.05/2=t26,0.025 =2.056. That means that, under the null
hypothesis, the test statistic of should with 95% certainty fall between -2.056 and +2.056.
Page 23
CHAPTER TWO: SIMPLE LINEAR REGRESSION 2022
Table 2.5: t-values given the degrees of freedom and the desired significance level
Page 24
CHAPTER TWO: SIMPLE LINEAR REGRESSION 2022
̂ −𝟎
𝜷 ̂
𝜷
𝒕𝒄 = ̂) = ̂) (2.30)
𝑺𝑬(𝜷 𝑺𝑬(𝜷
For example, you remember that in section 2.3 from a sample of size 𝑛 = 6, we have estimated the following
simple consumption function:
̂ 𝑖 = 1.53 + 0.607𝑖𝑛𝑐𝑖
𝐶𝑜𝑛𝑠
𝑆𝐸: (0.54) (0.056)
The values in the brackets are standard errors. Let’s test the hypothesis that income doesn’t affect consumption
expenditure using the t-test at 5% level of significance. In other words: let’s test the statistical significance of the
.
Step 1: We will do a two-tailed test, because we want if income affects consumption in general (not only if
income affects consumption positively).
Step 2: The hypothesis we want to test: 𝑯𝟎 : 𝜷 = 𝟎 against 𝑯𝑨 : 𝜷 ≠ 𝟎.
Step 3: The level of significance is 5%, because that is given above.
Step 4: The value from the table, tt, should be the value with 5%/2 significance and df of 6-2. t4,0.025=2.776
Step 5: The computed test statistic is
𝛽̂ − 0 𝛽̂ 0.607
𝑡𝛽̂ = = = ≅ 𝟏𝟎. 𝟖𝟒
𝑆𝐸(𝛽̂ ) 𝑆𝐸(𝛽̂ ) 0.056
̂ is statistically significant. It means is not zero, and
Step 6: t c = 10.84 > t t = 2.776 so that means that 𝛃
therefore we can conclude that income has a significant effect on consumption.
Page 25
CHAPTER TWO: SIMPLE LINEAR REGRESSION 2022
small p-values are evidence against the null hypothesis. That means, it’s possible to reject the null hypothesis if
the p-values are smaller relative to the level of significance.
In other words, the p-value is the lowest level of significance at which the observed value of a test statistic is
significant (i.e., one rejects 𝐻0 ). The p-value (or probability value) is the probability of getting a sample statistic
(such as 𝛼̂ 𝑜𝑟 𝛽̂ ) when the null hypothesis is true. The p-value for obtaining a sample outcome is compared to the
level of significance (𝛼). By checking if the p-value is less than the intended level of significance, e.g. 5%, an
econometrician can immediately conclude about the parameter’s significance, without following all the steps.
Especially if software is used to analyze the data it is common to immediately look at the p-value rather than
following all steps. See Figure 2.14.
Figure 2.14: Stata output displaying information relevant for conducting the t-test
Instead of reporting p-values by value, another common notation in published research work is the asterisks (*).
One * indicates significance at a level of 10% (that means the p-value is on the range 5%-10%, or 0.05-0.01), two
asterisks (**) indicates significance at a level of 5% (that means the p-value is on the range 1%-5%, or 0.01-0.05).
Lastly, three asterisks (***) indicates very strong significance, because the p-value is less than 1% or less than
0.01.
̂=𝜶
𝒀 ̂ 𝑿𝒑
̂+𝜷 (2.31)
Page 26
CHAPTER TWO: SIMPLE LINEAR REGRESSION 2022
Figure 2.15: Scatterplot showing the relation between sales (Y) and price (X) of second hand sport shoes.
SLR output of Stata is displayed in Figure 2.16 and reveals that 𝛼̂ = 243.2 and 𝛽̂ = −0.2198. Both
coefficients, and are significantly different from 0 (because p<0.05 for both coefficients).
Figure 2.16: SLR output for the relation between sales (Y) and prices of second-hand sport shoes (X).
Page 27
CHAPTER TWO: SIMPLE LINEAR REGRESSION 2022
This model can be used to forecast the sales of second-hand sport shoes, if the price is known. For example, if the
price of second-hand sport shoes is 500 Birr, the sales are expected to be:
𝑌̂𝑖 = 243.2 − 0.2198𝑋𝑖 = 243.2 − 0.2198 ∗ 500 = 133 pairs of shoes.
Another example, suppose you have an estimated model of sales (Y) of firms producing a particular product as
a function of advertisement expenditure (X):
̂ = 𝟑𝟎𝟎 + 𝟐. 𝟓𝑿𝟏
𝒀
(p:0.21) (p:0.003)
In addition, you have predicted value for the independent variable (advertisement expenditure in ETB), 𝐗 𝟏 =
𝟐𝟎𝟎. Then the predicted value for Sales is 800 Units. The fact that the intercept is insignificant and the slope is
significant does not affect if they should be included: all parameter estimates should be included in the predicted
value.
̂ 𝐢 = 𝟏. 𝟓𝟑 + 𝟎. 𝟔𝟎𝟕𝐢𝐧𝐜𝐢 . Based on this estimated
Last example: Recall our estimated consumption Model: 𝐂𝐨𝐧𝐬
model predict the consumption expenditure of a household whose income is 12 ETB?
̂ 𝐢 = 𝟏. 𝟓𝟑 + 𝟎. 𝟔𝟎𝟕(𝟏𝟐) = 𝟖. 𝟖𝟏𝟒. That means, a household with an income of 12 ETB will spend
Solution: 𝐂𝐨𝐧𝐬
8.814 ETB of its income for consumption.
Page 28