0% found this document useful (0 votes)
2 views28 pages

Chapter 2

Chapter Two discusses Simple Linear Regression, focusing on the relationship between a dependent variable and an independent variable. It introduces the concept of regression, its assumptions, and the mathematical formulation of the regression model, emphasizing the importance of identifying causal relationships in econometric analysis. The chapter also outlines key assumptions for valid regression analysis, such as linearity, correct model specification, and the independence of error terms.

Uploaded by

yabi1072
Copyright
© All Rights Reserved
We take content rights seriously. If you suspect this is your content, claim it here.
Available Formats
Download as PDF, TXT or read online on Scribd
0% found this document useful (0 votes)
2 views28 pages

Chapter 2

Chapter Two discusses Simple Linear Regression, focusing on the relationship between a dependent variable and an independent variable. It introduces the concept of regression, its assumptions, and the mathematical formulation of the regression model, emphasizing the importance of identifying causal relationships in econometric analysis. The chapter also outlines key assumptions for valid regression analysis, such as linearity, correct model specification, and the independence of error terms.

Uploaded by

yabi1072
Copyright
© All Rights Reserved
We take content rights seriously. If you suspect this is your content, claim it here.
Available Formats
Download as PDF, TXT or read online on Scribd

CHAPTER TWO: SIMPLE LINEAR REGRESSION

CHAPTER TWO
SIMPLE LINEAR REGRESSION
As you know, financial and theories are mainly concerned with the relationships between variables. This chapter
introduces the simplest possible regression analysis, namely, the two variables regression in which the dependent
variable is linearly related to one independent explanatory variable. This case is considered here, because it
presents the fundamental idea of regression analysis as simply as possible. We shall first discuss the basic concept
of regression, and then proceed to the core of simple linear regression analysis.

2.1 Concept of regression & assumptions


The main goal of any econometric analysis is to identify empirically if there is a causal relationship between
economic variables. The term regression was introduced by Francis Galton. In his famous paper “Family
Likeness in Stature”, Galton found that, although there was a tendency for tall parents to have tall children and
for short parents to have short children, the average height of children born of parents of a given height tended to
move or “regress” toward the average height in the population.

Regression analysis is concerned with the study of the dependence of one variable (the dependent variable) on
one or more other variables (the explanatory variable(s)). In other words, regression analysis is concerned with
describing and evaluating the relationship between a given variable (the dependent variable) and one or
more other variables (the independent variable(s)). The objective of regression analysis is to estimate and/or
predict the unknown (population) mean value of the dependent variable given known values of the explanatory
variables.

In Galton’s example, this is to estimate/predict the mean height of children (dependent variable), given the height
of their parents (explanatory variables). Galton’s example is not related to finance or economics, but the following
example is. A financial expert may be interested in studying the dependence of household monthly savings on
household monthly disposable income. That is, we want to predict average savings, knowing household monthly
disposable income. Such an analysis is helpful in estimating the marginal propensity to save (MPS), that is,
average change in saving for, say, a unit change in disposable income. To see how this can be done, consider
Figure 2.1.

As discussed in section 1.4, regression analysis deals with statistical dependence among variables, but not with
deterministic dependence among variables. In statistical relationships, we essentially deal with random
(stochastic) variables, i.e., variables that have probability distributions.

Page 1
CHAPTER TWO: SIMPLE LINEAR REGRESSION 2022

Figure 2.1 shows the distribution of household monthly savings in a hypothetical population corresponding to the
given or fixed values of the household monthly disposable income. Notice that corresponding to any given
household income is a range or distribution of the household savings. However, notice that despite the variability
of savings for a given value of household income, the average savings generally increases as the income of
household increases.

The line that passes through the average level of consumption expenditure for each level of household income is
known as the regression line. It shows how the average consumption expenditure increases with the household’s
income.

Regression Line

Figure 2.1: Scatter plot diagram with savings as dependent variable and disposable income as explanatory variable

In econometrics we exclusively deal with stochastic relationships. Although regression analysis deals with the
dependence of one variable on other variables, it does not necessarily imply causation. For example, consider
the regression line in Figure 2.2 which displays the relationship between hot drink sales and number of customers
of a resort. If it is busy in the resort (number of customers approaches 100-150, see the x-axis), the tea sales are
low. But if the number of customers is low, the tea sales are high. Does this mean that the number of customers
negatively affects tea sales? That would not make sense from a theoretical perspective. In this case, there is a third
variable, the weather, which affects both the tea sales and the number of customers. If the weather is cold, less
customers will come to the resort and those who come will likely order tea. If the weather is hot, more customers
will come to the resort, and they will likely order cold drinks (not tea).

Page 2
50
40
30
20
10 CHAPTER TWO: SIMPLE LINEAR REGRESSION 2022
0

0 50 100 150
number of customers

tea sales Fitted values

Figure 2.2: The scatter plot diagram with tea sales as dependent variable and number of customers as explanatory variable

Therefore, the determination of the direction of causation should come from outside of statistics, i.e., economic
or financial theory. In other words, statistical relationships by themselves cannot logically imply causation. To
ascribe causality, one must appeal to ‘a priori’ or theoretical considerations.

The variables in a regression relation consist of dependent and explanatory variables. The dependent variable is
the variable whose variation is being explained by the other variable(s). The explanatory variable is the variable
whose variation is used to explain the variation in the dependent variable. Other words for the dependent variable
are explained variable, or regressand, or, simply the letter Y. Other words for the explanatory variable are
independent variable, regressor, or, simply the letter X.

The population equation for Simple Linear Regression is


𝒀𝒊 = 𝜶 + 𝜷𝑿𝒊 + 𝑼𝒊 (2.1)
Where 𝑌𝑖 is the dependent variable value of observation i, 𝛼 is the constant (the intercept), 𝛽 is the coefficient (or
the slope), 𝑋𝑖 is the independent variable value of observation i and 𝑈𝑖 is the error term of observation i. Note that
(2.1) consists of three parts:
𝒀⏟𝒊 = 𝜶 + 𝜷𝑿𝒊
⏟ + 𝑼
⏟𝒊
𝒅𝒆𝒑𝒆𝒏𝒅𝒆𝒏𝒕 𝒗𝒂𝒓 𝒊𝒂𝒃𝒍𝒆 𝒓𝒆𝒈𝒓𝒆𝒔𝒔𝒊𝒐𝒏 𝒍𝒊𝒏𝒆 𝒓𝒂𝒏𝒅𝒐𝒎 𝒗𝒂𝒓 𝒊𝒂𝒃𝒍𝒆

 and  are true population parameters, but research rarely has sufficient resources to check values for the whole
population. Therefore, sample data is used to estimate/calculate 𝛼̂ (which is the best fitting  given the sample
data) and 𝛽̂ (which is the best fitting  given the sample data).

Page 3
CHAPTER TWO: SIMPLE LINEAR REGRESSION 2022

Consider this example: Cloth production is upcoming in Ethiopia. Some factories manage to make good profit,
while others make very small or no profits. Hana, a financial expert, expects that some make no profit because
they do not optimally use their workers and equipment (sewing machines). Therefore, she expects that technical
efficiency determines how much profit the cloth producing factory makes. In Hana’s analysis the dependent
variable is profitability (Yi) and the explanatory variable is technical efficiency (Xi). Profitability is measured as
𝑡𝑜𝑡𝑎𝑙 𝑝𝑟𝑜𝑓𝑖𝑡
net profit margin ratio (which is ). Technical efficiency is measured as
𝑡𝑜𝑡𝑎𝑙 𝑟𝑒𝑣𝑒𝑛𝑢𝑒
𝑎𝑐𝑡𝑢𝑎𝑙 𝑝𝑟𝑜𝑑𝑢𝑐𝑡𝑖𝑜𝑛 𝑜𝑢𝑡𝑝𝑢𝑡
. Note that both X and Y are ratios, which Hana collected
𝑚𝑎𝑥 𝑝𝑜𝑠𝑠𝑖𝑏𝑙𝑒 𝑝𝑟𝑜𝑑𝑢𝑐𝑡𝑖𝑜𝑛 𝑜𝑢𝑡𝑝𝑢𝑡 𝑔𝑖𝑣𝑒𝑛 𝑒𝑞𝑢𝑖𝑝𝑚𝑒𝑛𝑡 𝑎𝑛𝑑 𝑙𝑎𝑏𝑜𝑟

from 35 industrial cloth producers in Ethiopia (n=35).


The Simple Linear Regression output reveals the optimal parameter estimates, given the sample data, are 𝛼̂=-
0.031 and 𝛽̂ =0.19. That means that (fictional) factories with a technical efficiency of 0 are expected to have a
profitability of -.031, which is -3 percent. For each percent technical efficiency added, the profitability is expected
to increase by 0.22 percent. Factories with a technical efficiency of 50% (0.5) are expected to have a profit of
-0.031+0.5*0.19=0.064 (6.4 percent) and factories with a maximum technical efficiency (1, or 100%), are
expected to have a profitability of -0.031+1*0.19=0.159 (15.9%).
Of course, this is not a secure and mathematic relationship, because cloth factories deviate individually from this
prediction by 𝑈𝑖 , which is a stochastic term that includes determinants of profitability other than the technical
efficiency and other random issues (remember the discussion about econometric models in section 1.4). For a
regression analysis like this to be valid, there are nine assumptions about the data and the model parameters.

Assumption 1: The model is linear in parameters


Simple linear regression can only be used if the relationship between X and Y is linear, because 𝛽 is a constant
parameter, which should be relevant in predicting the change in Y for changes in all possible values of X. As an
example, let’s assume that Hana researches the effect of technical efficiency (X) on profitability of cloth factories
(Y) in Ethiopia. Suppose that the relationship between technical efficiency and profitability is as displayed in
Scatterplot 1 of Figure 2.3, the SLR model is not appropriate, because there is no effect of a change in X on Y.
All types of factories are available: those with high profitability and high technical efficiency, those with high
profitability and low technical efficiency, those with low profitability and low technical efficiency and those with
low profitability and high technical efficiency. Scatterplot 2 clearly shows a trend, but this is not a linear trend.
The effect of technical efficiency (X) on profitability (Y) is negative for low values of X, there is no relation
between X and Y for medium values of X and there is a positive relation between X and Y for high values of X.
Only the relationship in Scatterplot 3 shows a linear trend, which is suitable for analysis with SLR.

Page 4
CHAPTER TWO: SIMPLE LINEAR REGRESSION 2022

Figure 2.3: Three types of relations between technical efficiency and profitability

Assumption 2: The regression model is correctly specified


This means that the mathematical form of the model is correctly specified (in line with assumption 1) and all
important explanatory variables are included in it. In other words, there is no specification bias or error in the
model used in empirical analysis. Unfortunately, in practice one rarely specifies the correct model. Hence, an
econometrician would use some judgment in choosing the correct model, i.e., in determining the type and number
of variables entering the model, and functional form of the model s/he has to utilize on the basis of theoretical
grounds.
Assumption 3: 𝑼𝒊 is a random real variable with mean value of zero
This means that the value which 𝑼 𝒊 takes depends on chance. This means that for each value of x, the random
variable (U) may assume various values, some greater than zero and some smaller than zero, but if we considered
all the positive and negative values of U, for any given value of X, they would have an average value equal to
zero. In other words, the positive and negative values of 𝑼 cancel each other.
Mathematically,
𝑬(𝑼𝒊 ) = 𝟎 (2.2)
The implication of this assumption is that the factors not explicitly included in the model and therefore subsumed
in 𝑼𝒊 , do not systematically affect the mean value of Y. i.e., the positive 𝑼𝒊 values cancel out the negative 𝑈𝑖
values so that their average effect on Y is zero.
Given 𝒀𝒊 = 𝜶 + 𝜷𝑿𝒊 + 𝑼𝒊 , this assumption leads to the fact that: 𝑬(𝒀𝒊 ) = 𝜶 + 𝜷𝑿𝒊

Assumption 4: The variance of the error term (𝑼𝒊 ) is constant across observations
Another name of this assumption is the assumption of homoscedasticity. This means that, for all values of X, the
U’s will show the same dispersion around their mean. This implies that, given the value of X, the variance of 𝑈𝑖
is the same (constant) for all observations. Figure 2.4 displays a situation in which this assumption is met. Figure
Page 5
CHAPTER TWO: SIMPLE LINEAR REGRESSION 2022

2.5 shows a situation where the spread in U increases for higher values of X. Note: this assumption implies that
the values of Y corresponding to various values of X have constant variance.

Figure 2.4: Heteroskedastic Variance--- spread of errors depends on X Figure 2.5: Homoscedastic variance - the error has constant variance

Assumption 5: The random variable (U) has a normal distribution


This means the values of U (for each x) have a bell-shaped symmetrical distribution about their zero mean and
constant variance, 𝝈𝟐 , i.e.,
𝑼𝒊 ~𝑵(𝟎, 𝝈𝟐 ) (2.3)
The reason for this assumption is that if U is normally distributed, so will Y and the estimated regression
coefficients, and this will be useful in performing tests of hypotheses and constructing confidence intervals for
𝜶 and 𝜷.

Assumption 6: The random terms of different observations (𝑼𝒊 , 𝑼𝒋 ) are independent


This means, in other words, that there is no autocorrelation or no serial correlation. If the order of observations in
the dataset is random, there is likely no autocorrelation. However, whenever some ordering of sampling units is
present, autocorrelation may arise. In the cross-sectional data, the neighboring units tend to be similar with respect
to the characteristic under study. In time-series data, time is the factor that produces autocorrelation, because there
is a carryover effect. It means that:
𝑪𝒐𝒗(𝑼𝒊 , 𝑼𝒋 ) = 𝑬(𝑼𝒊 𝑼𝒋 ) = 𝟎 (2.4)
Assumption 7: Exogeneity: X is uncorrelated with the error term.
The explanatory variable is determined outside the model (exogenous). This means that there is no correlation
between the random variable and the explanatory variable.

Page 6
CHAPTER TWO: SIMPLE LINEAR REGRESSION 2022

What happens if this assumption is violated?


Suppose we have the model: 𝒀𝒊 = 𝜶 + 𝜷𝑿𝒊 + 𝑼𝒊 . Suppose 𝑋𝑖 and 𝑈𝑖 are positively correlated, i.e., when 𝑋𝑖 is
large, 𝑈𝑖 tends to be large as well.
Why would 𝑋𝑖 and 𝑈𝑖 be correlated? Suppose you are trying to study the relationship between the price of Burger
and the quantity sold across various restaurants of a given city. Suppose you estimate the following model: 𝒀𝒊 =
𝜶 + 𝜷 𝑷𝒓𝒊𝒄𝒆𝒊 + 𝑼𝒊 . The problem in this model, however, is that quality of Burger differs across restaurants.
Thus, quality should be included as explanatory variable, but you fail to do so it becomes part of 𝑼𝒊 . However,
price and quality are highly positively correlated. Therefore, 𝑋𝑖 and 𝑈𝑖 are positively correlated. This means that
𝛽̂1 will be too high. This is called Omitted Variables Bias. The consequence of omitting quality from this model
is visible in Figure 2.6.

Figure 2.6: The effect of omitting a relevant variable. Note that the estimated line over-estimates X's true impact on Y

Thus, to have unbiased estimator, the exogeneity assumption is important. If two variables are unrelated their
covariance shall be zero.
𝑪𝒐𝒗(𝑿𝒊 𝑼𝒊 ) = 𝟎 (𝟐. 𝟓)
Assumption 8: Variability in X values
The x values in a given sample must not all be the same. This means that x assumes different values in a given
sample; but it assumes fixed values in hypothetically repeated samples. This assumption is very critical since
without this assumption it would be impossible to estimate the parameters and hence, regression analysis would
fail. For example, if there is little variation in household income, we will not be able to explain much of the
variation in the consumption expenditure of the households.
Assumption 9: The distribution of the dependent variable, Y
Based on the assumptions we discussed so far about the distributions of 𝑿 and 𝑼, we can establish that Y is
normally distributed with Mean: 𝑬(𝒀𝒊 ) = 𝜶 + 𝜷𝑿𝒊 , and variance: 𝑽𝒂𝒓(𝒀𝒊 |𝑿𝒊 ) = 𝑽𝒂𝒓(𝑼𝒊 ) = 𝝈𝟐 .

Page 7
CHAPTER TWO: SIMPLE LINEAR REGRESSION 2022

2.2 Methods of Estimation: OLS, MoM and MLM


Recall our simple Regression Model: 𝑌𝑖 = 𝛼 + 𝛽𝑋𝑖 + 𝑈𝑖 with 𝑌𝑖 = household savings and 𝑋𝑖 = household
disposable income. Given the sample data, how do we find the optimal values for the parameters ( and )? The
parameters of the simple linear regression model can be estimated by various methods. Three of the most
commonly used methods are:
1. Ordinary Least Square Method (OLS)
2. Method of Moments (MoM)
3. Maximum Likelihood Method (MLM)
In this course we will discuss the OLS method in depth, and for the other methods you will get a reading
assignment.
2.2.1. The OLS method
This chapter will firstly show the OLS estimation method. The true population parameters ( and ) do exist, but
are unknown, because we only collect data from a sample. Therefore, we are going to estimate the population
parameters by calculating optimal sample parameters. For the model, we use the following notation:

𝒀𝒊 = 𝜶 ̂ 𝑿𝒊 + 𝒆𝒊
̂ +𝜷 (2.6)

Note that 𝛼 changes into 𝛼̂, 𝛽 changes into 𝛽̂ and 𝑈𝑖 changes into 𝑒𝑖 because we do our analysis with sample data.
OLS is the technique used to estimate a line that will minimize the error (the difference between the predicted
and the actual values of a dependent variable, i.e. the e).

From the estimated relationship, 𝑌𝑖 = 𝛼̂ + 𝛽̂ 𝑋𝑖 + 𝑒𝑖 , we obtain 𝑒𝑖 = 𝑌𝑖 − (𝛼̂ + 𝛽̂ 𝑋𝑖 ). That means that we should
find the values for 𝛼̂ and 𝛽̂ which minimize ∑ 𝑒𝑖2 = ∑(𝑌𝑖 − 𝛼̂ − 𝛽̂ 𝑋𝑖 )2 . Therefore, we must partially differentiate
∑ 𝑒𝑖2 with respect to 𝛼̂ and 𝛽̂ and set the partial derivatives equal to zero.
Partial derivative with respect to 𝛼̂:
𝝏 ∑ 𝒆𝟐𝒊
= −𝟐 ∑(𝒀𝒊 − 𝜶 ̂ 𝑿𝒊 ) = 𝟎
̂−𝜷 (2.7)
̂
𝝏𝜶

Or, rearranged, normal equation 1: ∑ 𝒀𝒊 = 𝒏𝜶 ̂ 𝜮𝑿𝒊 . Note that, if we divide both sides by n and rearrange,
̂+𝜷
normal equation 1 is
̂=𝒀
𝜶 ̂𝑿
̅−𝜷 ̅ (2.8)
Partial derivative with respect to 𝛽̂ :
𝝏 ∑ 𝒆𝟐𝒊
= −𝟐 ∑ 𝑿𝒊 (𝒀𝒊 − 𝜶 ̂ 𝑿𝒊 ) = 𝟎
̂−𝜷 (2.9)
̂
𝝏𝜷

Page 8
CHAPTER TWO: SIMPLE LINEAR REGRESSION 2022

Or, rearranged, normal equation 2:


∑ 𝒀𝒊 𝑿𝒊 = 𝜶 ̂ 𝜮𝑿𝟐𝒊
̂ 𝜮𝑿𝒊 + 𝜷 (2.10)
Substituting the 𝛼̂ based on (2.8) in (2.10), we obtain:
̂𝑿
̅−𝜷
∑ 𝒀𝒊 𝑿𝒊 = (𝒀 ̂ 𝜮𝑿𝟐𝒊 = 𝜮𝑿𝒊 (𝒀̄ − 𝜷
̅ )𝜮𝑿𝒊 + 𝜷 ̂ 𝑿̄) + 𝜷
̂ 𝜮𝑿𝟐𝒊 = 𝒀̄𝜮𝑿𝒊 − 𝜷
̂ 𝑿̄𝜮𝑿𝒊 + 𝜷
̂ 𝜮𝑿𝟐𝒊

̂ (𝜮𝑿𝟐𝒊 − 𝑿̄𝜮𝑿𝒊 )
∑ 𝒀𝒊 𝑿𝒊 − 𝒀̄𝜮𝑿𝒊 = 𝜷

̂ (𝜮𝑿𝟐𝒊 − 𝒏𝑿̄2)
𝜮𝑿𝒊 𝒀𝒊 − 𝒏𝑿̄𝒀̄ =𝜷
̂ = 𝜮𝑿𝒊 𝒀𝟐 𝒊 −𝒏𝑿̄𝒀̄
𝜷 𝜮𝑿 −𝒏𝑿̄𝟐
𝒊

or, rearranged,
̄ ̄
̂ = 𝜮(𝑿𝒊−𝑿)(𝒀𝒊−𝒀)
𝜷 (2.11)
̄ 𝟐
𝜮(𝑿𝒊 −𝑿)

After finding the value of 𝛽̂ , it is easy to find the value of 𝛼̂ by substituting 𝜷


̂ and the mean values for X and Y
in (2.8).
Example 1: Estimation of a Keynesian consumption function
Given the Keynesian consumption function, 𝑪𝒐𝒏𝒔𝒊 = 𝜶 + 𝜷𝒊𝒏𝒄𝒊 + 𝑼𝒊
where, 𝑪𝒐𝒏𝒔𝒊 , and 𝒊𝒏𝒄𝒊 −stands for consumption expenditure and monthly income of the ith
household, respectively. And 𝜶 and 𝜷 are the intercept and marginal propensity to consume,
respectively, which we are interested to compute.

Given consumption (Y) and income (X) data, both in thousands of Birr, of six households, obtain the OLS
estimators of 𝛼 and 𝛽.
42 54
𝑌̅ = = 7, and 𝑋̅ = = 9
6 6
Observations 𝑌𝑖 𝑋𝑖 𝑌𝑖 𝑋𝑖 𝑋𝑖2 𝑦𝑖 = 𝑌𝑖 − 𝑌̅ 𝑥𝑖 = 𝑋𝑖 − 𝑋̅ (𝑌𝑖 − 𝑌̅) (𝑋𝑖 − 𝑋̅)2

(𝑋𝑖 − 𝑋̅)
1. 4 5 20 25 -3 -4 12 16
2. 4 4 16 16 -3 -5 15 25
3. 7 8 56 64 0 -1 0 1
4. 8 10 80 100 1 1 1 1
5. 9 13 117 169 2 4 8 16
6. 10 14 140 196 3 5 15 25
Sums 42 54 429 570 0 0 51 84
∑(𝑋𝑖 − 𝑋̅)(𝑌𝑖 − 𝑌̅) 51
𝛽̂ = , 𝛽̂ = = 0.607
∑(𝑋𝑖 − 𝑋̅)2 84
𝛼̂ = 𝑌̅ − 𝛽̂ 𝑋̅ ∴ 𝛼̂ = 7 − 0.607(9) = 1.53
Therefore, the fitted regression equation (i.e., OLS regression line) is:

Page 9
CHAPTER TWO: SIMPLE LINEAR REGRESSION 2022

̂𝑖 = 1.53 + 0.607𝑖𝑛𝑐𝑖
𝐶𝑜𝑛𝑠
Example 2: Effect of advertisement on number of customers
The marketing manager of Safaricom, the new competitor of EthioTelecom in the Ethiopian telecommunication
market wants to know the relation between expenses on advertisement and new clients. To this end, the team
collects the following data:
Month January February April May June July August Sept Oct
X: Advertisement
2.5 3 2.25 4 3.5 10 3.5 2.5 1.2
expenses (in 100,000 Birr)
Y: New clients (in
4 5 3 5.5 4 10 5 3 1 Sum
100,000)
𝑋𝑖 − 𝑋̄ -1.11 -0.61 -1.36 0.39 -0.11 6.39 -0.11 -1.11 -2.41
𝑌𝑖 − 𝑌̄ -0.50 0.50 -1.50 1.00 -0.50 5.50 0.50 -1.50 -3.50
(𝑋𝑖 − 𝑋̄) ∗ ( 𝑌𝑖 − 𝑌̄) 0.55 -0.30 2.03 0.39 0.05 35.17 -0.05 1.66 8.42 47.93
(𝑋𝑖 − 𝑋̄)2 1.22 0.37 1.84 0.16 0.01 40.89 0.01 1.22 5.79 51.50
Note that the relationship displayed in Figure 2.7 indicates a positive linear relation between advertisement
expenses and new clients.
2.5+3+2.25+4+3.5+10+3.5+2.5+1.2
𝑋̄ = = 3.606
9

4 + 5 + 3 + 5.5 + 4 + 10 + 5 + 3 + 1
𝑌̄ = = 4.5
9
Next, we need to find (𝑋𝑖 − 𝑋̄)(𝑌𝑖 − 𝑌̄) and (𝑋𝑖 − 𝑋̄)2
for each i. These are displayed in the table above. We
find that the sum of (𝑋𝑖 − 𝑋̄)2 is 51.5, and the sum of
(𝑋𝑖 − 𝑋̄)(𝑌𝑖 − 𝑌̄) is 47.93. This, we use to calculate 𝛽̂ .
𝛴(𝑋𝑖 − 𝑋̄)(𝑌𝑖 − 𝑌̄) 47.93
𝛽̂ = = = 0.9305
𝛴(𝑋𝑖 − 𝑋̄)2 51.5
That means that, if advertisement expenses increase by Figure 2.7: Scatterplot displaying the relation between advertisement
expenditure and new clients.
100,000 Birr, Safaricom gets, on average, 93,050 extra
clients. Next, we can calculate 𝛼̂
𝛼̂ = 𝑌̅ − 𝛽̂ 𝑋̅ = 4.5 − 0.9305 ∗ 3.606 = 1.1449
Summarizing, the relation between advertisement expenses (X) and extra clients (Y) for Safaricom can be
expressed by the following equation:
𝑌𝑖 = 1.1449 + 0.9305𝑋𝑖 + 𝑢𝑖
Usually, datafiles consist of many more observations and OLS is rarely calculated by hand. Most econometric
analyses are conducted by software, like Stata. An example of Stata output for this example is displayed in Figure
2.8. Note that the Coef. shows the parameter estimates for the model. 𝛼̂ is the _cons value in the output, and 𝛽̂ is

Page 10
CHAPTER TWO: SIMPLE LINEAR REGRESSION 2022

the coefficient for the variable expenditure. The other information in the output will be discussed in the sections
that follow.

Figure 2.8: Parameter estimates from Stata output. The Coef. of _cons is 𝛼̂ and the Coef. of expenditure is 𝛽̂

2.2.2 The MoM method


For this method you will get a reading assignment.
2.2.3 The MLM
For this method you’ll get a reading assignment.

2.3 Functional Forms and Interpretation of Estimates


So far, we primarily dealt with models that are linear in parameters and in variables. In many cases, however,
linear relationships are not adequate for econometric applications. Now, we will consider some commonly used
regression models that may be non–linear in variables but are linear in the parameters or can be made so by
̂ in each case.
suitable transformations of the variables. We will look at the interpretation of the coefficient 𝜷
2.3.1 The Linear Model
The coefficient measures the effect of the regressor X on Y. Let us look at this in detail. The observation i of the
sample regression function is given by:
̂𝒊 = 𝜶
𝒀 ̂ 𝑿𝒊
̂+𝜷 (𝟐. 𝟏𝟐)
Taking the first order difference both sides of (2.12) and rearranging, we obtain:
̂𝒊
𝒅𝒀
𝒅𝒀 ̂ 𝒅𝑿𝒊 ,
̂𝒊 = 𝜷 ̂
=𝜷 (2.13)
𝒅𝑿𝒊

̂ is the change in Y (in the units in which Y is measured) by a unit change of X (in the units in which
Therefore, 𝜷
̂ 𝒊 = 𝟏. 𝟓𝟑 +
X is measured). For example, in our consumption example above the fitted model is 𝑪𝒐𝒏𝒔
̂ = 𝟎. 𝟔𝟎𝟕, can be interpreted as: if income increases by 1
𝟎. 𝟔𝟎𝟕𝒊𝒏𝒄𝒊 , which is a linear model. Therefore, 𝜷
thousands Birr, consumption will increase by 0.607 thousands Birr.
The linearity of this model implies that a one-unit change in X always has the same effect on Y, regardless of the
value of X considered.
2.3.2 Log-Linear Model (or Exponential Model)

Page 11
CHAPTER TWO: SIMPLE LINEAR REGRESSION 2022

Some functional relationships are exponential which take the form = 𝒂𝑿 , where a (a constant) is the “base”, X is
the exponent and Y is “growing exponentially” so long as 𝑿 > 𝟎. The most common base for exponential
functions is the constant e (where, 𝒆 = 𝟐. 𝟕𝟏𝟖𝟑).
Suppose the underlying relationship between two variables (X and Y) is given as: 𝒀𝒊 = 𝒆(𝜶+𝜷𝑿𝒊 +𝑼𝒊). By taking
natural logs on both sides of the above model, we obtain the following log-linear model:
𝒍𝒏(𝒀𝒊 ) = 𝜶 + 𝜷𝑿𝒊 + 𝑼𝒊 (2.14)
The corresponding sample regression function to (𝒄) be the following:
̂𝒊 ) = 𝜶
𝒍𝒏(𝒀 ̂ 𝑿𝒊
̂+𝜷 (2.15)
The slope coefficient in this model measures relative change in Y for a given absolute change in X.
𝐑𝐞𝐥𝐚𝐭𝐢𝐯𝐞 𝐜𝐡𝐚𝐧𝐠𝐞 𝐢𝐧 𝐭𝐡𝐞 𝐝𝐞𝐩𝐞𝐧𝐝𝐞𝐧𝐭 𝐯𝐚𝐫𝐢𝐚𝐛𝐥𝐞 ∆𝒀⁄
̂=
I.e., 𝜷 = 𝒀
𝐀𝐛𝐬𝐨𝐥𝐮𝐭𝐞 𝐜𝐡𝐚𝐧𝐠𝐞 𝐢𝐧 𝐭𝐡𝐞 𝐞𝐱𝐩𝐥𝐚𝐧𝐚𝐭𝐨𝐫𝐲 𝐚𝐫𝐢𝐚𝐛𝐥𝐞 ∆𝑿

If we multiply the relative change in Y by 100, (2.15) will give the percentage change or the growth rate in Y for
̂ gives the growth rate in Y.
an absolute change in the explanatory variable, i.e., 100 times 𝜷

In other words, taking first order differences in (2.15), and then multiplying both sides by 100%, we obtain
̂ 𝒊 )% = 𝜶
𝟏𝟎𝟎∆𝒍𝒏(𝒀 ̂ %∆𝑿𝒊
̂ + 𝟏𝟎𝟎𝜷
̂ %.
̂ will increase by 𝟏𝟎𝟎 ∗ 𝜷
Therefore, if 𝑿 increases by 1 unit, then 𝒀
For example, suppose an econometrician is interested to study the effect of years of education on hourly wage.
He expects that each year of education increases wage by a constant percentage. Therefore, based on a sample
of 100 individuals from a given city, the following model is estimated to explain wages:
̂ 𝒊 ) = 𝟎. 𝟕𝟓 + 𝟎. 𝟏𝟐𝟓𝑬𝑫𝑼𝑪𝒊
𝒍𝒏(𝑾𝑨𝑮𝑬
Where, EDUC (education) is measured in years of schooling and Wage is hourly wage in ETB. Note that the
coefficient on educ will have a percentage interpretation when it is multiplied by 100. For every additional year
of education, wage increases by 12.5%, on average.

2.3.3 Linear-Log Model


A linear-Log model is a function where the dependent variable is defined as a linear function of the logarithm of
explanatory variable as given below:
𝒀𝒊 = 𝜶 + 𝜷𝒍𝒏𝑿𝒊 + 𝑼𝒊 (2.16)
The corresponding fitted function is the following:
̂𝒊 = 𝜶
𝒀 ̂ 𝒍𝒏𝑿𝒊
̂ +𝜷 (2.17)
The slope coefficient in this model measures absolute change in Y for relative change in X.

Page 12
CHAPTER TWO: SIMPLE LINEAR REGRESSION 2022

𝐀𝐛𝐬𝐨𝐥𝐮𝐭𝐞 𝐜𝐡𝐚𝐧𝐠𝐞 𝐢𝐧 𝐘 ∆𝒀
̂=
I.e., 𝜷 = ∆𝑿 or ̂ ∆𝑿⁄
∆𝒀 = 𝜷
𝐑𝐞𝐥𝐚𝐭𝐢𝐯𝐞 𝐜𝐡𝐚𝐧𝐠𝐞 𝐢𝐧 𝐗 ⁄𝑿 𝑿
Therefore, taking first order differences in (2.17) and then multiplying and dividing the right hand side by 100,
̂
̂ 𝒊 = 𝜷 𝟏𝟎𝟎%∆𝒍𝒏𝑿𝒊. Therefore, if X increases by 1%, then 𝑌̂ will increase by (𝛽̂ /100) units.
we have ∆𝒀 𝟏𝟎𝟎

For example, suppose an estimated model of expenditure on Dairy products (in ETB) as a function of income
(in ETB) is given as:
𝑫𝒂𝒊𝒓𝒚 = −𝟏𝟐 + 𝟕. 𝟓𝒍𝒏(𝒊𝒏𝒄)
If the consumer’s income increases by 1%, on average, the demand of dairy products will increase by 0.075 ETB.
2.3.4 Log-Log Model (Double Log Model)
Sometimes, potential models are postulated in economic theory, such as the well-known Cobb-Douglas
functional form. This form is very popular for estimating production and demand functions. A potential model
with a unique explanatory variable is given by:
𝒀𝒊 = 𝒆𝜶 𝑿𝜷 𝒆𝑼
This model is not linear in the parameters, but it is linearizable by taking natural logarithms, and the following
is obtained:
𝒍𝒏(𝒀𝒊 ) = 𝜶 + 𝜷𝒍𝒏𝑿𝒊 + 𝑼𝒊 (2.18)
The corresponding fitted model to (2.18) is the following:
̂𝒊 ) = 𝜶
𝒍𝒏(𝒀 ̂ 𝒍𝒏𝑿𝒊
̂+𝜷 (2.19)
Taking first order differences in (2.19), we obtain
̂ 𝒊) = 𝜶
∆𝒍𝒏(𝒀 ̂ ∆𝒍𝒏𝑿𝒊
̂+𝜷 (2.20)
The slope coefficient in this model measures relative change in Y for a given relative change in X.
𝐑𝐞𝐥𝐚𝐭𝐢𝐯𝐞 𝐜𝐡𝐚𝐧𝐠𝐞 𝐢𝐧 𝐭𝐡𝐞 𝐝𝐞𝐩𝐞𝐧𝐝𝐞𝐧𝐭 𝐯𝐚𝐫𝐢𝐚𝐛𝐥𝐞 ∆𝒀⁄
̂=
I.e., 𝜷 = 𝒀
𝐑𝐞𝐥𝐚𝐭𝐢𝐯𝐞 𝐜𝐡𝐚𝐧𝐠𝐞 𝐢𝐧 𝐭𝐡𝐞 𝐞𝐱𝐩𝐥𝐚𝐧𝐚𝐭𝐨𝐫𝐲 𝐚𝐫𝐢𝐚𝐛𝐥𝐞 ∆𝑿⁄
𝑿
Therefore, multiplying both sides of (2.20) by 100%, we obtain percentage relationships
̂ 𝒊 )𝟏𝟎𝟎% = 𝜶
∆𝒍𝒏(𝒀 ̂ ∆𝒍𝒏𝑿𝒊 (𝟏𝟎𝟎%)
̂+𝜷
In case you remember the term elasticity: in this case 𝛽̂ represents elasticity of Y for X. If X increases by 1%,
then 𝑌̂ will increase by 𝛽̂ %. It is important to remark that, in this model, 𝛽̂ is the estimated elasticity of Y with
respect to X, for any value of X and Y. Consequently, in this model the elasticity is constant.
For example, suppose that to examine the effect of coffee price on quantity demanded of coffee, an investigator
estimated the following Log-Log Model:
̂ 𝒍𝒏(𝑪𝒐𝒇𝒇𝒑𝒓𝒊𝒄𝒆) + 𝑼𝒊
𝒍𝒏(𝑸𝑫𝑪𝒐𝒇𝒇) = 𝜶 + 𝜷

Page 13
CHAPTER TWO: SIMPLE LINEAR REGRESSION 2022

The fitted log-log model has been given as:


̂ ) = 𝟑. 𝟓 − 𝟓. 𝟐𝟓𝒍𝒏(𝒄𝒐𝒇𝒇𝒑𝒓𝒊𝒄𝒆)
𝒍𝒏(𝑸𝑫𝒄𝒐𝒇𝒇
This means that, if the price of coffee increases by 1%, the quantity demanded of coffee will decrease by 5.25%.
In this case, 𝛽̂ represents the estimated price elasticity of demand. Consider Table 2.1 for a summary of the
interpretation of the slope for the various functional forms.
̂ in different models
Table 2.1: Summary, interpretation of 𝜷

Model If X increases by Then Y will change by


Linear 1 unit 𝛽̂ 𝑢𝑛𝑖𝑡𝑠
Linear-Log 1% (𝛽̂ /100) 𝑢𝑛𝑖𝑡𝑠
Log-Linear 1 unit (𝛽̂ ∗ 100)%
Log-Log 1% 𝛽̂ %

2.4 Decomposition of Variation and Goodness of Fit


Once the parameter estimation has been done, we can see how well our sample regression line fits our data. In
other words, we can measure how well the explanatory variable, X, explains the dependent variable, Y. Usually
this is done by decomposing (that means categorizing) the variance of Y into groups. To avoid negative values
cancelling out positive values, the values are squared.
The variation of Y is measured as the sums of squares of the difference between 𝑌𝑖 and 𝑌̅ (the mean value of Y).
The Total Sum of Squares includes all the variation in Y.
TSS Total Sum of Squares ̅ )𝟐
∑(𝒀𝒊 − 𝒀 (2.21)
This variance of Y is partially explained by the model (based on the variation in X) and partially unexplained by
the model (included in U, the random error term). The part explained by the model, the Model Sum of Squares
(MSS), is the squared difference between 𝑌̂𝑖 and 𝑌̅. Note that 𝑌̂𝑖 is the predicted Y value for observation i, given
𝑋𝑖 . Mathematically, 𝑌̂𝑖 = 𝛼̂ + 𝛽̂ 𝑋𝑖 .
MSS Model Sum of Squares ̂𝒊 − 𝒀
∑(𝒀 ̅ )𝟐 (2.22)
Lastly, the unexplained variation is the squared difference between the observed value of Y and the predicted
value of Y according to the model, the difference between 𝑌𝑖 and 𝑌̂𝑖 .
RSS Residual Sum of Squares ̂ 𝒊 )𝟐
∑(𝒀𝒊 − 𝒀 (2.23)
Note that, because the variation in Y is either explained in the model (part of MSS) or not explained in the model
(part of RSS), it must be that TSS=MSS+RSS
For example, while predicting the effect of initial capital (X) on mini-enterprise profitability (Y), we may use
the following data of 5 mini enterprises that all started 2013 E.C. The profitability is measured as ROA of 2014
E.C.

Page 14
CHAPTER TWO: SIMPLE LINEAR REGRESSION 2022

Initial Capital ROA in


in 100,000 Birr percent
X Y ̅
𝒀𝒊 − 𝒀 ̅ )𝟐
(𝒀𝒊 − 𝒀 ̂𝒊
𝒀 ̂𝒊 − 𝒀
(𝒀 ̅) ̂𝒊 − 𝒀
(𝒀 ̅ )𝟐 𝒀 − 𝒀
̂𝒊 ̂ 𝒊 )𝟐
(𝒀 − 𝒀
6 4 -2 4 4.2 -1.8 3.24 -0.2 0.04
5 2 -4 16 2.4 -3.6 12.96 -0.4 0.16
8 8 2 4 7.8 1.8 3.24 0.2 0.04
7 7 1 1 6 0 0 1 1
9 9 3 9 9.6 3.6 12.96 -0.6 0.36
sum 34 32.4 1.6

To calculate the total sum of squares, first the mean of Y should be calculated, which is 6. The total sum of
squares (TSS) is ∑(𝑌𝑖 − 𝑌̅)2, which is 34 (see the fourth column the table). That is the total variation for which
explanation is needed. When using the model 𝑌𝑖 = 𝛼̂ + 𝛽̂ 𝑋𝑖 + 𝑒𝑖 , with Y being ROA and X being initial capital,
you will find that the OLS method leads to 𝛼̂=-6.6 and 𝛽̂ =1.8. This helps us to find 𝑌̂𝑖 for each i. In this way, we
can calculate the model sum of squared (MSS), which is ∑(𝑌̂𝑖 − 𝑌̅)2 , is 32.4 (see the 7th column of the table).
This MSS states that from all the variation (TSS = 34), 32.4 could be explained by the linear regression model.
Lastly, the residual sum of squared (RSS) can be calculated by squaring the error terms for each observation,
which is 1.6 (see the last column of the table). Note that TSS=MSS+RSS, because 34=32.4+1.6. We decomposed
the total variation of 34 into two elements: the part explained by the model (32.4) and the part not explained by
the model (1.6).

This decomposition of variation can be used to make statements about the goodness of fit. The goodness of fit of
a model is measured by a statistical index called coefficient of determination (𝐑𝟐 ). The Coefficient of
Determination (R2 ) is the proportion of total variation of (Y) which is explained by the variation of the explanatory
variable (X) included in the model.
𝑬𝒙𝒑𝒍𝒂𝒊𝒏𝒆𝒅 𝑽𝒂𝒓𝒊𝒂𝒕𝒊𝒐𝒏 𝒊𝒏 𝒀
𝑪𝒐𝒆𝒇𝒇𝒊𝒄𝒊𝒆𝒏𝒕 𝒐𝒇 𝑫𝒆𝒕𝒆𝒓𝒎𝒊𝒏𝒂𝒕𝒊𝒐𝒏 ( 𝑹𝟐 ) =
𝑻𝒐𝒕𝒂𝒍 𝑽𝒂𝒓𝒊𝒂𝒕𝒊𝒐𝒏 𝒊𝒏 𝒀
𝑴𝒐𝒅𝒆𝒍 𝑺𝒖𝒎 𝒐𝒇 𝑺𝒒𝒖𝒂𝒓𝒆𝒔 𝑴𝑺𝑺 ∑(𝒀 ̅ )𝟐
̂ 𝒊 −𝒀
= = ∑(𝒀𝒊 −𝒀̅ )𝟐
(2.24)
𝑻𝒐𝒕𝒂𝒍 𝑺𝒖𝒎 𝒐𝒇 𝑺𝒒𝒖𝒂𝒓𝒆𝒔 𝑻𝑺𝑺
The notion of this index is straightforward in a sense that a model will be good fit if the explanatory variable (X)
included in the model determines large part of the actual values of Y. But if X is irrelevant variable, then model
would predict no part of the actual values of X.

Note that the total variation in the observed Y values about their mean value can be partitioned into two parts, one
attributable to the regression line and the other to random forces because not all actual Y observations lie on the
fitted line. Geometrically, this breakdown is displayed in Figure 2.9.

Page 15
CHAPTER TWO: SIMPLE LINEAR REGRESSION 2022

Figure 2.9: Breakdown of the variation of Y into two components

Note that, because TSS=MSS + RSS,


𝑴𝑺𝑺 𝑻𝑺𝑺−𝑹𝑺𝑺 𝑹𝑺𝑺
𝑹𝟐 = = = 𝟏 − 𝑻𝑺𝑺 (2.25)
𝑻𝑺𝑺 𝑻𝑺𝑺

𝑹𝟐 measures part of total variation in Y which is explained by the model. As a result, it is used as an indicator to
measure the explanatory power or goodness of fit of a model. It shows the proportion of total variation in Y which
is attributable for the variation in X.

For example, in the above example where the ROA of mini-enterprises was predicted based on their start capital,
32.4 34−1.6
the TSS was 34, the MSS was 32.4 and the RSS was 1.6. Therefore, the R 2 for this case is 𝑅 2 = = =
34 34
1.6
1 − 34 ≈ 0.95.

How can this measure of the goodness of fit be interpreted? An 𝑅 2 of 0.95 means that 95% of the variation in Y
(the variation in ROA) can be explained by the model. That means that approximately 5% of the variation in Y is
not explained by the model. Equivalently, it would mean that, 90% of the total variation in the values of Y is
explained or determined (caused) by the variation in the values of X (the start up capital size). Therefore, the
model has a good fit.

Note that the maximum value for R2 is 1, which would indicate that 100% of the variation in Y is explained by
the variation in X. This would be a deterministic non-stochastic model, without any residual term. The lowest
value for R2 is 0. If 0% of the variation in Y is explained by the variation in X it means that the model there is no
difference between the mean value of Y (𝑌̅ ) and the predicted value of Y (𝑌̂𝑖 ) for each i.

Page 16
CHAPTER TWO: SIMPLE LINEAR REGRESSION 2022

Decomposition of the variation.


Note that
ModelSS+ResidualSS=TotalSS
Value of
R2

̂ and 𝜶
𝜷 ̂

Figure 2.10: Stata output of the SLR between ROA and Initial_Capital

Figure 2.10 displays the calculated parameter estimates and decomposition of the sums of square (SS). Note that,
also, the R2 is included in the output.

2.5 Properties of OLS estimates and the Gauss-Markov theorem


When using the OLS estimation method to find parameter estimates for a model based on a sample, we must
wonder if the parameter estimates are good. We would like estimates, for example the calculated value for 𝛽̂ , to
be very close to the value of the true population parameters (𝛽). Very close means that possible 𝛽̂ ’s based on
different samples vary within only a small range around 𝛽, the true population parameter.

If the assumptions as described in section 2.2 are met, OLS estimates are known from being very good. The ideal
of optimum properties that the OLS estimates possess may be summarized by the well-known theorem known as
the Gauss-Markov Theorem, named after Carl Friedrich Gauss (a German Mathematician) and Andrey Markov
(a well-known Russian Mathematician).

The Gauss-Markov Theorem can be stated as: “Given the assumptions of the classical linear regression model,
the OLS estimators, in the class of linear and unbiased estimators, have minimum variance. In other words,
the OLS estimators are BLUE.”

An estimator is called BLUE if it is Best Linear Unbiased Efficient


a. Linear: a linear function of a random variable, such as, the dependent variable, 𝒀.
An estimator is linear if it is a linear function of the dependent variable of the model. Linearity of an estimate
does not determine the closeness of an estimate to the true population parameter. However, it is desirable in
Page 17
CHAPTER TWO: SIMPLE LINEAR REGRESSION 2022

econometric analysis because it is useful to deduce the probability distribution of an estimate from the known
probability distribution of the dependent variable of a model.

b. Unbiased: its average or expected value is equal to the true population parameter.
According to this criterion a good estimator is one that produces an unbiased estimate. An estimate is said to be
unbiased if its bias is zero. The bias of an estimate is defined by the difference between the expected value of the
estimate and the value of the population parameter.
That is,
̂) − 𝜷
𝑩𝒊𝒂𝒔 𝒐𝒇 𝒂𝒏 𝑬𝒔𝒕𝒊𝒎𝒂𝒕𝒆, 𝜷 = 𝑬(𝜷
̂ ) is unbiased if
Thus, an estimate 𝑜𝑓(𝜷
̂) = 𝜷
𝑬(𝜷 (2.26)

Figure 2.11: Biased and unbiased estimates of 𝜽

Unbiasedness is desirable but not sufficient alone. Unbiasedness is good to be attained along with minimum
variance. This is so because even an unbiased estimate could be far from the true population parameter unless it
has minimum variance. Minimum variance means that the estimate has a minimum variance in the class of linear
and unbiased estimators.

An estimator is best if it has the smallest variance as compared to any other estimators obtained from other
̂ is best if:
econometric method. That is, an estimator 𝜷
̂ ) < 𝑽𝒂𝒓(𝜷
𝑽𝒂𝒓(𝜷 ̃)

𝑬[𝜷 ̂ )]𝟐 < 𝑬[𝜷


̂ − 𝑬(𝜷 ̃ )]𝟐
̃ − 𝑬(𝜷 (2.27)
̃ is any other estimate of 𝛽 obtained from other econometric technique.
Where: 𝜷

c. An unbiased estimator with the least variance is known as an efficient estimator.

𝑬𝒇𝒇𝒊𝒄𝒊𝒆𝒏𝒄𝒚 = 𝑼𝒏𝒃𝒊𝒂𝒔𝒆𝒅𝒏𝒆𝒔𝒔 + 𝑴𝒊𝒏𝒊𝒎𝒖𝒎 𝑽𝒂𝒓𝒊𝒂𝒏𝒄𝒆

Page 18
CHAPTER TWO: SIMPLE LINEAR REGRESSION 2022

Efficiency is a desirable statistical property because from two unbiased estimators of the same population
parameter, we prefer the one that has the smaller variance, i.e., the one that is statistically more precise.
Let β̂ and β̃ be two unbiased estimators of the population parameter 𝛃, such that 𝐄(𝛃
̂ ) = 𝛃 and 𝐄(𝛃
̃) = 𝛃.
̂ is efficient relative to the estimator 𝛃
Then the estimator 𝛃 ̃ if the variance of the finite-sample distribution of
̂ is less than the variance of the finite-sample distribution of 𝛃
𝛃 ̃;

Reading assignment: read, as background material, the mathematical proof for the fact that, if the assumptions of
SLR are met, the OLS estimates are BLUE.

2.6 Evaluation of Estimates: Confidence Intervals and Hypothesis Testing


2.6.1 Introduction into estimate evaluation
In this section, we shall develop statistical criteria for the evaluation of an estimated model. Statistical criteria are
developed based on statistical and probability theories, and they are called first order tests of a model. In doing
econometrics we are interested to know the true parameters ( and ), but we can only get estimates for the
parameters based on the sample (𝛼̂ and 𝛽̂ ). Sample values of 𝜶 ̂ might be different from zero by chance.
̂ 𝑎𝑛𝑑 𝜷
̂ can be different from zero not for the fact that 𝜶 and 𝜷 are actually different
̂ 𝑎𝑛𝑑 𝜷
That is, the sample values of 𝜶
from zero but by a mere chance that happened in a sample. Therefore, in econometric applications it is mandatory
to test whether the true population values of 𝜶 and 𝜷 are different from zero. The two approaches that can be
used for evaluating estimates are confidence interval calculation and hypothesis testing.

2.6.2 Confidence interval


In the confidence interval evaluation approach we estimate the interval of values between which the true values
of the population parameters are expected to lie within a certain “degree of confidence”. In statistics, the process
of estimating an interval of values between which the true values of the population parameters are expected to lie
based on the sampling distribution of the sample estimates is called interval estimation.

An interval of values constructed (estimated) to predict ranges of values for the unknown population parameters
̂ & the actual sample values of 𝛂
̂&𝛃
𝛂 and 𝛃 based on some clues (i.e., the sampling distribution of 𝛂 ̂
̂&𝛃
obtained from the sample) is known as confidence interval for the population parameters 𝛂 and 𝛃. And the
degree of certainty we may assign to the interval to contain the true value of the population parameter within its
range is known as confidence level or confidence coefficient. The confidence interval can be computed from a
Z-distribution (if the sample is large n>30) or from a t-distribution (if the sample is small, n<30).

Page 19
CHAPTER TWO: SIMPLE LINEAR REGRESSION 2022

I. Confidence interval from the Standard Normal Distribution (Z−Distribution)


If the sample is large, we can use the Z-transformation formula together with Z-distribution probability table
(Table 2.2) to construct confidence interval for 𝜷. To do this we must follow the following steps:
Step – 1: Choose the confidence coefficient (confidence level 𝛅). This is the probability for the interval to contain
the true value of 𝜷 within its range. In other words: it is how sure you want to be about being right. It is customary
in econometric analysis to choose 95% as confidence coefficient, but one can Table 2.2: z-values given the
confidence interval
choose any level of confidence coefficient as he/she may wish.
Step – 2: Construct a 𝜹% standardized confidence interval. Determine
standardized interval of values within which there are 𝜹% standardized
values. That is:
𝑷𝒓(𝒁𝟏 < 𝒁 < 𝒁𝟐 ) = 𝜹
For instance, if 𝜹 is 95% then the standardized confidence interval is
(−𝟏. 𝟗𝟔, 𝟏. 𝟗𝟔). That is: 𝑷𝒓(−𝟏. 𝟗𝟔 < 𝒁 < 𝟏. 𝟗𝟔) = 𝟎. 𝟗𝟓. See Table 2.2 for the z-values that correspond to
other confidence levels.
Step -3: Change the standardized confidence interval into actual value confidence interval. The actual value
confidence interval for 𝜷 based on a Z-distribution can be calculated by using:
̂ ± 𝒁𝒕 . 𝑺𝑬(𝜷
𝜷=𝜷 ̂) (2.28)
̂ will be:
For instance, if 𝜹% is 95% then the actual 95% confidence interval for 𝜷
̂ ± 𝟏. 𝟗𝟔. 𝑺𝑬(𝜷
𝜷=𝜷 ̂)

𝒁𝟏 = −𝟏. 𝟗𝟔 𝒁𝟐 = 𝟏. 𝟗𝟔
Figure 2.12: Graphical representation of a 95% confidence interval

Step-4: Interpretation – if we are curious if X is affecting Y significantly, we have to see if  could be 0. If the
95% confidence interval includes 0, it is possible that  is 0, indicating that there is no significant effect of X on
Y. However, if the confidence interval does not include 0, it is likely that X affects Y.

Page 20
CHAPTER TWO: SIMPLE LINEAR REGRESSION 2022

For example, let’s construct a 95% confidence interval for the population slope parameter 𝜷, for a supply
function.
𝑸 = 𝟏𝟎𝟎 + 𝟒𝑷
𝑺𝑬: (𝟏𝟎) (𝟏. 𝟓)

Where, the values in the bracket are standard errors and 𝒏 = 𝟕𝟎𝟎 producers

The 95% confidence interval can be calculated by 𝛽̂ ± 1.96. 𝑆𝐸(𝛽̂ ). The lower limit is 𝛽̂ − 1.96 ∗ 𝑆𝐸(𝛽̂ ) = 4 −
1.96 ∗ 1.5 = 1.06 and the upper limit is 𝛽̂ + 1.96 ∗ 𝑆𝐸(𝛽̂ ) = 4 + 1.96 ∗ 1.5 = 6.94. This interval is a random
interval. If the sampling process and confidence interval calculation will be repeated, it will include the true
population  in at least 95% of all cases. Because 0 is not included in the confidence interval, we can say that 
is statistically significantly different from 0. Therefore, P does significantly positively affect Q.

II. Confidence interval from the Student’s t-distribution.


For smaller sample sizes the Z-distribution is not applicable. If 𝑛 ≤ 30, the student’s t-distribution can be used to
construct a confidence interval. The procedure for constructing confidence interval with t-distribution is similar
to the one outlined for Z-distribution the only difference is we must take into account the degrees of freedom
while using t-distribution probability table (Table 2.3). In simple linear regression the df is n-2. Therefore, for
any confidence level, 𝜹%, the confidence interval for 𝜶 and 𝜷 based on t-distribution can be given as: 𝛼 = 𝛼̂ ±
𝑡𝛼⁄2 . 𝑆𝐸(𝛼̂) and 𝛽 = 𝛽̂ ± 𝑡𝛼⁄2 . 𝑆𝐸(𝛽̂ ). Take the t-value from Table 2.3.
Table 2.3: Table with t-values for different levels of significance and degrees of freedom

Stata regression output also includes 95% confidence intervals. Consider the output in Figure 2.13 where the
number of clients of Safaricom are predicted, based on their expenses on advertisement.

Page 21
CHAPTER TWO: SIMPLE LINEAR REGRESSION 2022
The confidence interval for  is 0.669---1.192
The confidence interval for  is 0.014---2.276
Both  and  are likely different from 0, because the confidence interval excludes 0.

Figure 2.13: Stata output including the Confidence interval of α and β

2.6.3 Hypothesis testing


Test of the statistical significance of the parameter estimates is determining whether sample values of the
parameter estimates are statistically different from zero or not. The true population values of 𝜶 and 𝜷 are
unobservable (can’t be estimated) so how can we determine whether their values are zero or not? In econometric
̂ together with sample values of 𝜶
̂ and 𝜷
applications we use the sampling distributions of 𝜶 ̂ to compute
̂ and 𝜷
̂ to be zero. In this case, if it is less likely for 𝜶 and 𝜷 to be zero, then 𝜶 and 𝜷 are said
̂ and 𝜷
how likely is for 𝜶
to be statistically different from zero and the reverse will be the conclusion if it is highly likely for 𝜶 and 𝜷 to be
̂ are statistically significant (statistically
̂ and 𝜷
zero. If 𝜶 and 𝜷 are statistically different from zero, then we say 𝜶
different from zero).
Therefore, test of the statistical significance of the parameter estimates (𝛼̂ 𝑎𝑛𝑑 𝛽̂ ) refers to the process in which
̂ together with their sample values to compute how
̂ 𝑎𝑛𝑑 𝜷
we use the sampling distributions of OLS estimators 𝜶
likely is for 𝜶 and 𝜷 to be zero to determine whether they are statistically different from zero. In OLS the most
common hypothesis test for a parameter estimate is the t-test.
Steps in the t-test
Step – 1: Determine whether you need a one-tailed or a two-tailed test.
The test will be one tailed test if it is possible to determine the possible signs of 𝛼̂ & 𝛽̂ to be either positive or
negative using prior information. In this case the test will be right tail test if the signs of 𝛼̂ & 𝛽̂ are only positive
and it will be left tail test if the possible signs of 𝛼̂ & 𝛽̂ are only negative. On the other hand, the test must be a
two tailed test if prior information as regard to the possible signs of 𝛼̂ & 𝛽̂ is not available. For example, if you
want to know if age has an effect on the usage of mobile phones for financial transactions, you may want to test
the hypothesis that age has no effect at all (𝛽 = 0) or you might want to test if older people are less likely to use
mobile phones for financial transactions (𝛽 < 0). The first case requires a two-tailed test, the latter a left-tailed
test.

Page 22
CHAPTER TWO: SIMPLE LINEAR REGRESSION 2022

Step - 2: State the null hypothesis (𝑯𝟎 ) and the alternative hypothesis (𝑯𝑨 ) of the test.
Table 2.4 provides an overview of the null hypothesis and the alternative hypothesis for various standard t-tests.
Table 2.4: Overview of the hypotheses for three types of tests
For one-tailed test
For two-tailed test Right-tailed Left-tailed
𝐻0 : 𝛼 = 0 𝑜𝑟 𝛽 = 0 𝐻0 : 𝛼 = 0 𝑜𝑟 𝛽 = 0 𝐻0 : 𝛼 = 0 𝑜𝑟 𝛽 = 0
𝐻𝐴 : 𝛼 ≠ 0 𝑜𝑟 𝛽 ≠ 0 𝐻𝐴 : 𝛼 > 0 𝑜𝑟 𝛽 > 0 𝐻𝐴 : 𝛼 < 0 𝑜𝑟 𝛽 < 0

Step - 3: State the level of significance


The level of significance is the subjective probability that a researcher considers as too low to assume. It is the
probability of making ‘wrong’ decision, i.e., the probability of rejecting the hypothesis(𝑯𝟎 ) while it is
actually true or the probability of committing a type I error. It is customary in econometric research to choose
the 5% or the 1% level of significance. 5% level of significance means that in making our decision we allow
(tolerate) five times out of a hundred to be ‘wrong’ i.e., reject the hypothesis when it is actually true.

Step – 4: Get tt (given the level of significance and the degrees of freedom) from the table
tt is the critical value for the t-test, which is obtained from Table 2.5. The critical value/s is/are the maximum
and/or minimum standardized value/s beyond which there is only 𝛼% (level of significance) chance for the
estimate to assume values under the null hypothesis. The critical value/s of the test is/are obtained from the t-
table. As a result, the critical values are also known as table or theoretical values. 𝑡𝑡 should be found at 𝛼⁄2 and
with n-2 df for a two tailed test. For a one tailed test, use 𝛼. For example, for a two-tailed test with a sample size
of 28 respondents and a significance level of 5%, tn-2,0.05/2=t26,0.025 =2.056. That means that, under the null
hypothesis, the test statistic of  should with 95% certainty fall between -2.056 and +2.056.

Page 23
CHAPTER TWO: SIMPLE LINEAR REGRESSION 2022

Table 2.5: t-values given the degrees of freedom and the desired significance level

Step - 5: Compute the test statistic (𝒕𝒄 )


tc is called computed value of 𝑡 (or test-statistics of the estimator), by taking the value of 𝛼 𝑜𝑟 𝛽 in the null
hypothesis.
𝐄𝐬𝐭𝐢𝐦𝐚𝐭𝐨𝐫 – 𝐇𝐲𝐩𝐨𝐭𝐡𝐞𝐬𝐢𝐳𝐞𝐝 𝐯𝐚𝐥𝐮𝐞
𝒕𝒄 = (2.29)
𝐒𝐭𝐚𝐧𝐝𝐚𝐫𝐝 𝐄𝐫𝐫𝐨𝐫 𝐨𝐟 𝐭𝐡𝐞 𝐞𝐬𝐭𝐢𝐦𝐚𝐭𝐨𝐫
For 𝛽, for example

Page 24
CHAPTER TWO: SIMPLE LINEAR REGRESSION 2022

̂ −𝟎
𝜷 ̂
𝜷
𝒕𝒄 = ̂) = ̂) (2.30)
𝑺𝑬(𝜷 𝑺𝑬(𝜷

if the null hypothesis is two-tailed and stated as 𝛽 = 0.


Step-6: Compare the computed value 𝐭 𝐜 with the table value 𝐭 𝐭 & conclude about statistical significance
If 𝑡𝑐 > 𝑡𝑡 or if 𝑡𝑐 < −𝑡𝑡 , reject 𝐻0 and accept 𝐻𝐴 . That means that the parameter (e.g. 𝛽̂ ) is statistically significant.
If the calculated t is within the range -tt and +tt, (−𝑡𝑡 < 𝑡𝑐 < 𝑡𝑡 ) accept 𝐻0 and reject 𝐻𝐴 . That means that the
parameter (e.g. 𝛽̂ ) is statistically insignificant.

For example, you remember that in section 2.3 from a sample of size 𝑛 = 6, we have estimated the following
simple consumption function:
̂ 𝑖 = 1.53 + 0.607𝑖𝑛𝑐𝑖
𝐶𝑜𝑛𝑠
𝑆𝐸: (0.54) (0.056)
The values in the brackets are standard errors. Let’s test the hypothesis that income doesn’t affect consumption
expenditure using the t-test at 5% level of significance. In other words: let’s test the statistical significance of the
.
Step 1: We will do a two-tailed test, because we want if income affects consumption in general (not only if
income affects consumption positively).
Step 2: The hypothesis we want to test: 𝑯𝟎 : 𝜷 = 𝟎 against 𝑯𝑨 : 𝜷 ≠ 𝟎.
Step 3: The level of significance is 5%, because that is given above.
Step 4: The value from the table, tt, should be the value with 5%/2 significance and df of 6-2. t4,0.025=2.776
Step 5: The computed test statistic is
𝛽̂ − 0 𝛽̂ 0.607
𝑡𝛽̂ = = = ≅ 𝟏𝟎. 𝟖𝟒
𝑆𝐸(𝛽̂ ) 𝑆𝐸(𝛽̂ ) 0.056
̂ is statistically significant. It means  is not zero, and
Step 6: t c = 10.84 > t t = 2.776 so that means that 𝛃
therefore we can conclude that income has a significant effect on consumption.

Another way to do the t-test


When doing parameter estimation by using a software program like Stata or Excel, the t-value is displayed.
However, next to the t-value, output includes a p-value for each estimated parameter. The p-value is the actual
probability of committing type-I error in a particular test. That is, the actual probability of rejecting the null-
hypothesis, while it’s true. It shows the possibility (chance) being wrong if we reject the null hypothesis. Thus,

Page 25
CHAPTER TWO: SIMPLE LINEAR REGRESSION 2022

small p-values are evidence against the null hypothesis. That means, it’s possible to reject the null hypothesis if
the p-values are smaller relative to the level of significance.

In other words, the p-value is the lowest level of significance at which the observed value of a test statistic is
significant (i.e., one rejects 𝐻0 ). The p-value (or probability value) is the probability of getting a sample statistic
(such as 𝛼̂ 𝑜𝑟 𝛽̂ ) when the null hypothesis is true. The p-value for obtaining a sample outcome is compared to the
level of significance (𝛼). By checking if the p-value is less than the intended level of significance, e.g. 5%, an
econometrician can immediately conclude about the parameter’s significance, without following all the steps.
Especially if software is used to analyze the data it is common to immediately look at the p-value rather than
following all steps. See Figure 2.14.

p-values of the estimates


Standard Calculated t-
error of the value of the note that  is significant at 1% (because
estimates estimates p<0.01) and  is significant at 5%
(because p<0.05)

Figure 2.14: Stata output displaying information relevant for conducting the t-test

Instead of reporting p-values by value, another common notation in published research work is the asterisks (*).
One * indicates significance at a level of 10% (that means the p-value is on the range 5%-10%, or 0.05-0.01), two
asterisks (**) indicates significance at a level of 5% (that means the p-value is on the range 1%-5%, or 0.01-0.05).
Lastly, three asterisks (***) indicates very strong significance, because the p-value is less than 1% or less than
0.01.

2.7 Predictions with Simple Linear Regression Models


Forecasting using regression involves making predictions about the dependent variable based on average
relationships observed in the estimated model. Predicted values are values of the dependent variable based on
the estimated model and a prediction about the values of the independent variables.
For a simple regression, the value of Y is predicted as:

̂=𝜶
𝒀 ̂ 𝑿𝒑
̂+𝜷 (2.31)

Page 26
CHAPTER TWO: SIMPLE LINEAR REGRESSION 2022

̂ is the predicted value of the dependent variable, and


Where, 𝒀
𝑿𝒑 is the predicted value of the independent variable (input).
The estimated intercept and all the estimated slopes are used in the prediction of the value of the dependent
variable, even if a slope is not statistically significantly different from zero.
For example: suppose that you’ve estimated and evaluated the following model:
𝑌𝑖 = 𝛼 + 𝛽𝑋𝑖 + 𝑢𝑖
Where Yi is the sales of second-hand sport shoes of shops in Arba Minch, and 𝑋𝑖 is the common price they charge
for second-hand sport shoes. Figure 2.15 displays the relation, which is in line with the law of demand that states
that demand is less if the price is high.

Figure 2.15: Scatterplot showing the relation between sales (Y) and price (X) of second hand sport shoes.

SLR output of Stata is displayed in Figure 2.16 and reveals that 𝛼̂ = 243.2 and 𝛽̂ = −0.2198. Both
coefficients,  and  are significantly different from 0 (because p<0.05 for both coefficients).

Figure 2.16: SLR output for the relation between sales (Y) and prices of second-hand sport shoes (X).

Page 27
CHAPTER TWO: SIMPLE LINEAR REGRESSION 2022

This model can be used to forecast the sales of second-hand sport shoes, if the price is known. For example, if the
price of second-hand sport shoes is 500 Birr, the sales are expected to be:
𝑌̂𝑖 = 243.2 − 0.2198𝑋𝑖 = 243.2 − 0.2198 ∗ 500 = 133 pairs of shoes.
Another example, suppose you have an estimated model of sales (Y) of firms producing a particular product as
a function of advertisement expenditure (X):
̂ = 𝟑𝟎𝟎 + 𝟐. 𝟓𝑿𝟏
𝒀
(p:0.21) (p:0.003)
In addition, you have predicted value for the independent variable (advertisement expenditure in ETB), 𝐗 𝟏 =
𝟐𝟎𝟎. Then the predicted value for Sales is 800 Units. The fact that the intercept is insignificant and the slope is
significant does not affect if they should be included: all parameter estimates should be included in the predicted
value.
̂ 𝐢 = 𝟏. 𝟓𝟑 + 𝟎. 𝟔𝟎𝟕𝐢𝐧𝐜𝐢 . Based on this estimated
Last example: Recall our estimated consumption Model: 𝐂𝐨𝐧𝐬
model predict the consumption expenditure of a household whose income is 12 ETB?
̂ 𝐢 = 𝟏. 𝟓𝟑 + 𝟎. 𝟔𝟎𝟕(𝟏𝟐) = 𝟖. 𝟖𝟏𝟒. That means, a household with an income of 12 ETB will spend
Solution: 𝐂𝐨𝐧𝐬
8.814 ETB of its income for consumption.

Critical note in forecasting/predicting


Note that the forecasted/predicted values are only valid for conditions similar as under which the sample was
collected. In the case of the first example about shoes sales, it is impossible to predict the sales for the case in
which the shop will charge a price of 1,000,000 Birr for a pair of shoes, because such high price observation is
absent in the datafile. Also, it would predict negative shoes sales, which is, of course, impossible.

Page 28

You might also like