Understanding Regression Analysis Basics
Understanding Regression Analysis Basics
REGRESSION ANALYSIS
UNIT INTRODUCTION
The term “regression” was originally employed by Galton in 1877 for indicating certain
relationships in the theory of heredity3 but it has come to imply the statistical method developed to
investigate those relationships. There are problems in business and industry where two or more
variables show a mutual relationship and are hence capable of simultaneous analysis. One of the
variables is of primary interest, which may not be measured directly. On certain other occasions
the variable of primary interest can be measured but direct measurement of this variable is very
expensive. In these problems we first try to build up a suitable functional relationship between the
primary variable and one or more auxiliary variables. On the basis of this functional relationship
we attempt to predict the value of the primary variables for given values of the auxiliary variables.
Several illustrations can be given.
A) A marketing research unit may be interested in knowing whether increased advertising will
result in increased sales. The variable, which is of primary interest, in this case is sales.
The auxiliary variable considered in this case is advertising expenditure. If on the basis of
past record, it is possible to build up a suitable functional relationship between advertising
expenditure and sales, it can be used to predict sales for a given increase in advertising
expenditure.
B) A manufacturer of farm tools wishes to study whether the volume of sales can be predicted
from the corresponding farm income. On the basis of the data on sales and farm income a
prediction model can be obtained which may be used to predict sales from the farm income.
As illustrated in the above examples, there are a lot of variables in a business that vary together in
some direction. Those relationships will be studied will be studied under econometrics using a
concept of regression. In regression, our task is usually to estimate the population regression
3 Galton come up with, the height of the children of unusually tall or unusually short parents tends to move toward
the average height of the population. That is, the average height of sons of a group of tall fathers was less than their
fathers’ height and the average height of sons of a group of short fathers was greater than their fathers’ height, thus
“regressing” tall and short sons alike toward the average height of all men.
1
function (PRF) on the basis of sample regression function (SRF) as accurately as possible. To do
so, this unit is organized into five sections so as you understand regression.
The first section deals with the concept of regression. Regression is defined as a statistical method
that helps us to analyze and understand the relationship between two or more variables of interest.
The process that is adapted to perform regression analysis helps to understand which factors are
important, which factors can be ignored, and how they are influencing each other. The second
section of the unit discusses the two most important types of relationship; stochastic and
deterministic relationship. A relationship between X and Y, characterized as Y = f(X) is said to be
deterministic if for each value of the X there is one and only one corresponding value of Y, while
stochastic relationship is when for each value of X there is probability distribution of Y.
In section three we introduce the simple linear regression model which shows the relationship
between dependent and one independent variable in linear form. Within this section, we will
discuss the method of classical linear regression model with their assumptions. In addition to this,
we introduce the method of least squares that requires minimizing sum of square of residual ∑ 𝑒𝑖
and we will discuss the desirable statistical properties of least square estimates and the Gauss-
Markov theorem. The fourth section discusses the concept of covariance, correlation and
regression. Covariance measures the direction of the relationship between two variables, whereas,
correlation measures the extent in which those variables are related. Finally, the fifth section
illustrates the procedures for testing of hypothesis, performing the student’s test and confidence
interval. It also provides a comparison between regression analysis and ANOVA.
2
2.1. Definition of Regression Analysis
After having established the fact that two variables are closely related we may be interested in
estimating (predicting) the values of one variable given the value of another. Regression analysis
is concerned with the study of the dependence of one variable (the dependent/explained variable),
on one or more other variables (explanatory/independent variables), with a view to estimating
and/or predicting the (population) mean or average value of the former in terms of the known or
fixed (in repeated sampling) values of the latter. The dependent variable is a variable which its
average value is computed using the already known values of the explanatory variable(s). But the
values of the explanatory variables are obtained from fixed or in repeated sampling of the
population.
In other words, regression analysis is a mathematical measure of the average relationship between
two or more variables in terms of the original units of the data. In short regression analysis is a
statistical analysis designed to determine the extent to which one factor changes with changes in
another factor or other factors.
Generally, regression analysis is concerned with describing and evaluating the relationship
between a given variable (often called the dependent variable) and one or more variables which
are assumed to influence the given variable (often called independent or explanatory variables).
We denote the dependent variable by 𝑌and the explanatory variables by𝑋1 , 𝑋2 , . . . , 𝑋𝑘 . If k = 1 ,
that is, there is only one explanatory variable (or one of the 𝑋 −variables), we have what is known
as simple regression. This is what we discuss in this unit. On the other hand, if 𝑘 > 1, that is, there
are more than one explanatory variable (or the X − variables), we have what is known as multiple
regression, which will be discussed in depth in unit 3. We shall first discuss two important forms
of relation: stochastic and non-stochastic, among which we shall be using the former in
econometric analysis.
3
only with some probability. Let’s illustrate the distinction between stochastic and non-stochastic
relationships with the help of a supply function.
Assuming that the demand for a certain commodity depends on its price (ceteris paribus) and the
function being linear, the relationship can be put as:
𝑄 = 𝑓(𝑃) = 𝛼 − 𝛽𝑃 − − − − − − − − − − − − − − − − − − − − − − − − − (2.1)
The above relationship between P and Q is such that for a particular value of P, there is only one
corresponding value of Q. This is, therefore, a deterministic relationship since for each price there
is always one and only one corresponding quantity demanded. This implies that all the variation
in Y is due solely to changes in X, and that there are no other factors affecting the dependent
variable.
If this were true all the points of price-quantity pairs, if plotted on a two-dimensional plane, would
fall on a straight line. However, if we gather observations on the quantity actually demanded in
the market at various prices and we plot them on a diagram we see that they do not fall on a straight
line.
Figure 2.1: scatter plot and straight line b/n quantity demand and price
20
15
10
5
0
0 5 10 15 20
Price
Qq Fitted values
The deviation of the observation from the line may be attributed to several factors.
d. Error of aggregation
e. Error of measurement
4
In order to consider, the above sources of errors we introduce in econometric functions a random
variable which is usually denoted by the letter ‘u’ or ‘𝜀’ and is called error term or random
disturbance or stochastic term of the function, so called because u is supposed to ‘disturb’ the exact
linear relationship which is assumed to exist between X and Y. By introducing this random variable
in the function, the model is rendered stochastic of the form:
𝑌𝑖 = 𝛼 + 𝛽𝑋 + 𝑈𝑖 ………………………………………………………. (2.2)
Thus, a stochastic model is a model in which the dependent variable is not only determined by the
explanatory variable(s) included in the model but also by others which are not included in the
model.
To sum up, in statistical relationships among variables we essentially deal with random or
stochastic variables, that is, variables that have probability distributions. In functional or
deterministic dependency, on the other hand, we also deal with variables, but these variables are
not random or stochastic. In regression analysis, we are concerned with statistical dependence
among variables (not Functional or Deterministic), we essentially deal with random or stochastic
variables (with the probability distributions).
𝑌
⏟𝑖 = 𝛼
⏟+ 𝛽𝑥𝑖 + 𝑈
⏟𝑖
𝑇ℎ𝑒 𝑑𝑒𝑝𝑒𝑛𝑑𝑒𝑛𝑡 𝑣𝑎𝑟𝑖𝑎𝑏𝑙𝑒 𝑇ℎ𝑒 𝑟𝑒𝑔𝑟𝑒𝑠𝑠𝑖𝑜𝑛 𝑙𝑖𝑛𝑒 𝑅𝑎𝑛𝑑𝑜𝑚 𝑣𝑎𝑟𝑖𝑎𝑏𝑙𝑒
5
The first component in the bracket is the part of Y explained by the changes in X and the second
is the part of Y not explained by X, that is to say the change in Y is due to the random influence
of 𝑢𝑖 . In the simple linear regression model, we typically refer to Y (Dependent Variable) as the:
Covariate
Additionally, α and β are unknown parameters, and the Ui is independent random variables with
zero mean and constant variance for all. The parameters α and β are called regression parameters
(or regression coefficients). The regression parameters are unknown, non-random parameters.
They are the intercept and the slope, respectively, of the straight-line relating Y to X.
Figure 2.2: The regression line and scatter plots b/n CGPA and studying hour
4
3.5
3
2.5
2
0 5 10 15 20
SH
6
The scatter of observations represents the true relationship between CGPA and studying hour. The
line represents the exact part of the relationship and the deviation of the observation from the line
represents the random component of the relationship.
The classicals assumed that the model should be linear in the parameters regardless of whether the
explanatory and the dependent variables are linear or not. A function is said to be linear in the
parameter, say, β, if β appears with a power of 1 only and is not multiplied or divided by any other
parameter. Example:
This means that the value which u may assume in any one period depends on chance; it may be
positive, negative or zero. Every value has a certain probability of being assumed by u in any
particular instance.
3. The mean value of the random variable(U) in any particular period is zero
This means that for each value of x, the random variable(u) may assume various values, some
greater than zero and some smaller than zero, but on average value equal to zero. In other words,
the positive and negative values of u cancel each other.
7
For all values of X, the u’s will show the same dispersion around their mean. In Fig.2.c this
assumption is denoted by the fact that the values that u can assume lie with in the same limits,
irrespective of the value of X. For 𝑋1, u can assume any value with in the range AB; for 𝑋2, u can
assume any value with in the range CD which is equal to AB and so on. Graphically;
Mathematically;
𝑉𝑎𝑟(𝑈𝑖 ) = 𝐸[𝑈𝑖 − 𝐸(𝑈𝑖 )]2 = 𝐸(𝑈𝑖 )2 = 𝜎 2 (Since𝐸(𝑈𝑖 ) = 0).
This constant variance is called homoscedasticity assumption and the constant variance itself is
called homoscedastic variance. Data that does not show constant error variance is called
heteroscedasticity and must be corrected by a data transformation or other methods.
𝑈𝑖 𝑁(0, 𝜎 2 ) ………………………………………...……2.4
That is, 𝑈𝑖 are normally distributed for all i . In conjunction with assumption 3 and 4, this implies
that u i are independently and normally distributed with mean zero and a common variance 𝜎 2 .
Values taken by the regressor X are considered fixed in repeated samples. More technically, X
is assumed to be non-stochastic.
8
7. The random terms of different observations (𝑼𝒊 , 𝑼𝒋 )are independent (The assumption of
no autocorrelation)
This means the value which the random term assumed in one period does not depend on the value
which it assumed in any other period. Algebraically,
Correlated error terms sometimes occur in time series data and is known as auto-correlation.
Autocorrelation is the degree of similarity between a given time series and a lagged version of
itself over successive time intervals. It's conceptually similar to the correlation between two
different time series, but autocorrelation uses the same time series twice: once in its original form
and once lagged one or more time periods.
This means there is no correlation between the random variable and the explanatory variable.
Correlation of the error terms with the independent variable is called endogeneity. Endogeneity
refers to situations in which an explanatory variable is correlated with the error term. If two
variables are unrelated their covariance is zero.
9
Variance: 𝑉𝑎𝑟(𝑌𝑖 ) = 𝛦(𝑌𝑖 − 𝛦(𝑌𝑖 ))2
= 𝛦(𝛼 + 𝛽𝑋𝑖 + 𝑢𝑖 − (𝛼 + 𝛽𝑋𝑖 ))2
= 𝛦(𝑢𝑖 )2
= 𝜎 2 (since 𝛦(𝑢𝑖 )2 = 𝜎 2 )
∴ 𝑣𝑎𝑟( 𝑌𝑖 ) = 𝜎 2 ………………………………………. 2.8
The shape of the distribution of 𝑌𝑖 is determined by the shape of the distribution of 𝑢𝑖 which is
normal by assumption 5. Since 𝛼 𝑎𝑛𝑑 𝛽, being constant, they don’t affect the distribution of 𝑦𝑖 .
Furthermore, the values of the explanatory variable, 𝑥𝑖 , are a set of fixed values by assumption 6
and therefore don’t affect the shape of the distribution of 𝑦𝑖 .
∴ 𝑌𝑖 ~N(𝛼 + 𝛽𝑥𝑖 , 𝜎 2 )
Specifying the model and stating its underlying assumptions are the first stage of any econometric
application. The next step is the estimation of the numerical values of the parameters of economic
relationships. The parameters of the simple linear regression model can be estimated by three of
the most commonly methods:
10
3. Method of moments (MM)
The model 𝑌𝑖 = 𝛼 + 𝛽𝑋𝑖 + 𝑈𝑖 is called the true relationship between Y and X because Y and X
represent their respective population value, and 𝛼 𝑎𝑛𝑑 𝛽 are called the true parameters since they
are estimated from the population value of Y and X. But it is difficult to obtain the population
value of Y and X because of technical or economic reasons. So, we are forced to take the sample
value of Y and X. The parameters estimated from the sample value of Y and X are called the
estimators of the true parameters 𝛼 𝑎𝑛𝑑 𝛽 and are symbolized as 𝛼̂ 𝑎𝑛𝑑 𝛽̂.
To find the values of 𝛼̂ 𝑎𝑛𝑑 𝛽̂ that minimize this sum, we have to partially differentiate ∑ 𝑒𝑖2 with
respect to 𝛼̂ 𝑎𝑛𝑑 𝛽̂ and set the partial derivatives equal to zero.
𝜕 ∑ 𝑒𝑖2
1. ̂
= −2 ∑(𝑌𝑖 − 𝛼̂ − 𝛽̂ 𝑋𝑖 ) = 0. . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . .2.12
𝜕𝛼
Note: at this point that the term in the parenthesis in equation 2.12 and 2.15 is the residual, 𝑒 =
𝑌𝑖 − 𝛼̂ − 𝛽̂ 𝑋𝑖 . Hence it is possible to rewrite them as −2 ∑ 𝑒𝑖 = 0 and −2 ∑ 𝑋𝑖 𝑒𝑖 = 0. It follows
that;
11
∑ 𝑒𝑖 = 0 and ∑ 𝑋𝑖 𝑒𝑖 = 0. . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . .2.16
If we rearrange equation 2.15 we obtain;
∑ 𝑌𝑖 𝑋𝑖 = 𝛼̂𝛴𝑋𝑖 + 𝛽̂ 𝛴𝑋𝑖2 ………………………………………………………………. 2.17
Equation (2.13) and (2.17) are called the Normal Equations. Substituting the values of 𝛼̂ from
(2.14) to (2.17), we get:
The expression in (2.21) to estimate the parameter coefficient is termed is the formula in deviation
form.
Example 2.1: Given the following sample data on daily consumption of an individual(Y) and his
income (X) for a sample of 10 individuals. Given the information below, answer question A-C
hereunder.
12
Observation Yi Xi
1 85 180
2 110 250
3 200 500
4 120 260
5 68 175
6 107 235
7 75 190
8 100 250
9 111 215
10 115 380
Solution
13
∑ 𝑋𝑖 𝑌𝑖 − 𝑛𝑋̅𝑌̅
𝛽=
∑ 𝑋𝑖 2 − 𝑛𝑋̅ 2
30,430
≅ 0.3263
93,252.5
Hence, 𝛽 = 0.3263
𝛼 ≅ 23.01995
̂𝑖 = 23.01995 + 0.3263𝑋̂𝑖
Thus, the fitted regression function is given by: 𝑌
The value of the intercept term (autonomous consumption) is 23.02, which implies that the value
of the dependent variable ‘consumption’ is 23.02 when the income of the individual is zero. That
is, the intercept measures the amount of consumption undertaken if income is zero. The intercept
is generally assumed and empirically documented to be positive this is because households
consume something even if their income is zero for survival. The regression slope or β is 0.3263
which is marginal propensity to consume that shows by how much consumption changes when
income changes by 1 Birr. From this regression we can infer that, when income increases by 1
Birr, on average consumption of a given household increases by 0.3263 Birr, ceteris paribus.
This means on average an individual whose income is 450 Birr will consume approximately 170
Birr and the remaining 280 Birr will be saved.
14
population parameters i.e. to vary within only a small range around the true parameter. How are
we to choose among the different econometric methods, the one that gives ‘good’ estimates? We
need some criteria for judging the ‘goodness’ of an estimate.
‘Closeness’ of the estimate to the population parameter is measured by the mean and variance or
standard deviation of the sampling distribution of the estimates of the different econometric
methods. We assume the usual process of repeated sampling i.e. we assume that we get a very
large number of samples each of size ‘n’; we compute the estimates 𝛽̂ ’s from each sample, and for
each econometric method and we form their distribution. We next compare the mean (expected
value) and the variances of these distributions and we choose among the alternative estimates the
one whose distribution is concentrated as close as possible around the population parameter.
The ideal or optimum properties that the OLS estimates possess may be summarized by well-
known theorem known as the Gauss-Markov Theorem.
Statement of the theorem: “Given the assumptions of the classical linear regression model, the
OLS estimators, in the class of linear and unbiased estimators, have the minimum variance, i.e.
the OLS estimators are BLUE.
Sometimes the theorem referred as the BLUE theorem i.e. Best, Linear, Unbiased Estimator. An
estimator is called BLUE if:
b. Unbiased: its average or expected value is equal to the true population parameter.
c. Minimum variance: It has a minimum variance in the class of linear and unbiased
estimators. An unbiased estimator with the least variance is known as an efficient
estimator.
According to the Gauss-Markov theorem, the OLS estimators possess all the BLUE properties.
The detailed proof of these properties is presented below
a. Linearity: (for ˆ )
15
Proof: From (2.21) of the OLS estimator of ˆ is given by:
(but xi = ( X − X ) = X − nX = nX − nX = 0 )
x Y xi
ˆ = i 2 ; Now, let = Ki (i = 1,2,..... n)
xi xi2
∴ 𝛽̂ = 𝛴𝐾𝑖 𝑌 − − − − − − − − − − − − − − − − − − − − − − − − − −(2.22)
̂ is linear in Y
̂)
Linearity: (for 𝜶
̂ is linear in Y.
Proposition: 𝜶
̂ is given by:
Proof: From (2.14) of the OLS estimator of 𝜶
∑𝑦
𝛼̂ = 𝑌̄ − 𝛽̂ 𝑋̄ = − 𝑋̄𝛴𝐾𝑖 𝑌
𝑛
𝛼̂ = 𝛴(1⁄𝑛 − 𝑋̄𝑘𝑖 )𝑌𝑖 ………………………….………...2.23
Hence 𝛼̂ is linear in Y.
b. Unbiasedness:
Proposition: ˆ & ˆ are the unbiased estimators of the true parameters &
To show ˆ & ˆ are the unbiased estimators of their respective parameters means to prove that:
= k i + k i X i + k i u i ,
but k i = 0 and k i X i = 1
xi ( X − X ) X − nX nX − nX
k i = = = = =0
xi
2
xi2
xi2 xi2
16
k i = 0 …………………………………………………………………………… (2.24)
xi X i ( X − X ) Xi
k i X i = =
xi2 xi2
X 2 − XX X 2 − nX 2
= = =1
X 2 − nX 2 X 2 − nX 2
k i X i = 1.......... .......... ......... …………………………………………………………(2.25)
ˆ = + k i u i ˆ − = k i u i − − − − − − − − − − − − − − − − − − − − − − − − − ( 2 .26 )
= + 1
n X i + 1n ui − Xki − Xki X i − Xki ui
= + 1 n ui − Xki ui , ˆ − = 1
n ui − Xki ui
𝛼̂ − 𝛼 = ( 1 n − Xk i )u i ………………………………..… (2.27)
(ˆ ) = − − − − − − − − − − − − − − − − − − − − − − − − − − − − − ( 2 . 28)
̂ is an unbiased estimator of .
c. Minimum variance of ˆ and ˆ
Now, we have to establish that out of the class of linear and unbiased estimators of and ,
ˆ and ˆ possess the smallest sampling variances. For this, we shall first obtain variance of
ˆ and ˆ and then establish that each has the minimum variance in comparison of the variances
of other linear and unbiased estimators obtained by any other econometric methods than OLS.
17
a. Variance of ˆ
var( ˆ ) = E ( k i u i ) 2
= [k12 u12 + k 22 u22 + .......... .. + k n2 un2 + 2k1k 2 u1u2 + ....... + 2k n−1k n un−1un ]
x i xi2 1
k i = 2 , and therefore, k i =
2
= 2
x i (xi )
2 2
xi
2
var( ˆ ) = 2 k i2 = ………………………………………………………...(2.30)
xi2
b. Variance of ̂
= (ˆ − )2 − − − − − − − − − − − − − − − − − − − − − − − − − − ( 2 .31 )
var(ˆ ) = (1 n − Xk i ) u i2
2
= (1 n − Xk i ) (ui ) 2
2
= 2 ( 1 n − Xki ) 2
= 2 ( 1 n2 − 2 n Xk i + X 2 k i2 )
= 2 ( 1 n + X 2 ki2 )
18
1 X2 xi2 1
= ( +
2
) , Since k i =
2
= 2
n xi 2
(xi )
2 2
xi
Again:
1 X 2 xi2 + nX 2 X 2
+ = =
n xi2 nxi2 nxi
2
X2 X i2
var(ˆ ) = 2 1 n + 2 = 2 ………………………………………… (2.32)
xi nxi
2
We have computed the variances OLS estimators. Now, it is time to check whether these variances
of OLS estimators do possess minimum variance property compared to the variance’s other
estimators of the true and , other than ˆ and ˆ .
To establish that ˆ and ˆ possess minimum variance property, we compare their variances with
that of the variances of some other alternative linear and unbiased estimators of and , say
* and * . Now, we want to prove that any other linear and unbiased estimator of the true
population parameter obtained from any other econometric method has larger variance that that
OLS estimators.
1. Minimum variance of ˆ
But, wi = k i + ci
𝛴𝑤𝑖 = 𝛴(𝑘𝑖 + 𝑐𝑖 ) = 𝛴𝑘𝑖 + 𝛴𝑐𝑖
19
Therefore, ci = 0 since k i = wi = 0
To prove whether ˆ has minimum variance or not let’s compute var( *) to compare with
var( ˆ )
var( *) = var( wi Yi )
= wi var(Yi )
2
ci xi
wi2 = ki2 + ci2 Since k i ci = =0
xi2
ˆ = ( 1n − Xki )Yi
By analogy with that the proof of the minimum variance property of ˆ , let’s use the weights wi
= ci + ki Consequently;
20
* = ( 1n − Xwi )Yi
Since we want * to be on unbiased estimator of the true , that is, ( *) = , we substitute
for Y = + xi + u i in * and find the expected value of * .
* = ( 1n − Xwi )( + X i + ui )
X ui
= ( + + − Xwi − XX i wi − Xwi u i )
n n n
i.e., if wi = 0, and wi X i = 1 . These conditions imply that ci = 0 and ci X i = 0 .
= 2 ( n n 2 + X 2 wi − 2 X wi )
2 1
n
var( *) = 2 (
1
n + X 2 wi
2
) , Since wi = 0
but wi = k i2 + ci2
2
var( *) = 2 (
1
n + X 2 (ki2 + ci2 )
1 X2
var( *) = 2 + + 2 X 2 ci2
n x i
2
X i2
= 2 + 2 X 2 ci2
nxi
2
The first term in the bracket it var(ˆ ) , hence
var(*) = var(ˆ ) + 2 X 2 ci2 ………………………………………………2.35
⇒ 𝑣𝑎𝑟( 𝛼 ∗) > 𝑣𝑎𝑟( 𝛼̂) , Since 𝜎 2 𝑋̄ 2 𝛴𝑐𝑖2 > 0
Therefore, we have proved that the least square estimators of linear regression model are best,
linear and unbiased (BLUE) estimators.
21
2.4. Covariance Vs. correlation
Dear students, what is covariance? What is correlation? Elaborate also the difference and
similarity among covariance, correlation and regression analysis, if any?
Covariance is a statistical tool that is used to determine the relationship between the movements
of two random variables. Covariance measures the direction of the relationship between two
variables. Specifically, covariance measures the total variation of two random variables from their
expected values. Using covariance, we can only gauge the direction of the relationship (whether
the variables tend to move in tandem or show an inverse relationship). However, it does not
indicate the strength of the relationship, nor the dependency between the variables. This
relationship is determined by the sign (positive or negative) of the covariance value, which is
calculated by:
A positive covariance between two variables indicates that these variables tend to be higher or
lower at the same time. In other words, a positive covariance between variables X and Y indicates
that X is higher than average at the same times that Y is higher than average, and vice versa. When
charted on a two-dimensional graph, the data points will tend to slope upwards. As depicted in
figure 2.3 below, the household monthly expenditure and family size are positively related.
Figure 2.3: positive covariance between family size and household expenditure
15000
household expenditure
5000 0 10000
0 5 10
family size
22
A Negative Covariance is when the calculated covariance is less than zero, this indicates that the
two variables have an inverse relationship. In other words, an X value that is lower than average
tends to be paired with a Y that is greater than average, and vice versa. That is, the two variables
tend to move in inverse directions. For instance, as the temperature decreases (becomes colder),
the sale of heaters will increase. This example is shown in graphical representation by the
downward sloped relationship between temperature in degree Celsius (X) and the sale of heaters
by firm in a day (Y).
Figure 2.4: negative covariance between temperature and the sale of heaters
40
30
20
10
0
10 15 20 25
Temperature in Co
In finance, the concept is primarily used in portfolio theory. One of its most common applications
in portfolio theory is the diversification method, using the covariance between assets in a portfolio.
By choosing assets that do not exhibit a high positive covariance with each other, the unsystematic
risk can be partially eliminated. While the covariance does measure the directional relationship
between two assets, it does not show the strength of the relationship between the two assets; the
correlation is a more appropriate indicator of this strength.
Correlation is a statistical measure that expresses the extent to which two variables are linearly
related (meaning they change together at a constant rate). Correlation is a statistical measure that
indicates the extent to which two or more variables fluctuate in relation to each other. In other
words, correlation refers to the degree to which a pair of variables are linearly related. Familiar
examples of correlation include the correlation between the height of parents and their offspring,
23
and the correlation between the price of a good and the quantity the consumers are willing to
purchase, as it is depicted in the so-called demand curve. It's a common tool for describing simple
relationships without making a statement about cause and effect i.e. correlation does not
necessarily imply causation. On the basis of the nature of relation between the variables,
correlation may be categorized as follow:
Simple correlation is a measure used to determine the strength and the direction of the relationship
between two variables, X and Y. A simple correlation coefficient can range from –1 to 1. However,
maximum (or minimum) values of some simple correlations cannot reach unity (i.e., 1 or –1). The
farther the value is from 0, the greater the strength of the relationship, which will be practically
perfect when it reaches -1 or 1. At 0, which is the null value, in theory there will be no correlation
between the two variables.
Example 2.2: A manger may wish to see whether the number of years the sales person has been
working for the company has anything to do with the amount of sales of the representatives. The
only two variables are: years of experience and amount of sales
In multiple correlation three or more variables are studied simultaneously. For instance,
performance of students in university exam may depend upon his/her IQ, mother’s qualification,
father’s qualification, parent’s income, number of hours of studies, etc. Whenever we are
interested in studying the joint effect of two or more variables on a single variable, multiple
correlation gives the solution of our problem. In fact, multiple correlation is the study of
combined influence of two or more variables on a single variable.
Example 2.3: A manager may wish to study the extent of correlation among demand of his output,
preference of consumer, price of his product and price of his competitors. This type of study
involves several variables: demand of the product, preference of consumer, price of his product
24
and price of his competitors. Hence, the extent of relationship among those variables could be
studied under multiple correlation.
A correlation is also said to be partial if it studies the degree of relationship between two
variables keeping all other variables connected with these two are constant. Partial correlation
is a process in which we measure of the strength and also direction of a linear relationship
between two continuous variables while controlling for the effect of one or more other
continuous variables it is called ‘covariates’ and also ‘control’ variables.
Example 2.4: demand for a product is dependent on factors such as price of the product, price of
related products, income of the consumer, taste and preference, expectation of future price etc. to
name a few. However, in the analysis of the law of demand, we only analyze the relationship
between price of the product and quantity demand of it, by assuming other factors are held
constant. This analysis is called partial correlation analysis.
B) Positive and Negative Correlation
Two variables may have a positive correlation, negative correlation, or they may be uncorrelated.
Two variables are said to be positively correlated if they tend to change together in the same
direction, that is, if they tend to increase or decrease together. Such positive correlation is
postulated by economic theory for the quantity of a commodity supplied and its price. When the
price increases the quantity supplied increases and when price falls the quantity supplied decreases,
ceteris paribus. In contrast to this, negative correlation exists when two variables tend to change
in the opposite direction: when X increases Y decreases, and vice versa. For example, saving and
household size are negatively correlated. When price increases, demand for the commodity
decreases and when price falls demand increases, ceteris paribus.
C) Linear and Non-Linear Correlation
If the amount of change in one variable is accompanied by the same amount of change in the other
variable, it is known as linear correlation. In other words, Linear correlation is when the ratio of
proportion of two given variables are same/constant. Example- every time when the income
increases by 20% there is a rise in expenditure of 5%.
Example 2.5:
25
If the amount of change in one variable is not accompanied by the same amount of change in the
other variable, the correlation is said to be non-linear. In other words, non-linear correlation is
when the ratio of variations between two given variables changes. Example- with the 20% increase
in the income the expenditure sometimes increases by 5% sometimes by 10%.
Example 2.6:
The variable X: 10, 12, 30, 70, 74
The variable Y: 40, 46, 50, 75, 90
Methods of Measuring Correlation
There are three methods of measuring correlation. These are:
1. The Scattered Diagram or Graphic Method
2. The Simple Linear Correlation coefficient
3. The coefficient of Rank Correlation
1. The Scattered Diagram or Graphic Method
The scatter diagram is a rectangular diagram which can help us in visualizing the relationship
between two phenomena. The closer the data points come together and make a straight line, the
higher the correlation between the two variables, or the stronger the relationship. If the data points
make a straight line going from the origin out to high x- and y-values, then the variables are said
to have a positive correlation. If the line goes from a high-value on the y-axis down to a high-value
on the x-axis, the variables have a negative correlation.
26
Figure 2.5: Perfect linear correlations
A perfect positive correlation is given the value of 1. A perfect negative correlation is given the
value of -1. If there is absolutely no correlation present the value given is 0. The closer the number
is to 1 or -1, the stronger the correlation, or the stronger the relationship between the variables.
The closer the number is to 0, the weaker the correlation.
2. The Population Correlation Coefficient ‘’ and its Sample Estimate ‘r’
In the light of the above discussions it appears clear that we can determine the kind of correlation
between two variables by direct observation of the scatter diagram. In addition, the scatter diagram
indicates the strength of the relationship between the two variables. This section is about how to
determine the type and degree of correlation using a numerical result. For a precise quantitative
measurement of the degree of correlation between Y and X we use a parameter which is called the
correlation coefficient and is usually designated by the Greek letter ‘’. Having as subscripts the
variables whose correlation it measures, refers to the correlation of all the values of the
population of X and Y. Its estimate from any particular sample (the sample statistic for correlation)
is denoted by ‘r’ with the relevant subscripts. For example, if we measure the correlation between
X and Y the population correlation coefficient is represented by xy and its sample counterpart by
rxy. The simple correlation coefficient is used to measure relationships which are simple and linear
only. It cannot help us in measuring non-linear as well as multiple correlations. Sample correlation
coefficient is defined by the formula:
𝑛 ∑ 𝑋𝑖 𝑌𝑖 −∑ 𝑋𝑖 ∑ 𝑌𝑖
𝑟𝑥𝑦 = ………………………………………………….2.37
√𝑛 ∑ 𝑋𝑖 2 −(∑ 𝑋𝑖 )2 √𝑛 ∑ 𝑌𝑖 2 −(∑ 𝑌𝑖 )2
Where, xi = X i − X and y i = Yi - Y
We will use a simple example from the theory of supply. Economic theory suggests that the
quantity of a commodity supplied in the market depends on its price, ceteris paribus. When price
increases the quantity supplied increases, and vice versa. When the market price falls producers
offer smaller quantities of their commodity for sale.
27
Example 2.7: The following table shows the quantity supplied for a commodity with the
corresponding price values. Determine the type of correlation that exists between these two
variables.
Time period (in days) Quantity supplied Yi (in tons) Price Xi (in shillings)
1 10 2
2 20 4
3 50 6
4 40 8
5 50 10
6 60 12
7 80 14
8 90 16
9 90 18
10 120 20
To estimate the correlation coefficient we, compute the following results.
Y X xi = X i − X yi = Yi − Y x2 y2 xiyi XY X2 Y2
10 2 -9 -51 81 2601 459 20 4 100
20 4 -7 -41 49 1681 287 80 16 400
50 6 -5 -11 25 121 55 300 36 2500
40 8 -3 -21 9 441 63 320 64 1600
50 10 -1 -11 1 121 11 500 100 2500
60 12 1 -1 1 1 -1 720 144 3600
80 14 3 19 9 361 57 1120 196 6400
90 16 5 29 25 841 145 1440 256 8100
90 18 7 29 49 841 203 1620 324 8100
120 20 9 59 81 3481 531 2400 400 14400
Sum=610 110 0 0 330 10490 1810 8520 1540 47700
Mean=61 11
n XY − X Y
10(8520) − (110)(610)
r= = = 0.975
10(1540) − (110)(110) 10(47700) − (610)(610)
n X − ( X )
2 2
n Y − ( Y )
2 2
Or using the deviation form (Equation 2.38), the correlation coefficient can be computed as:
1810
r= = 0.975
330 10490
28
This result shows that there is a strong positive correlation between the quantity supplied and the
price of the commodity under consideration.
The simple correlation coefficient has the value always ranging between -1 and +1. That means
the value of correlation coefficient cannot be less than -1 and cannot be greater than +1. Its
minimum value is -1 and its maximum value is +1. If r= -1, there is perfect negative correlation
between the variables. If 0 r +1 , there is positive correlation between the two variables and
movement from zero to positive one increases the degree of positive correlation. If r= +1, there is
perfect positive correlation between the two variables. If the correlation coefficient is zero, it
indicates that there is no linear relationship between the two variables.
Though, correlation coefficient is most popular in applied statistics and econometrics, it has its
own limitations. The major limitations of the method are:
1. The correlation coefficient always assumes linear relationship regardless of the fact whether
the assumption is true or not.
2. Great care must be exercised in interpreting the value of this coefficient as very often the
coefficient is misinterpreted. For example, high correlation between lung cancer and
smoking does not show us smoking causes lung cancer.
3. The value of the coefficient is unduly affected by the extreme values
29
4. The coefficient requires the quantitative measurement of both variables. If one of the two
variables is not quantitatively measured, the coefficient cannot be computed.
6 ∑ 𝐷2
𝑟 = 1 − 𝑛(𝑛2 −1)………………………………… ………………………………2.39
Where,
n = number of observations.
Two points are of interest when applying the rank correlation coefficient. Firstly, it does not matter
whether we rank the observations in ascending or descending order. However, we must use the
same rule of ranking for both variables. Second, if two (or more) observations have the same value
we assign to them the mean rank. Let’s use example to illustrate the application of the rank
correlation coefficient.
Example 2.8: A market researcher asks experts to express their preference for twelve different
brands of soap. Their replies are shown in the following table.
30
Brands of soap A B C D E F G H I J K L
Person I 9 10 4 1 8 11 3 2 5 7 12 6
Person II 7 8 3 1 10 12 2 6 5 4 11 9
The figures in this table are ranks but not quantities. We have to use the rank correlation coefficient
to determine the type of association between the preferences of the two persons. This can be done
as follows.
6 D 2
6(50)
r' = 1 − = 1− = 0.827
n(n − 1)
2
12(122 − 1)
This figure, 0.827, shows a marked similarity of preferences of the two persons for the various
brands of soap.
2.5. Statistical test of Significance of the OLS Estimators (First Order tests)
After the estimation of the parameters and the determination of the least square regression line, we
need to know how ‘good’ is the fit of this line to the sample observation of Y and X, that is to say
we need to measure the dispersion of observations around the regression line. This knowledge is
essential because the closer the observation to the line, the better the goodness of fit, i.e. the better
is the explanation of the variations of Y by the changes in the explanatory variables.
We divide the available criteria into three groups: the theoretical a priori criteria, the statistical
criteria, and the econometric criteria. Under this section, our focus is on statistical criteria (first
order tests). The two most commonly used first order tests in econometric analysis are:
i. The coefficient of determination (the square of the correlation coefficient i.e. R2): This test
is used for judging the explanatory power of the independent variable(s).
31
ii. Tests of statistical significance: This test is used for judging the statistical reliability of the
estimates of the regression coefficients.
This is the measure of dispersion of observations around the regression line. The closer the
observations to the regression line the better the goodness of fit. It is also the square of the
correlation coefficient, which shows the percentage of the total variation of the dependent variable
(Y) that can be explained by the independent variable, (X). To elaborate this let’s draw a horizontal
line corresponding to the mean value of the dependent variable 𝑌̄. By fitting the line 𝑌̂ = 𝛼̂0 +
𝛽̂1 𝑋 we try to obtain the explanation of the variation of the dependent variable Y produced by the
changes of the explanatory variable X.
Yi
(1) ui = Yi − Y Yi = 0 + 1 X i
−
(3) y i = Yi − Y
−
(2) y i = Yi − Y
− − −
(Y ) (X ,Y)
−
(X ) Xi
As can be seen from figure 2.6 above, 𝑌𝑖 − 𝑌̄represents measures the variation of the sample
observation value of the dependent variable around the mean. However, the variation in Y that
can be attributed the influence of X, (i.e. the regression line) is given by the vertical distance 𝑌̂ −
𝑌̄. The part of the total variation in Y about 𝑌̄ that can’t be attributed to X is equal to 𝑌 − 𝑌̂ which
is referred to as the residual variation/unexplained variation.
In summary:
32
𝑦𝑖 = 𝑌 − 𝑌̄= deviation of Y from its mean.
𝑦̂ = 𝑌̂ − 𝑌̄= deviation of the regressed (predicted) value (𝑌̂) from the mean.
Now, we may write the observed Y as the sum of the predicted value (𝑌̂) and the residual term
(ei.).
𝑌
⏟𝑖 = ⏟̂
𝑌 + 𝑒⏟𝑖
𝑂𝑏𝑠𝑒𝑟𝑣𝑒𝑑 𝑌𝑖 𝑝𝑟𝑒𝑑𝑖𝑐𝑡𝑒𝑑 𝑌𝑖 𝑟𝑒𝑠𝑖𝑑𝑢𝑎𝑙
From above equation we can have the deviation form 𝑦 = 𝑦̂ + 𝑒. By squaring and summing both
sides, we obtain the following expression:
𝛴𝑦 2 = 𝛴(𝑦̂ + 𝑒)2
𝛴𝑦 2 = 𝛴(𝑦̂ 2 + 𝑒𝑖2 + 2𝑦𝑒𝑖)
= 𝛴𝑦𝑖 2 + 𝛴𝑒𝑖2 + 2𝛴𝑦̂𝑒𝑖
But 𝛴𝑦̂𝑒𝑖 =𝛴𝑒(𝑌̂ − 𝑌̄) = 𝛴𝑒(𝛼̂ + 𝛽̂ 𝑥𝑖 − 𝑌̄)
= 𝛼̂𝛴𝑒𝑖 + 𝛽̂ 𝛴𝑒𝑥𝑖 − 𝑌̂𝛴𝑒𝑖
(but 𝛴𝑒𝑖 = 0, 𝛴𝑒𝑥𝑖 = 0)
⇒ ∑ 𝑦̂𝑒 = 0……………………………………………..……… (2.40)
Therefore;
2 2
𝛴𝑦
⏟𝑖 = ⏟̂2
𝛴𝑦 + 𝛴𝑒
⏟𝑖 ………………………………... (2.35)
𝑇𝑜𝑡𝑎𝑙 𝐸𝑥𝑝𝑙𝑎𝑖𝑛𝑒𝑑 𝑈𝑛 𝑒𝑥𝑝 𝑙𝑎𝑖𝑛𝑒𝑑
𝑣𝑎𝑟𝑖𝑎𝑡𝑖𝑜𝑛 𝑣𝑎𝑟𝑖𝑎𝑡𝑖𝑜𝑛 𝑣𝑎𝑟𝑖𝑎𝑡𝑖𝑜𝑛
OR,
𝑇𝑜𝑡𝑎𝑙 𝑠𝑢𝑚 𝑜𝑓 𝐸𝑥𝑝𝑙𝑎𝑖𝑛𝑒𝑑 𝑠𝑢𝑚 𝑅𝑒 𝑠 𝑖𝑑𝑢𝑎𝑙 𝑠𝑢𝑚
= +
⏟ 𝑠𝑞𝑢𝑎𝑟𝑒 ⏟ 𝑜𝑓𝑠𝑞𝑢𝑎𝑟𝑒 ⏟ 𝑜𝑓𝑠𝑞𝑢𝑎𝑟𝑒
𝑇𝑆𝑆 𝐸𝑆𝑆 𝑅𝑆𝑆
i.e. 𝑇𝑆𝑆 = 𝐸𝑆𝑆 + 𝑅𝑆𝑆 ……………………………………..………. (2.41)
Mathematically; the explained variation as a percentage of the total variation (r2) is explained as:
𝐸𝑆𝑆 𝛴𝑦̂ 2
𝑟 2 = 𝑇𝑆𝑆 = 𝛴𝑦 2 ……………………………………………………..………. (2.42)
33
We can substitute (2.43) in (2.42) and obtain:
̂ 2 𝛴𝑥 2
𝛽
𝑟 2 = 𝐸𝑆𝑆/𝑇𝑆𝑆 = …………………………………………………… 2.44
𝛴𝑦 2
We can also express r2 in other form by substituting 𝛽̂ in equation 2.44 above as follows:
𝛴𝑥𝑦 2 𝛴𝑥𝑖2 𝛴𝑥 𝑦
= (𝛴𝑥 2 ) , Since𝛽̂ = 𝛴𝑥𝑖 2 𝑖
𝛴𝑦 2 𝑖
34
A) Compute the Explained, unexplained, and total sum of squares
Obser. Yi Xi ̅
𝒀−𝒀 ̅ )𝟐
𝑻𝑺𝑺 = (𝒀 − 𝒀 ̅
𝑿−𝑿 ̅ )𝟐
(𝑿 − 𝑿
1 85 180 -24 576 -83.5 6972.25
2 110 250 1 1 -13.5 182.25
3 200 500 91 8,291 236.5 55932.25
4 120 260 11 121 -3.5 12.25
5 68 175 -41 1,681 -88.5 7832.25
6 107 235 -2 4 -28.5 812.25
7 75 190 -34 1,156 -73.5 5402.25
8 100 250 -9 81 -13.5 182.25
9 110 215 1 1 -48.5 2352.25
10 115 380 6 36 116.5 13572.25
Sum ΣYi=1090 ΣXi=2,635 ∑(𝑌 − 𝑌̅)2 =11,948 ∑(𝑋 − 𝑋̅)2 =93,252.5
Mean 109 263.5
As shown in the table above the total sum of square (TSS) equals to ∑(𝑌 − 𝑌̅)2 =11,948 and this
is the total variation of Y from its mean. The explained sum of square (ESS) becomes:
𝐸𝑆𝑆 = 𝛴𝑦̂ 2 = 𝛽̂ 2 𝛴𝑥 2 , we have calculated the value of 𝛽 in example 2.1 and obtained 𝛽 = 0.3263,
hence:
𝑅𝑆𝑆 = 2019.2487
35
is, out of 100 percent variation in consumption, 83.1 percent will be explained by the change in
income of an individual and the rest 16.9 percent will be explained by factors other than income
of consumer which are in the residual (i.e. Consumer’s expectations, the rate of interest, tastes and
preferences, the stock of wealth etc.).
To test the significance of the OLS parameter estimators we need the following:
The OLS estimates 𝛼̂ 𝑎𝑛𝑑 𝛽̂ are obtained from a sample of observations on Y and X. Since
sampling errors are inevitable in all estimates, it is necessary to apply test of significance in order
to measure the size of the error and determine the degree of confidence in order to measure the
validity of these estimates. This can be done by using various tests. The most common ones are:
i) Standard error test iii) Confidence interval
The standard error is an indispensable tool in the kit of a researcher and econometrician, because
it is used in testing the validity of statistical hypothesis. The standard deviation of the sampling
distribution of a statistic is called the standard error. The standard error is important in dealing
with statistics (measures of samples) which are normally distributed. Here the use of the word
“error” is justified in this connection by the fact that we usually regard the expected value to be
true value and the divergence from it as error of estimation due to sampling fluctuations.
This test helps us decide whether the estimates 𝛼̂ 𝑎𝑛𝑑 𝛽̂ are significantly different from zero, i.e.
whether the sample from which they have been estimated might have come from a population
whose true parameters are zero. 𝛼 = 0 𝑎𝑛𝑑/𝑜𝑟 𝛽 = 0. In other words, the standard error may
36
fairly be taken to measure the reliability/unreliability of the sample estimate. The greater the
standard error the greater the difference between observed and expected values and greater the
unreliability of the sample estimate. On the other hand, the smaller value of the standard error, the
smaller the difference between observed and expected values and greater the reliability of the
sample estimate. The reciprocal of the standard error may be regarded as a measure of reliability.
The reliability of an observed proportion varies as the square root of the number of observations
on which it is based. The larger the sample size, the smaller the standard error and the smaller the
sample size, the larger the standard error.
Formally, in standard error and other tests of significance, we test the null hypothesis 𝑯𝟎 : 𝜷𝒊 = 𝟎
against the alternative hypothesis 𝑯𝟏 : 𝜷𝒊 ≠ 𝟎. Null hypothesis (often symbolized H0 and read as
“H-naught”) states that there is no relationship in the population and that the relationship in the
sample reflects only sampling error. That is, the sample relationship “occurred by chance.” On the
other hand, alternative hypothesis (often symbolized as H1) states that there is a relationship in the
population and that the relationship in the sample reflects this relationship in the population. The
standard error test may be outlined as follows.
∑ 𝑒𝑖 2
∑𝑋 2
̂ 2 𝛴𝑋 2
𝜎 ∑ 𝑒𝑖 .∑ 𝑋 2 2
𝑆𝐸(𝛼̂) = √𝑣𝑎𝑟( 𝛼̂) = √ 𝑛𝛴𝑥 2 = √ 𝑛−2∑ 2 or √(𝑛 ∑ 𝑥𝑖 2 )(𝑛−2)……………….2.48
𝑛 𝑥
Second: compare the standard errors with the numerical values of 𝛼̂𝑎𝑛𝑑𝛽̂ .
Decision rule:
• If 𝑆𝐸(𝛽̂𝑖 ) > 1⁄2 𝛽̂𝑖 , accept the null hypothesis and reject the alternative hypothesis. We
37
Economic interpretation of standard error test
If we have 𝑌𝑖 = 𝛼̂ + 𝛽̂ 𝑥𝑖 + 𝑒𝑖
Case A) Acceptance of the null hypothesis for regression slope (β) and rejecting the null for
intercept (𝛼)
Accept Ho: 𝛽=0 and reject the alternative that H1 = 𝛽 ≠ 0, means the estimation are statistically
insignificant or
- We accept the null hypothesis that the true parameter is zero/ independent variable is
insignificant.
- The independent variable does not influence the dependent variables (no relationship
between the dependent & independent variables)
On the other hand, if we reject the null hypothesis H0: 𝛼 =0 & accept the alternative 𝐻1 : 𝛼 ≠ 0
means 𝛼 is statistically significant or the equation will have intercept or the intercept is differed
from zero. That is, the inclusion of an intercept term in regression model is meaningful.
Then the above equation will be 𝑌𝑖 = 𝛼̂ and geometrically/ graphically can be shown as:
𝑌 = 𝛼̂𝑖
Case B) Acceptance of the null hypothesis for intercept 𝛼̂ and rejecting for regression slope (β)
If we accept the null hypothesis H0: 𝛼 =0 & reject the alternative 𝐻1 : 𝛼 ≠ 0 means 𝛼 is statistically
insignificant or the equation will not have intercept or the intercept is not differed from zero. On
37
the other hand, when we reject the null hypothesis 𝐻0 : 𝛽̂ = 0 & accept the alternative that 𝐻1 : 𝛽̂ ≠
0 means:
𝑌𝑖 = 𝛽̂ 𝑥𝑖
X
Case C) Rejecting of the null hypothesis for both regression slope (β) and intercept (𝛼):
In this case both the regression slope and intercept are statistically different from zero/significant.
That is, the independent variables will influence the dependent variable or there is a relationship
between the dependent & independent variables. Beside this, the regression line do have an
intercept that is statistically different from zero.
𝑌 = 𝛼 + 𝛽𝑋
X
Case D) when we accept the null hypothesis for both regression slope and intercept
This is when we fail to reject the null hypothesis that states the sample statistics comes from a
population whose parameters are zero and there is no relationship between the dependent &
60
independent variables. That is, both the regression slope and intercept do not have statistical
significance or the change in independent variable leaves the dependent variable unchanged.
Like the standard error test, this test is also important to test the significance of the parameters.
Student’s t-test is a method of testing hypotheses about the mean of a small sample drawn from a
normally distributed population when the population variance is unknown. In 1908 William Sealy
Gosset, developed the t-test and t distribution. The t distribution is a family of curves in which the
number of degrees of freedom (the number of independent observations in the sample minus one)
specifies a particular curve. As the sample size (and thus the degrees of freedom) increases, the t
distribution approaches the bell shape of the standard normal distribution. In practice, for tests
involving the mean of a sample of size greater than 30, the normal distribution is usually applied.
From statistics, any variable X can be transformed into t using the general formula:
𝑋−𝜇
𝑡= , with n-1 degree of freedom.
𝑠𝑥
Where:
SE = is standard error
k = number of parameters in the model.
Since we have two parameters in simple linear regression with intercept different from zero, our
degree of freedom is n-2. Like the standard error test, we formally test the hypothesis: 𝑯𝟎 : 𝜷𝒊 = 𝟎
against the alternative 𝑯𝟏 : 𝜷𝒊 ≠ 𝟎 for the slope parameter; and 𝑯𝟎 : 𝜶 = 𝟎 against the alternative
𝑯𝟏 : 𝜶 ≠ 𝟎 for the intercept.
38
Step 1: Compute t*, which is called the computed value of t, by taking the value of 𝛽 in the null
hypothesis. In our case 𝛽 = 0, then t* becomes:
̂ −0
𝛽 ̂
𝛽
𝑡 ∗= 𝑆𝐸(𝛽̂) = 𝑆𝐸(𝛽̂)…………………………………………………………….2.49
Step 2: Choose level of significance. Level of significance is the probability of making ‘wrong’
decision, i.e. the probability of rejecting the hypothesis when it is actually true or the probability
of committing a type I error. It is real that, in making decision one can never be 100% sure that
one will make the right decision; because we are liable to commit one of the following types of
errors:
We choose a level of significance for deciding whether to accept or reject our hypothesis. It is
customary in econometrics to choose the 5% or 1% level of significance. This means that in
making our decision we allow (tolerate) five times of a hundred to be “wrong”, that is to reject the
probability when it is actually true. Since we want the probability of being wrong to be small, we
assign a lower value to this value and we call it the level of significance of our test.
Step 3: Choose the location of the critical region (Check whether there is one tail or two tail tests)
The critical region may be chosen either at the right end (right tail) of the distribution of the
variable or at the left tail or half at each end of the distribution. The decision of whether to choose
a one tail or a two-tail critical region depends on the form in which the alternative hypothesis is
expressed. If the inequality sign in the alternative hypothesis is ≠, then it implies a two-tail test
and divide the chosen level of significance by two; decide the critical region or critical value of t
called tc. But if the inequality sign is either > or < then it indicates one tail test and there is no need
to divide the chosen level of significance by two to obtain the critical value of tc from the t-table.
Example:
If we have 𝐻0 : 𝛽𝑖 = 0
against: 𝐻1 : 𝛽𝑖 ≠ 0
39
Then this is a two-tail test. If the level of significance is 5%, divide it by two to obtain critical
value of t from the t-table.
Note that, in econometrics it has become customary to perform a two-tail test. The choice of two-
tail implies no a priori knowledge regarding the sign of the coefficient whose significance is being
tested. However, a one-tail test would be appropriate for variables with a priori expectation
regarding the sign of the coefficients of economic relationship.
Step 4: Obtain critical value of t, (tc) at 𝛼⁄2and n-2 degree of freedom for two tail tests. For
example, if we choose the 5% level of significance ( ) , each tail will include the area (probability)
0.025 ( / 2) .
• If t*> tc, reject H0 and accept H1. The conclusion is 𝛽̂ 𝑜𝑟 𝛼 is statistically significant.
• If t*< tc, accept H0 and reject H1. The conclusion is 𝛽̂ 𝑜𝑟 𝛼 is statistically insignificant.
Rejection of the null hypothesis doesn’t mean that our estimate 𝛼̂ 𝑎𝑛𝑑 𝛽̂ is the correct estimate of
the true population parameter 𝛼 𝑎𝑛𝑑 𝛽. It simply means that our estimate comes from a sample
drawn from a population whose parameter 𝛽 is different from zero. In order to define how close,
the estimate to the true parameter, we must construct confidence interval for the true parameter. In
other words, we must establish limiting values around the estimate with in which the true parameter
is expected to lie within a certain “degree of confidence”. A confidence interval is a range of
values that is likely to contain a population parameter with a certain level of confidence. It is
written as:
40
In a two-tail test at level of significance, the probability of obtaining the specific t-value either
–tc or tc is 𝛼⁄2at n-2 degree of freedom. The probability of obtaining any value of t which is equal
̂ −𝛽
𝛽
to 𝑆𝐸(𝛽̂) at n-2 degree of freedom is 1 − (𝛼⁄2 + 𝛼⁄2)𝑖. 𝑒. 1 − 𝛼.
The limit within which the true 𝛽 lies at (1 − 𝛼)% degree of confidence is:
̂ − 𝑺𝑬(𝜷
[𝜷 ̂ )𝒕𝒄 , 𝜷
̂ + 𝑺𝑬(𝜷
̂ )𝒕𝒄 ];………………………………………………….………..2.53
where 𝑡𝑐 is the critical value of t at 𝛼⁄2confidence interval and n-2 degree of freedom.
𝐻0 : 𝛽 = 0
𝐻1 : 𝛽 ≠ 0
Decision rule: If the hypothesized value of 𝛽 in the null hypothesis is within the confidence
interval, accept H0 and reject H1. The implication is that 𝛽̂ is statistically insignificant; while if
the hypothesized value of 𝛽 in the null hypothesis is outside the limit, reject H0 and accept H1.
This indicates 𝛽̂ is statistically significant.
i. In the long run 95 out of 100 cases will contain the true parameter i in the limit or
intervals. OR
41
ii. We are 95% confident that the unknown population parameters ( i ) will lie within the
limit /interval. OR
iii. In 5% of the cases the population parameter will lie outside the confidence limits.
Analysis of variance (ANOVA) uses F-tests to statistically assess the overall (joint) significance
of parameters. The F-Test of overall significance in regression is a test of whether or not your
linear regression model provides a better fit to a dataset than a model with no predictor variables.
The F * ratio (the observed variance ratio) is obtained by dividing the two ‘mean square errors’
that is the MSE of the ESS (between the means) and RSS (within the sample). The steps involved
in F-test includes:
H0: β=0 (The model with no predictor variables (also known as an intercept-only model) fits the
data as well as the regression model), which is checked against
H1: β≠0 (Your regression model fits the data better than the intercept-only model)
ESS is explained sum of square and RSS is residual (unexplained) sum of square
Degree of freedom for nominator will be (df1) = k-1(the number of parameters to be estimated less
1) and Degree of freedom for denominator is (df2) = n-k-1(sample size less the parameters to be
estimated less 1).
42
Step 4: obtain F-critical from F-distribution at a specified significance level (customarily we use
1%,5% or at most 10%) at k-1 degree of freedom for numerator and n-k degree of freedom for
denominator.
Step 5: Compare the F statistic obtained in Step 2 with the critical value obtained in Step 4. If the
F statistic is greater than the critical value at the required level of significance, we reject the null
hypothesis. If the F statistic obtained in Step 2 is less than the critical value at the required level
of significance, we cannot reject the null hypothesis.
Example 2.9: A researcher is using data from a sample of 30 firms to investigate the relationship
between sales of the firm Yi (measured in Birr) and expenditure on promotion of their product Xi
(measured in Birrs per year). Preliminary analysis of the sample data produces the following
sample information:
Where xi = X i − X and yi = Yi − Y and yˆ i = Yˆi − Y for i = 1, ..., n. Use the above sample
𝛴𝑥𝑖 𝑦𝑖 367.18
𝛽̂ = = ≅ 0.94
𝛴𝑥𝑖2 390.67
Hence, 𝛽 = 0.94
And to calculate the value of α estimate, we have to first calculate mean of X and Y
∑ 𝑌𝑖 1,034.97
𝑌̅ = = ≅ 34.5
𝑛 30
∑ 𝑋𝑖 528.6
𝑋̅ = = ≅ 17.62
𝑛 30
Therefore, α becomes: 𝛼 = 𝑌̅ − 𝛽𝑋̅ = 34.5 − 0.94(17.62) = 34.5 − 16.56
𝛼 ≅ 17.94
̂𝑖 = 17.94 + 0.94𝑋̂𝑖
Thus, the fitted regression function is given by: 𝑌
43
(b) Interpret 𝛼 and 𝛽
Interpretation of β: the value of regression slope is 0.94 which implies when expenditure on
advertise increases by 1 Birr then, on average the sales of firms increases by 0.94 Birr, ceteris
paribus.
Interpretation of α: the value intercept is 17.94 which implies when advertisement and all other
factors that affect sales are valued at zero (constant), then firms’ sales becomes 17.94 Birr.
(c) What is the part of the variation in sales which is not explained by the regression?
Solution: to find residual sum of square, we have to find total sum of square and explained sum
of square. That is,
TSS= 𝛴𝑦 2 which is given as: ∑𝑛𝑖=1 𝑦𝑖 2 = 1256.24 and the explained sum of square can be
calculated as:
ESS= 𝛴𝑦̂ 2 = 𝛽̂ 2 𝛴𝑥 2 =(0.94)2 (390.67) ≅345.2
ESS=345.2
2 𝑛
ESS= 𝛴𝑦̂ = 𝛽 ∑𝑖=1 𝑥𝑖 𝑦𝑖 = (0.94)(367.18) ≅345.2
ESS=345.2
Therefore, the residual sum of square can easily obtain as:
RSS = 𝛴𝑒 2 = 911.04
(d) Calculate the coefficient of determination for the data and interpret its value?
2
𝐸𝑆𝑆 𝛽̂ 2 𝛴𝑥 2
𝑟 = =
𝑇𝑆𝑆 𝛴𝑦 2
345.2
𝑟2 =
1256.24
𝑟 2 ≅ 0.275 = 27.5%
Interpretation of 𝑟 2 : 𝑟 2 = 0.275, this means that the variation in sales can be explained by
advertisement by 27.5%. The remaining 62.3% of the total variation in sales is attributed to the
factors included in the disturbance variable 𝑢𝑖 .
44
𝛴𝑒 2 𝛴𝑋 2
𝑣𝑎𝑟( 𝛼̂) =
𝑛 − 2 𝑛𝛴𝑥 2
911.04 11,432.92 10,415,847.4368
𝑣𝑎𝑟( 𝛼̂) = =
30 − 2 30(390.67) 328,162.8
𝑣𝑎𝑟( 𝛼̂) ≅
𝜎̂ 2
𝑣𝑎𝑟( 𝛽̂ ) =
𝛴𝑥 2
𝛴𝑒 2
𝑣𝑎𝑟( 𝛽̂ ) =
𝑛 − 2𝛴𝑥 2
911.04 911.04
𝑣𝑎𝑟( 𝛽̂ ) = =
(30 − 2)(390.67) 10,938.76
𝑣𝑎𝑟( 𝛽̂ ) = 0.0833
(f) Formulate hypothesis and test the significance of the estimates using standard error test?
The null hypothesis 𝑯𝟎 : 𝛃 = 𝟎 against the alternative hypothesis 𝑯𝟏 : 𝛃 ≠ 𝟎, then we test this
hypothesis using those steps:
𝑆𝐸(𝛽̂ ) = 0.28862
To be significant (to reject the null hypothesis) 𝑆𝐸(𝛽̂𝑖 ) < 1⁄2 𝛽̂𝑖 hence,
Therefore, we reject the null hypothesis which implies the estimates are statistically significant. In
other words, the expenditure on advertisement will influence the sales or there is a relationship
between the dependent & independent variables. This also checked for 𝛼 or intercept as follows:
The null hypothesis 𝑯𝟎 : 𝛂 = 𝟎 against the alternative hypothesis 𝑯𝟏 : 𝛂 ≠ 𝟎, then we test this
45
hypothesis using those steps:
𝑆𝐸(𝛼̂) = √31.74
𝑆𝐸(𝛼̂) = 5.6338
Step 2: compare the standard errors with the numerical values of 𝛼̂.
To be significant (to reject the null hypothesis) 𝑆𝐸(𝛼̂) < 1⁄2 𝛼̂ hence,
Hence, we reject the null hypothesis H0: 𝛼 =0 & accept the alternative 𝐻1 : 𝛼 ≠ 0 means 𝛼 is
statistically significant or the equation will have intercept or the intercept is differed from zero.
𝛽̂ 0.94
𝑡 ∗= = = 3.2569
𝑆𝐸(𝛽̂ ) 0.28862
Step 2: Choose level of significance. For this example, we are asked to test the hypothesis at 5%.
Step 3: Check whether there is one tail test or two tail tests. In the alternative hypothesis the sign
is ≠, then it implies a two-tail test and we divide the chosen level of significance by two (hence
we check at 0.025).
Step 4: Obtain critical value of t, (tc) at 5⁄2 %and n-2(30-2) degree of freedom for two tail tests.
𝑡𝑐 = 2.048
If t*> tc, reject H0 and accept H1. Here 3.2569 > 2.048, hence we can conclude that 𝛽̂ is
statistically significant.
46
Step 1: calculate t-calculated
𝛼̂ 17.94
𝑡 ∗= = = 3.1843
𝑆𝐸(𝛼̂) 5.6338
Step 2: choosing the level of significance (5%)
Step 3: similarly, the alternative hypothesis for intercept is also two tailed test.
Step 4: Obtain critical value of t, (tc) at 5⁄2 % and 30-2 degree of freedom for two tail tests.
𝑡𝑐 = 2.048
[0.3489, 1.5311]
Since, the value of 𝛽 is within the confidence interval, we reject H0 and accept H1. The implication
is that 𝛽̂ is statistically significant.
Similarly, this test can also be undertaken for intercept term as follows:
̂ − 𝑺𝑬(𝜶
[𝜶 ̂ )𝒕𝒄 , 𝜶
̂ + 𝑺𝑬(𝜶
̂ )𝒕𝒄 ];
[6.402 , 29.478]
Since, the value of 𝛼 is within the confidence interval, we reject H0 and accept H1. The implication
is that 𝛼 is statistically significant.
(I) Test to determine if there is evidence of regression at the 0.01 level (F-test)?
H0: β=0 (The model with no predictor variables fits the regression model),
47
H1: β≠0 (the regression model fits the data better than the intercept-only model)
345.2⁄2 − 1 345.2
𝐹∗ = =
911.04⁄30 − 2 − 1 33.742
𝐹 ∗ ≅ 10.231
Step 3: Calculate the degrees of freedom.
Degree of freedom for nominator= (df1) = k-1= 2-1=1and
Degree of freedom for denominator is (df2) = 30-2-1=27
Step 4: obtain F-critical from F-distribution at 1% level of significance and 1 df1 & 27 df2
𝑭𝑪 = 𝟕. 𝟔𝟕𝟕
48
OLS estimators possess the minimum variance property because they are linear and unbiased estimators that also satisfy the Gauss-Markov theorem. This means that for any unbiased estimator, the variance of the OLS estimator is less than or equal to the variance of any alternative linear unbiased estimator. Therefore, OLS provides optimal precision in estimation, reducing the standard errors associated with coefficient estimates .
The classical linear regression model ensures accurate predictions by adhering to several key assumptions: linearity, independence, homoscedasticity, and normal distribution of errors. When these assumptions are met, the model can provide the best linear unbiased estimates. The method of least squares used in OLS further minimizes the sum of squared deviations, ensuring that the predicted values are as close to the actual dependent variable values as possible .
The Gauss-Markov theorem is significant because it states that in a linear regression model, as long as the assumptions of the classical linear regression are met, the ordinary least squares (OLS) estimator yields the best linear unbiased estimates (BLUE) of the coefficients. This implies that among all linear and unbiased estimators, the OLS estimator has the smallest variance, making it theoretically optimal for linear regression .
The primary difference is that in a deterministic relationship, for each value of X, there is one and only one corresponding value of Y, expressed as Y = f(X). In contrast, a stochastic relationship implies that for each value of X, there is a probability distribution for Y, signifying multiple potential outcomes for Y at each X .
One method used for measuring correlation is the scatter diagram or graphic method. This method visually represents data relationships by plotting data points on a rectangular coordinate system. If the points are close to forming a straight line, there is a strong linear correlation. A line from the origin to the high values suggests a positive correlation, while a downward slope implies a negative correlation. The scatter diagram thus helps in visualizing both the type and strength of correlation .
The method of least squares in regression analysis minimizes the sum of the squared residuals or deviations, ensuring the line of best fit is as close to all the data points as possible. This contributes to the accuracy of the estimates by reducing the overall distance between the observed and predicted values, thus improving the model's predictive power and precision .
The sample correlation coefficient, represented as 'r', estimates the population correlation coefficient 'ρ', which measures the strength and direction of a linear relationship in the entire population. The sample coefficient is calculated from the sample data and used to infer the population's correlation. While 'r' provides a practical measure, it is subject to sampling variability, thereby necessitating careful statistical inference to generalize results to the population .
Covariance measures the direction of the linear relationship between two variables, indicating whether they tend to increase or decrease together. Correlation, on the other hand, measures both the strength and direction of this relationship, providing a dimensionless index ranging from -1 to 1. A strong correlation value indicates a stronger linear relationship. Covariance can be any value between negative and positive infinity, making correlation a more intuitive measure for the strength of a relationship .
Partial correlation measures the strength and direction of a linear relationship between two continuous variables while controlling for the effect of one or more other variables, called covariates. This allows researchers to isolate the specific relationship between two variables, factoring out the influence of other correlated variables, leading to a clearer understanding of the direct connection between the primary variables of interest .
Testing hypotheses in regression models allows researchers to determine the statistical significance of the predictors, ensuring that observed relationships are not due to random chance. Performing confidence interval analyses provides a range within which we can expect the true parameter value to fall, enabling more robust decision-making. Both are critical for validating model results and making informed predictions or inferences based on the model .