Econometrics Module Part I 2018
Econometrics Module Part I 2018
MODULE PART I
March, 2026
Debre Tabor, Ethiopia
Chapter 1: Introduction
Econometrics is about how we can use economic or social science theory and data, along with
tools from statistics to answer “how much” type questions. It integrates mathematical knowledge,
statistical skills and economic theories to solve business and economic problems of agribusiness
firms. For instance, economics tells us that the demand for a good is a function of the good‟s price,
and that in most cases; the price elasticity of demand is negative. But for many practical purposes
one may be interested to quantify the elasticity more accurately. For such kind of questions
econometrics can provide the answer.
1.1. Definition and scope of econometrics
What is Econometrics?
Simply stated Econometric means economic measurement. The “metric” part of the word
signifies measurement and econometrics is concerned with the measuring of economic
relationships.
It is a social science in which the tools of economic theory, mathematics and statistical
inference are applied to the analysis of economic phenomena (Arthur Goldberger).
In the words of Maddala econometrics is “the application of statistical and mathematical
methods to the analysis of economic data, with a purpose of giving empirical content to
economic theories and verifying them or refuting them.”
Econometrics utilizes economic theory, as embodied in an econometric Model; facts, as
summarized by relevant data, and statistical theory, as refined into econometric techniques to
measure and to test empirically certain relationships among economic variables.
It is a special type of economic analysis and research in which the general economic
theory formulated in mathematical form (i.e. mathematical economics) is combined
with empirical measurement (i.e. statistics) of economic phenomena.
Why a Separate Discipline?
A model is any representation of an actual phenomenon such as an actual system or process. The
real-world system is represented by the model in order to explain it, to predict it, and to control it.
Any model represents a compromise between reality and manageability.
A given representation of real-world system can be a model if it fulfills the following requirements.
(1) It must be a “reasonable” representation of the real-world system and in that sense it
should be realistic.
(2) On the other hand, it must be “manageable” in that it yields certain insights or
conclusions.
A good model is both realistic and manageable. A highly realistic but too complicated model is a
“bad” model in the sense it is not manageable. A model that is highly manageable but so idealized
that it is unrealistic not accounting for important components of the real-world system, is a “bad”
model too.
In general, to find the proper balance between realism and manageability is the essence of good
Modeling. Thus, a good model should, on the one hand, specify the interrelationship among the
parts of a system in a way that is sufficiently detailed and explicit and, on the other hand, it should
be sufficiently simplified and manageable to ensure that the model can be readily analyzed and
conclusions can be reached concerning the real world.
Economic models
Any economic theory is an observation from the real world. For one reason, the immense
complexity of the real-world economy makes it impossible for us to understand all
interrelationships at once. Another reason is that all the interrelationships are not equally important
for the understanding of the economic phenomenon under study. The sensible procedure is
therefore, to pick up the important factors and relationships relevant to our problem and to focus
our attention on these alone. Such a deliberately simplified analytical framework is called on
economic model. It is an organized set of relationships that describes the functioning of an
economic entity under a set of simplifying assumptions. All economic reasoning is ultimately
based on models. Economic models consist of the following three basic structural elements.
1. A set of variables
Econometric models
The most important characteristic of economic relationships is that they contain a random element
which is ignored by mathematical economic models which postulate exact relationships between
economic variables.
Example: Economic theory postulates that the demand for a commodity depends on its price, on
the prices of other related commodities, on consumers‟ income and on tastes. This is an exact
relationship which can be written mathematically as:
𝑄 = 𝑏0𝑃 + 𝑏1𝑃 + 𝑏2𝑃𝑂 + 𝑏3 𝑡
The above demand equation is exact. However, many more factors may affect demand. In
econometrics the influence of these „other‟ factors is taken into account by the introducing random
variable. In our example, the demand function studied with the tools of econometrics would be of
the stochastic form:
𝑄 = 𝑏0𝑃 + 𝑏1𝑃 + 𝑏2 𝑃𝑂 + 𝑏3𝑡 + 𝜇 where 𝜇 stands for the random factors which affect the
quantity demanded. The random term (also called error term or disturbance term) is a surrogate
variable for important variables excluded from the model, errors committed and measurement
errors.
Desirable Properties of an Econometric Model
An econometric model is a model whose parameters have been estimated with some appropriate
econometric technique. The „goodness‟ of an econometric model is judged customarily based on
the following desirable properties.
1. Theoretical Plausibility: The model should be compatible with the postulates of
economic theory and adequately describe the economic phenomena to which it relates.
2. Explanatory ability: The model should be able to explain the observations of the actual
world. It must be consistent with the observed behavior of the economic variables whose
relationship it determines.
3. Accuracy of the estimates of the parameter: The estimates of the coefficients should be
accurate in the sense that they should approximate as best as possible the true parameters
of the structural model. The estimates should, if possible, possess the desirable properties
of unbiasedness, consistency and efficiency.
4. Forecasting ability: The model should produce satisfactory predictions of future values
of the dependent (endogenous) variables.
5. Simplicity: The model should represent the economic relationships with maximum
simplicity. The fewer the equations and the simpler their mathematical form, the better the
model provided that the other desirable properties are not affected by the simplifications
of the model.
1.2. Goals of Econometrics
relationships and with the predication of the values of economic variables. The relationships of
economic theory which can be measured with econometric techniques are relationships in which
some variables are postulated as causes of the variation of other variables. Starting with the
postulated theoretical relationships among economic variables, econometric research or inquiry
generally proceeds along the following lines/stages.
1. Statement of theory or hypothesis.
6. Hypothesis testing
7. Forecasting or prediction
To illustrate the preceding steps, let us consider the well-known Keynesian theory of consumption.
Keynes stated: “Consumption increases as income increases, but not as much as the increase in
income”. It means that “The marginal propensity to consume (MPC) for a unit change in income
is greater than zero but less than unit”
2. Specification of the mathematical model of the theory
Although Keynes postulated a positive relationship between consumption and income, he did not
specify the precise form of the functional relationship between the two. For simplicity, a
mathematical economist might suggest the following form of the Keynesian consumption function:
𝑌 = 𝛽1 + 𝛽2𝑋; 0 < 𝛽2 < 1 .................................................................................... (1.3.1)
Where Y = consumption expenditure and X = income, and where 𝛽1 and 𝛽2, known as the
parameters of the model, are, respectively, the intercept and slope coefficients.
The slope coefficient 𝛽2 measures the MPC. This equation, which states that consumption is
linearly related to income, is an example of a mathematical model of the relationship between
consumption and income that is called the consumption function in economics. A model is simply
a set of mathematical equations. If the model has only one equation, as in the preceding example,
it is called a single-equation model, whereas if it has more than one equation, it is known as a
multiple-equation model.
In Eq. (1.3.1) the variable appearing on the left side of the equality sign is called the dependent
variable and the variable(s) on the right side are called the independent, or explanatory,
variable(s). Thus, in the Keynesian consumption function, Eq. (1.3.1), consumption (expenditure)
is the dependent variable and income is the explanatory variable.
3. Specification of the econometric model of the theory
The purely mathematical model of the consumption function given in Eq. (1.3.1) is of limited
interest to the econometrician, for it assumes that there is an exact or deterministic relationship
between consumption and income. But relationships between economic variables are generally
inexact. For example, in the above example, in addition to income, other variables affect
consumption expenditure. Such as, size of family, ages of the members in the family, family
religion, etc., are likely to exert some influence on consumption. To allow for the inexact
relationships between economic variables, the econometrician would modify the deterministic
consumption function (1.3.1) as follows:
𝑌 = 𝛽1 + 𝛽2𝑋 + 𝜇 ......................................................................................................... (1.3.2)
Where 𝜇, known as the disturbance, or error term, is a random (stochastic) variable. The
disturbance term u may well represent all those factors that affect consumption but are not taken
into account explicitly. Equation (1.3.2) is an example of an econometric model. More
technically, it is an example of a linear regression model. The econometric consumption function
hypothesizes that the dependent variable Y (consumption) is linearly related to the explanatory
variable X (income) but that the relationship between the two is not exact; it is subject to individual
variation.
4. Obtaining Data
To estimate the econometric model given in (1.3.2), that is, to obtain the numerical values of 𝛽1,
we need data. Have a look at the data given in the table below!
Table 1.1: Twelve years consumption expenditure (Y) and income of (X)
Note that the statistical technique of regression analysis is the main tool used to obtain the
estimates. Using this technique and the data given in Table 1.1, we obtain the following estimates
of β1 and β2, namely, - 231.8 and 0.7194. Thus, the estimated consumption function is:
The hat on the Y indicates that it is an estimate. MPC was about 0.72 and it means that for the
sample period when real income increases by 1 USD, led (on average) real consumption
expenditure increases of about 72 cents.
Note: A hat symbol (^) above one variable will signify an estimator of the relevant population
value.
6. Hypothesis Testing
Are the estimates accords with the expectations of the theory that is being tested? Is MPC < 1
statistically? If so, it may support Keynes‟ theory. Confirmation or refutation of economic theories
based on sample evidence is object of Statistical Inference (hypothesis testing).
7. Forecasting or Prediction
Suppose we have the estimated consumption function given in (1.3.3). Suppose further the
government believes that consumer expenditure of about 4000 will keep the unemployment rate
at its current level of about 4.2 percent. What level of income will guarantee the target amount of
consumption expenditure?
𝑌 𝑋 → X =5882
Give MPC = 0.72, an income of $5882 Bill will produce an expenditure of $4000 Bill. By fiscal
and monetary policy, Government can manipulate the control variable X to get the desired level
of target variable Y.
1.4. Elements of Econometrics
1. Estimation
How big change in one variable tends to be associated with a unit change in another?
2. Testing hypothesis
sample from the population. If sample data are not consistent with the statistical hypothesis,
the hypothesis is rejected.
There are two types of statistical hypotheses.
i. Null hypothesis. The null hypothesis, denoted by H0, is usually the hypothesis that
sample observations result purely from chance.
ii. Alternative hypothesis. The alternative hypothesis, denoted by H1 or Ha, is the
hypothesis that sample observations are influenced by some non-random cause.
State the hypotheses. This involves stating the null and alternative hypotheses. The
hypotheses are stated in such a way that they are mutually exclusive. That is, if one is true,
the other must be false.
Formulate an analysis plan. The analysis plan describes how to use sample data to evaluate
the null hypothesis. The evaluation often focuses around a single test statistic.
Analyze sample data. Find the value of the test statistic (mean score, proportion, t statistic,
z-score, etc.) described in the analysis plan.
Interpret results. Apply the decision rule described in the analysis plan. If the value of the
test statistic is unlikely, based on the null hypothesis, reject the null hypothesis.
A forecast can be defined as a statement about an unknown and uncertain event most often,
but not necessarily, a future event. Such a statement may vary greatly in form and content: it
can be qualitative or quantitative, conditional or unconditional, explicit or silent on the
probabilities involved. Economic forecasts refer to the economic aspects of unknown events
How many more students would reed need to admit in order to fill its class if tuition
were $1000 higher?
What will happen to interest rates next year if the economy recovers?
3. Data
Data is an input for an econometric analysis. Econometric methods depend on the nature of the
data used. Use of inappropriate methods may lead to misleading results. There are two sources of
data.
i. Primary sources of data: Data collected from direct respondents using formal and
informal survey. Formal surveys are interviewing respondents using questionnaire.
Informal surveys are key informant interview and focus group discussion
ii. Secondary sources of data: Data from previous published and unpublished materials
from the internet and different offices.
Types of data
1. Continuous data
Continuous data can take on any value and are not confined to take specific numbers. For example,
the rental yield on a property could be 6.2%, 6.24%, or 6.238%.
2. Discrete data
Discrete data can only take on certain values, which are usually count numbers. For instance, the
number of adult family members in a given family.
Different kinds of economic data sets
1. Cross-sectional data
Sample of individuals, households, firms, cities, states, countries, or other units of interest at a
given point of time/in a given period. Cross-sectional observations are more or less independent.
For example, pure random sampling from a population. Sometimes pure random sampling is
violated, e.g. units refuse to respond in surveys, or if sampling is characterized by clustering.
Cross-sectional data typically encountered in applied microeconomics.
Table 1.1: A cross sectional data set on wage and other individual characteristics
Observations of a variable or several variables over time. e.g. stock prices, money supply,
consumer price index, gross domestic product, annual homicide rates, automobile sales, …
Ordering of observations conveys important information. Data frequency: daily, weekly, monthly,
quarterly, annually, etc. Typical features of time series: trends and seasonality. Typical
applications: applied macroeconomics and finance.
Table1.3:Minimum wage ,unemployment, and related data for Addis Ababa sub city
Obser no Year Avegmin avgcov prunemp Pr GNP
1 1950 0.20 20.1 15.4 878.7
2 1951 0.21 20.7 16.0 925.0
3 1952 0.23 22.6 14.8 1015.9
. . . . . .
. . . . . .
37 1986 3.35 58.1 18.9 4281.6
38 1987 3.35 58.2 16.8 4496.7
39 1988 4.25 62.5 22.4 5214.6
Two or more cross sections are combined in one data set. Cross sections are drawn independently
of each other. Pooled cross sections often used to evaluate policy changes. Example:
Evaluate effect of change in property taxes on house prices
. . . . . . .
. . . . . . .
. . . . . . .
The same cross-sectional units are followed over time. Panel data have a cross-sectional and a time
series dimension. Panel data can be used to account for time-invariant unobservable. Panel data
can be used to model lagged responses. Example
o City crime statistics; each city is observed in two years
Table 1.5: a two year panel data set on city crime statistics
Obs no City year murders population unem police
1 1 1986 5 350000 8.7 440
2 1 1990 8 359200 7.2 471
3 2 1986 2 64300 5.4 75
4 2 1990 1 65100 5.5 75
. . . . . . .
. . . . . . .
297 14 1986 10 260700 9.6 286
298 14 1990 6 245000 9.8 334
299 20 1986 25 543000 4.3 520
300 20 1990 32 546200 5.2 493
Econometrics may be divided into two broad categories: theoretical econometrics and applied
econometrics. In each category, one can approach the subject in the classical or Bayesian
tradition. Theoretical econometrics is concerned with the development of appropriate methods for
measuring economic relationships specified by econometric models. In this aspect, econometrics
leans heavily on mathematical statistics. For example, one of the methods used extensively in
econometrics is least squares. Theoretical econometrics must spell out the assumptions of this
method, its properties, and what happens to these properties when one or more of the assumptions
of the method are not fulfilled.
Econometric methods may be classified in to two groups: (1) single-equation techniques, which
are methods that are applied to one relationship at a time; and simultaneous-equation techniques,
which are methods applied to all the relationships of a model simultaneously.
In applied econometrics we use the tools of theoretical econometrics to study some special field(s)
of economics and business, such as the production function, investment function, demand and
supply functions, etc.
Correlation is the relationship between more than one variable is considered as correlation.
Correlation is considered as a number which can be used to describe the relationship between two
variables. Correlation is a statistical measure that indicates the extent to which two or more
variables fluctuate together. The Degree and type of relationship between any two or more
quantities (variables) in which they vary together over a period; for example, variation in the level
of expenditure or savings with variation in the level of income. A positive correlation exists where
the high values of one variable are associated with the high values of the other variable(s). A
negative correlation means association of high values of one with the low values of the other(s).
Correlation can vary from +1 to -1. Values close to +1 indicate a high-degree of positive
correlation, and values close to -1 indicate a high degree of negative correlation. Values close to
zero indicate poor correlation of either kind, and 0 indicates no correlation at all. While correlation
is useful in discovering possible connections between variables, it does not prove or disprove any
cause-and-effect or causal relationships between them. See also regression. Simple correlation is
defined as a variation related amongst any two variables.
The multiple correlation and partial correlation are categorized as related variation among three
or more variables. Two variables are correlated only when they vary in such a way that the higher
and lower values of one variable corresponds to the higher and lower values of the other variable.
We might also get to know if they are correlated when the higher value of one variable corresponds
with the lower value of the other.
It is a statistical method which enables the researcher to find whether two variables are related and
to what extent they are related. Correlation is considered as the sympathetic movement of two or
more variables. We can observe this when a change in one particular variable is accompanied by
changes in other variables as well, and this happens either in the same or opposite direction, then
the resultant variables are said to be correlated. Considering a data where we find two or more
variables getting valued then we might study the related variation for these variables.
Correlation a mutual relationship or connection, the process of correlating two or more things.
The extent to which two variables are interdependent. Unlike regression, this calculation is not
used to predict the value of one variable from the other. It is Statistics interdependence of variable
quantities.
2.2. Coefficient of Linear Correlation
direction of a linear relationship between two variables. The linear correlation coefficient
is sometimes referred to as the Pearson product moment correlation coefficient in honor
of its developer Karl Pearson.
The mathematical formula for computing r is:
𝑛 ∑ 𝑋𝑌 ∑𝑋∑𝑌 ∑𝑥 𝑦
𝑟
√ 𝑛∑𝑋 ∑𝑋 𝑛∑𝑌 ∑𝑌
√ ∑𝑥 ∑𝑦
𝑤 𝑟 𝑥 𝑋 𝑋̅ 𝑛 𝑦 𝑌 𝑌̅
An increase in one variable may cause an increase in the other variable, or a decrease in one
variable may cause decrease in the other variable. When the variables move in the same direction
like this they are said to be positively correlated. The positive correlation may be termed as direct
correlation. If a decrease in one variable causes an increase in the other variable or vice versa, the
variables are said to be negatively correlated. The negative correlation may be termed as inverse
correlation. In case the two variables are not at all related they are said to be independent or
uncorrelated.
The value of r is such that -1 < r < +1. The + and – signs are used for positive
linear correlations and negative linear correlations, respectively.
Positive correlation: If x and y have a strong positive linear correlation, r is close to
+1. r value of exactly +1 indicates a perfect positive fit. Positive values indicate a
relationship between x and y variables such that as values for x increase, values
for y also increase.
16 | Prepared by: Bishaw A. (MSc) Email Address: bish200821@[Link]
Debre Tabor University, CAES, Department of Agricultural Economics
A perfect correlation of ± 1 occurs only when the data points all lie exactly on a straight
line.
If r = +1, the slope of this line is positive. If r = -1, the slope of this line is negative.
A correlation greater than 0.8 is generally described as strong, whereas a correlation less
than
0.5 is generally described as weak. These values can vary based upon the "type" of data
being examined. A study utilizing scientific data may require a stronger correlation
than a study using social science data.
Properties of simple correlation coefficient
Coefficient of correlation lies between –1≤ r ≤1
If r = -1 or +1 indicate that there is perfect negative (inverse) or positive (direct) linear Relationship
between two variables respectively.
A coefficient of correlation(r) that is closes to zero shows the relationship is quite weak,
Whereas r is closest to +1 or -1, shows that the relationship is strong.
Note that;
The strength of correlation does not depend on the positiveness and negativeness of r
The correlation between two variables is linear if a unit changes in one variable result in a
Constant change in the other variable. Correlation can be studied through plotting scattered
Diagrams.
Example: If the following data is given for you as data collected from a given market on quantity
and price find the simple correlation coefficient and discuss the type of relationship between the
variables. Based on your knowledge of economics and simple correlation coefficient, what are the
variables expressed as quantity?
Q P 𝑄 𝑃 𝑄𝑃 𝑞 𝑝 𝑞 𝑝 𝑞𝑝
10 2 100 4 20 -51 -9 2601 81 459
20 4 400 16 80 -41 -7 1681 49 287
50 6 2500 36 300 -11 -5 121 25 55
40 8 1600 64 320 -21 -3 441 9 63
50 10 2500 100 500 -11 -1 121 1 11
60 12 3600 144 720 -1 1 1 1 -1
80 14 6400 196 1120 19 3 361 9 57
90 16 8100 256 1440 29 5 841 25 145
90 18 8100 324 1620 29 7 841 49 203
120 20 14400 400 2400 59 9 3481 81 531
Sum=610 110 47700 1540 8520 0 0 10490 330 1810
coefficient. We define
r12.3 = partial correlation coefficient between Y and X1, holding X2 constant
r13.2 = partial correlation coefficient between Y and X2, holding X1 constant
r23.1 = partial correlation coefficient between X1 and X2, holding Y constant
These partial correlations can be easily obtained from the simple or zero order, correlation coefficients
as follows.
𝑟 𝑟 𝑟
𝑟
√ 𝑟 𝑟
𝑟 𝑟 𝑟
𝑟
√ 𝑟 𝑟
𝑟 𝑟 𝑟
𝑟
√ 𝑟 𝑟
The partial correlations given in Equations (2.2) to (2.4) are called first order correlation coefficients.
By order we mean the number of secondary subscripts. Thus r1 2.3 4 would be the correlation coefficient
of order two, r1 2.3 4 5 would be the correlation coefficient of order three, and so on. As noted previously,
r12, r13, and so on are called simple or zero-order correlations. The interpretation of, say, r 1 2.3 4 is that it
gives the coefficient of correlation between Y and X1, holding X2 and X3 constant.
Interpretation of Simple and Partial Correlation Coefficients
In the two-variable case, the simple r had a straightforward meaning: It measured the degree of (linear)
association (and not causation) between the dependent variable Y and the single explanatory variable X.
But once we go beyond the two-variable case, we need to pay careful attention to the interpretation of
the simple correlation coefficient. From (4.4.18), for example, we observe the following:
1. Even if r12 = 0, r12.3 will not be zero unless r13 or r23 or both are zero.
2. If r12 = 0 and r13 and r23 are nonzero and are of the same sign, r 1 2.3 will be negative, whereas if they
are of the opposite signs, it will be positive. An example will make this point clear. Let Y = crop yield,
X1 = rainfall, and X2 = temperature. Assume r12 = 0, that is, no association between crop yield
and rainfall. Assume further that r 13 is positive and r23 is negative. Then, as (2.2) shows, r1 2.3 will be
positive; that is, holding temperature constant, there is a positive association between yield and rainfall.
This seemingly paradoxical result, however, is not surprising. Since temperature X2 affects
both yield Y and rainfall X1, in order to find out the net relationship between crop yield and rainfall, we
need to remove the influence of the “nuisance” variable temperature. This example shows how one
might be misled by the simple coefficient of correlation.
3. The terms r12.3 and r12 (and similar comparisons) need not have the same sign.
4. In the two-variable case we have seen that 𝑟 lies between 0 and 1. The same property holds true of
the squared partial correlation coefficients. Using this fact, the reader should verify that one can obtain
the following expression from (4.4.18):
This gives the interrelationships among the three zero-order correlation coefficients. Similar
expressions can be derived from Equations (2.3) and (2.4).
5. Suppose that r13 = r23 = 0. Does this mean that r12 is also zero? The answer is obvious from (2.5).
The fact that Y and X2 and X1 and X2 are uncorrelated does not mean that Y and X1 are uncorrelated.
In passing, note that the expression r21 2 .3 may be called the coefficient of partial determination and
may be interpreted as the proportion of the variation in Y not explained by the variable X2 that has been
explained by the inclusion of X1 into the model . Conceptually it is similar to 𝑅
Multiple correlation coefficients
Before moving on, note the following relationships between multiple regression coefficient (R2),
simple correlation coefficients, and partial correlation coefficients:
𝑅
𝑅
In concluding this section, consider the following: It was stated previously that 𝑅 will not decrease if
an additional explanatory variable is introduced into the model, which can be seen clearly from (2.7).
This equation states that the proportion of the variation in Y explained by X1 and X2 jointly is the sum
of two parts: the part explained by X1 alone ( 𝑟 ) and the part not explained by X2 ( 𝑟 )
times the proportion that is explained by X2 after holding the influence of X1 constant. Now 𝑅
𝑟 so long as 𝑟 . At worst, 𝑟 will be zero, in which case 𝑅 = 𝑟
Rank Correlation
Sometimes we come across statistical series in which the variables under consideration are not
capable of quantitative measurement, but can be arranged in serial order. This happens when we
dealing with qualitative characteristics (attributes) such as beauty, efficient, honest, intelligence
…. etc. in such case one may rank the different items and apply the spearman method of rank
difference for finding out the degree of relationship. The greatest use of this method (rank
correlation) lies in the fact that one could use it to find correlation of qualitative variables, but
since the method reduces the amount of labor of calculation, it is sometimes used also where
quantitative data is available. It is used when statistical series are ranked according to their
magnitude and the exact size of individual item is not known. Spearman‟s correlation coefficient
is denoted by r'. Steps of r'
i. Rank the different items in X and Y.
Example: A market researcher asks two experts to express the preferences for 12 different brands of
Soap.
Brand of soap X Y Di Di2
A 9 7 2 4
B 10 8 2 4
C 4 3 1 1
D 1 1 0 0
E 8 10 -2 4
F 11 12 -1 1
G 3 2 -1 1
H 2 6 -4 16
I 5 5 0 0
J 7 4 3 9
K 12 11 1 1
L 6 9 -3 9
Sum 50
𝑟
𝑛 𝑛
∑
𝑟
REVIEW QUESTION
In the literature the terms dependent variable and explanatory variable are described variously. A
representative list is:
Dependent Explanatory
variable variable (s)
Independent
Explained variable(s)
variable
Predictor(s)
Predictand
Regressor(s)
Regressand
If we are studying the dependence of a variable on only a single explanatory variable, such as that
As noted in Section 2.1.1, regression analysis is largely concerned with estimating and/or
predicting the (population) mean value of the dependent variable on the basis of the known or
fixed values of the explanatory variable(s). To understand this, consider the data given on 2.1.
The data in the table refer to a total population of 60 families in a hypothetical community and
their weekly income (X) and weekly consumption expenditure (Y), both in dollars. The 60 families
are divided into 10 income groups (from $80 to $260) and the weekly expenditures of each family
in the various groups are as shown in the table. Therefore, we have 10 fixed values of X and the
corresponding Y values against each of the X values.
There is considerable variation in weekly consumption expenditure in each income group, which
can be seen clearly from Figure 2.1. But the general picture that one gets is that, despite the
variability of weekly consumption expenditure within each income bracket, on the average,
weekly consumption expenditure increases as income increases.
Fig 3.1.3. Conditional distribution of expenditure for various level of income (data of table 2.1)
To see this clearly, in Table 2.1 we have given the mean, or average, weekly consumption
expenditure corresponding to each of the 10 levels of income. Thus, corresponding to the weekly
income level of $80, the mean consumption expenditure is $65, while corresponding to the income
level of $200, it is $137. In all we have 10 mean values for the 10 subpopulations of Y. We call
these mean values conditional expected values, as they depend on the given values of the
(conditioning) variable X. Symbolically, we denote them as E(Y |X), which is read as the expected
value of Y given the value of X.
It is important to distinguish these conditional expected values from the unconditional expected
value of weekly consumption expenditure, E(Y). If we add the weekly consumption expenditures
for all the 60 families in the population and divide this number by 60, we get the number $121.20
($7272/60), which is the unconditional mean, or expected, value of weekly consumption
expenditure, E(Y); it is unconditional in the sense that in arriving at this number we have
disregarded the income levels of the various families. Obviously, the various conditional expected
values of Y given in Table 2.1 are different from the unconditional expected value of Y of $121.20.
When we ask the question, “What is the expected value of weekly consumption expenditure of a
family,” we get the answer $121.20 (the unconditional mean). But if we ask the question, “What
is the expected value of weekly consumption expenditure of a family whose monthly income is,
say, $140,” we get the answer $101 (the conditional mean).
Geometrically, a population regression curve (line) is simply the locus of the conditional means
of the dependent variable for the fixed values of the explanatory variable(s).
From the preceding discussion, it is clear that each conditional mean E(Y | Xi) is a function of Xi,
where Xi
(𝑌⎹𝑋𝑖) = (𝑋𝑖).......................................................3.1.1
Where f (Xi) denotes some function of the explanatory variable X. In the above example, E(Y |
Xi) is a linear function of Xi. Equation (3.1.1) is known as the conditional expectation function
(CEF) or population regression function (PRF) or population regression (PR) for short. It
states merely that the expected value of the distribution of Y given Xi is functionally related to Xi.
In simple terms, it tells how the mean or average response of Y varies with X.
As a first approximation or a working hypothesis, we may assume that the PRF E(Y | Xi) is a linear
function of Xi, say, of the type;
(𝑌⎹𝑋𝑖) = 𝛽𝑋𝑖 … … …… … … …… … …… … … 3.1.2
Where β1 and β2 are unknown but fixed parameters known as the regression coefficients
It is clear from Figure 2.1 that, as family income increases, family consumption expenditure on
the average increases, too. But what about the consumption expenditure of an individual family in
relation to its (fixed) level of income? It is obvious from Table 2.1 and Figure 2.1 that an
individual family‟s consumption expenditure does not necessarily increase as the income level
increases. For example, from Table 2.1 we observe that corresponding to the income level of $100
there is one family whose consumption expenditure of $65 is less than the consumption
expenditures of two families whose weekly income is only $80. But notice that the average
consumption expenditure of families with a weekly income of $100 is greater than the average
consumption expenditure of families with a weekly income of $80 ($77 versus $65).
We see from Figure 2.1 that, given the income level of Xi, an individual family‟s consumption
expenditure
is clustered around the average consumption of all families at that Xi, that is, around its
conditional
expectation. Therefore, we can express the deviation of an individual Yi around its expected
value as follows:
𝜇𝑖 = 𝑌𝑖 − 𝐸(𝑌⎹𝑋𝑖)
Or
𝑌𝑖 = (𝑌⎹𝑋𝑖) + 𝜇𝑖 …… …… …… …… …… …… …… …… … 3.1.3.
where the deviation ui is an unobservable random variable taking positive or negative values.
Technically, ui is known as the stochastic disturbance or stochastic error term.
How do we interpret (3.1.3)? We can say that the expenditure of an individual family, given its
income level, can be expressed as the sum of two components: (1) E(Y | Xi), which is simply the
mean consumption expenditure of all the families with the same level of income. This component
is known as the systematic, or deterministic, component, and (2) ui, which is the random, or
nonsystematic, component. Stochastic disturbance term is a surrogate or proxy for all the omitted
or neglected variables that may affect Y but are not (or cannot be) included in the regression model.
If E(Y | Xi) is assumed to be linear in Xi, as in Eq. (3.1.2), Eq. (3.1.3) may be written as:
𝑌𝑖 = 𝐸(𝑌⎹𝑋𝑖) + 𝜇𝑖
= 𝛽𝑋𝑖 𝜇𝑖 … …… … … …… … ….3.1.4
It is about time to face up to the sampling problems, for in most practical situations what we have
is but a sample of Y values corresponding to some fixed X’s. Therefore, the task now is to estimate
the PRF on the basis of the sample information. As an illustration, pretend that the population of
Table 2.1 was not known to us and the only information we had was a randomly selected sample
of Y values for the fixed X’s as given in Table 2.2. The question is: From the sample of Table 2.2
can we predict the average weekly consumption expenditure Y in the population as a whole
corresponding to the chosen X’s? In other words, can we estimate the PRF from the sample data?
As one surely suspects, we may not be able to estimate the PRF “accurately” because of sampling
fluctuations.
Plotting the data of Tables 2.2 a and 2.2 b, we obtain the scattergram given in Figure 2.2. In the
scattergram two samples regression lines are drawn so as to “fit” the scatters reasonably well:
SRF1 is based on the first sample, and SRF2 is based on the second sample. Which of the two
regression lines represents the “true” population regression line? There is no way we can be
absolutely sure that either of the regression lines shown in Figure 2.2 represents the true population
regression line (or curve).
Table 2.2a: random sample from the Table 2.2b: another random sample from the
population population
Observation Y X Observation Y X
1 70 80 1 55 80
2 65 100 2 88 100
3 90 120 3 90 120
4 95 140 4 80 140
5 110 160 5 118 160
6 115 180 6 120 180
7 120 200 7 145 200
8 140 220 8 135 220
9 155 240 9 145 240
10 150 260 10 175 260
Weekly consumption
expenditure(Y) SRF1
SRF2
Weekly income(X)
Fig 2.2
The regression lines in Figure 2.2 are known as the sample regression lines. They represent the
population regression line, but because of sampling fluctuations they are at best an approximation
of the true PR. Now, analogously to the PRF that underlies the population regression line, we can
develop the concept of the sample regression function (SRF) to represent the sample regression
line. The sample counterpart of (3.1.2) may be written as;
𝑌̂ ̂ 𝛽̂ 𝑋 ……………………………………………………………………….3.1.5
𝑤 𝑟 𝑌̂ 𝑖𝑠 𝑟 𝑠 𝑜𝑟
𝑌̂ 𝑠𝑡𝑖𝑚 𝑡𝑜𝑟 𝑜 𝑌 𝑋
̂ 𝑠𝑡𝑖𝑚 𝑡𝑜𝑟 𝑜
𝛽̂ 𝑠𝑡𝑖𝑚 𝑡𝑜𝑟 𝑜 𝛽
Note that an estimator, also known as a (sample) statistic, is simply a rule or formula or method
that tells how to estimate the population parameter from the information provided by the sample
at hand. A particular numerical value obtained by the estimator in an application is known as an
estimate.
Now just as we expressed the PRF in two equivalent forms, (3.1.2) and (3.1.4), we can express the
SRF (3.1.5) in its stochastic form as follows:
𝑌̂ ̂ 𝛽̂ 𝑋 𝑢̂……………………………………………………………………….3.1.6
Where, in addition to the symbols already defined, 𝑢 ^i denotes the (sample) residual term.
Conceptually 𝑢 ^i is analogous to ui and can be regarded as an estimate of ui.
To sum up, our primary objective in regression analysis is to estimate the PRF
𝑌 𝛽𝑋 𝑢……………………………………………………………………….3.1.4
𝑜𝑛𝑡 𝑏 𝑠𝑖𝑠 𝑜 𝑡 𝑆𝑅
𝑌̂ ̂ 𝛽̂ 𝑋 𝑢̂……………………………………………………………………….3.1.6
Because more often than not our analysis is based upon a single sample from some population.
The deviations of the observations from the line may be attributed to several factors.
In economic reality each variable is influenced by a very large number of factors. However, not
all the factors influencing a certain variable can be included in the function for various reasons.
2. Random behavior of the human beings
The scatter of points around the line may be attributed to an erratic element which is inherent in
human behavior. Human reactions are to a certain extent unpredictable and may cause deviations
from the normal behavioral pattern depicted by the line.
3. Imperfect specification of the mathematical form of the model
We may have linearized a possibly nonlinear relationship. Or we may have left out of the model
some equations.
4. Errors of aggregation
We often use aggregate data (aggregate consumption, aggregate income), in which we add
magnitudes referring to individuals whose behavior is dissimilar. In this case we say that variables
expressing individual peculiarities are missing.
5. Errors of measurement
This refers to errors of measurement of the variables, which are inevitable due to the methods of
collecting and processing statistical information.
The first four sources of error render the form of the equation wrong, and they are usually referred
to as error in the equation or error of omission. The fifth source of error is called error of
measurement or error of observation. In order to take in to account the above sources of error we
introduce in econometric functions a random variable u called random disturbance term of the
function, so called because u is supposed to disturb the exact linear relationship which is assumed
to exist between X and Y.
The first and perhaps more “natural” meaning of linearity is that the conditional expectation of Y
is a linear function of Xi, such as, for example, (3.1.2). Geometrically, the regression curve in this
case is a straight line. In this interpretation, a regression function such as E(Y | Xi) = β1 + β2 𝑋𝑖2
is not a linear function because the variable X appears with a power or index of 2.
ii. Linearity in the Parameters
The second interpretation of linearity is that the conditional expectation of Y, E(Y | Xi), is a linear
function of the parameters, the β‟s; it may or may not be linear in the variable X. In this
interpretation E(Y | Xi) = β1 + β2𝑋𝑖2 is a linear (in the parameter) regression model. Of the two
interpretations of linearity, linearity in the parameters is relevant for the development of the
regression theory to be presented shortly. Therefore, from now on the term “linear” regression will
always mean a regression that is linear in the parameters; the β‟s (that is, the parameters are raised
to the first power only). It may or may not be linear in the explanatory variables, the X‟s.
3.1.5. The Ordinary Least Squares Methods (OLS)
To estimate the coefficients β1 and β2 we need observations on X, Y and u. yet u is never observed
like the other explanatory variables, and therefore in order to estimate the function Yi = β1 +
β2Xi
+ ui, we should guess the values of u, that is we should make some reasonable assumptions about
the shape of the distribution of each ui (its means, variance and covariance with other u‟s). These
assumptions are guesses about the true, but unobservable, value of ui.
The linear regression model is based on certain assumptions, some of which refers to the
32 | Prepared by: Bishaw A. (MSc) Email Address: bish200821@[Link]
Debre Tabor University, CAES, Department of Agricultural Economics
distribution of the random variable u, some to the relationship between u and the explanatory
variables, and some refers to the relationship between the explanatory variables themselves.
1. 𝑢̂ is a random real variable and has zero mean value: E(𝑢̂ ) = 0 (or E(𝑢̂ |Xi) = 0)
o This implies that for each value of X, 𝜇 may assume various values, some
positive, and some negative but on average zero.
o Further E(Yi) = α + βXi gives the relationship between X and Y on the
average, i.e. when X takes on value Xi , then Y will on the average take on
E(Yi) (or E(Yi|Xi))
2. The variance of 𝑢̂ is constant for all i, i.e., var(𝑢̂ 𝑖|Xi) = E(ui2|Xi) = 𝜎2, and is
called the assumptions of common variance or homoscedasticity.
o The implication is that for all values of X, the values of u show the same
dispersion around their mean.
o The consequence of this assumption is that var(yi|Xi) =𝜎2
o If on the other hand the variance of Y population varies as X changes, a situation
of non- constancy of the variance of Y, called heteroskedasticity arises.
3. 𝑢̂ has a normal distribution, i.e., 𝑢̂ ∼ N (0, 𝜎2), which also implies Yi ∼ N ( 𝛽𝑋 , 𝜎2).
4. The random terms of different observations are independent, cov (uiuj) =E (uiuj) = 0 for i
≠ j where i and j run from 1 to n. This is called the assumption of no autocorrelation
(serial) among the error terms.
o The consequence of this assumption is that cov (YiYj) = 0, for i ≠ j i.e. no
autocorrelation among the Y‟s.
5. Xi‟s are a set of fixed values in the process of repeated sampling which underlies the
linear regression model, i.e. they are non-stochastic.
6. 𝑢̂ is independent of the explanatory variables, i.e., cov (𝑢̂ Xi) = E(𝑢̂ Xi) = 0.
7. Variability in X values. The X values in a given sample must not all be the same.
Technically, var(X) must be a finite positive number.
8. The regression model is correctly specified.
3.2 The Least Square Criterion and Normal Equations of OLS
Thus far we have completed the work involved in the first stage of any econometric application,
namely we have specified the model and stated explicitly its assumptions. The next step is the
estimation of the model, that is, the computation of the numerical values of its parameters.
The linear relationship 𝑌 𝛽𝑋 𝑢 holds for the population of the values of X and Y, so
that we could obtain the numerical values of 𝑛 𝛽 only if we could have all the possible
values of X, Y and u which form the population of these variables. Since this is impossible in
practice, we get a sample of observed values of Y and X, specify the distribution of the u‟s and try
to get satisfactory estimates of the true parameters of the relationship. This is done by fitting a
regression line through the observations of the sample, which we consider as an approximation to
the true line.
The method of ordinary least squares is one of the econometric methods which enable us to find
the estimate of the true parameter and is attributed to Carl Friedrich Gauss, a German
mathematician. To understand this method, we first explain the least squares principle.
Recall the two-variable PRF:
𝑌 𝛽𝑋 𝑢……………………………………………………………………….3.1.4
However, as noted in earlier, the PRF is not directly observable. We estimate it from the SRF:
𝑌 ̂ 𝛽̂ 𝑋 𝑢̂……………………………………………………………………….3.1.6
𝑌 𝑌̂ 𝑢̂………………………………………………………………………………...(3.1.7)
𝑤 𝑟 𝑌̂ 𝑖𝑠 𝑡 𝑠𝑡𝑖𝑚 𝑡 𝑜𝑛 𝑖𝑡𝑖𝑜𝑛 𝑙 𝑚 𝑛 𝑜 𝑌
But how is the SRF itself determined? To see this, let us proceed as follows. First, express (3.1.7) as
𝑢̂ 𝑌 𝑌̂ 𝑌 ̂ 𝛽̂ 𝑋
Which shows that the 𝑢̂ (the residuals) are simply the differences between the actual and estimated
Y values. Now given n pairs of observations on Y and X, we would determine the SRF in such a
manner that it is as close as possible to the actual Y. To this end, we adopt the least-squares
criterion, which states that the SRF can be fixed in such a way that;
∑ 𝑢̂ = ∑(𝑌𝑖 − 𝑌𝜄 ) 2
= ∑(𝑌𝑖 ̂ ̂ 𝑋𝑖 )2 … … … … … … … … … … … … … … … … … … … … .3.1.9
𝛽
It is obvious from (3.1.8) that i2= f (α, 𝛽) that is, the sum of the squared residuals is some
function of the estimators α and 𝛽. For any given set of data, choosing different values for α and
The principle or the method of least squares chooses α and 𝛽 in such a manner that, for a given
sample or set of data, i2 is as small as possible. In other words, for a given sample, the method
of least squares provides us with unique estimates of α and β that give the smallest possible value
of i2.
The process of differentiation yields the following equations for estimating α and β.
Differentiating Eq. (3.1.8) partially with respect to α and 𝛽, we obtain;
∑ 𝑢̂
∑ 𝑌 ̂ 𝛽̂ 𝑋
̂
∑ 𝑢̂
∑ 𝑌 ̂ 𝛽̂ 𝑋 𝑋
𝛽̂
∑ 𝑌𝑖 = 𝑛α + 𝛽 ∑ 𝑋 … … … … … … … … … … … …… … … … … … … … … … . .3.1.9
∑𝑌 𝑋 ̂ ∑𝑋 ̂
𝛽 ∑𝑋
Where n is the sample size. These simultaneous equations are known as the normal equations.
Solving the normal equations simultaneously, we obtain;
∑ 𝑌 𝑌̅ 𝑋 𝑋̅
𝛽 𝑛 ̂ 𝑌̅ 𝑌̂𝑋̅
∑ 𝑋 𝑋̅
Where 𝑋 and 𝑌are the sample means of X and Y and where we define xi = (Xi − 𝑋) and yi = (Yi
− 𝑌). The above lowercase letters in the formula denote deviations from mean values. Equation (
3.1.11) can be obtained directly from (3.1.9) by simply dividing both sides of the equation by n.
35 | Prepared by: Bishaw A. (MSc) Email Address: bish200821@[Link]
Debre Tabor University, CAES, Department of Agricultural Economics
Note that, by making use of simple algebraic identities, formula (3.1.11) for estimating β can be
alternatively expressed as;
∑𝑥 𝑦 ∑𝑥 𝑦
𝛽̂
∑𝑥 ∑𝑥 𝑛𝑋̅
The estimators obtained previously are known as the least-squares estimators, for they are derived
from the least-squares principle. We finally write the regression line equation as 𝑌i = 𝛽 + 𝛽 Xi.
Interpretation of estimates
Estimated intercept, α: The estimated average value of the dependent variable
when the independent variable takes on the value zero
Estimated slope 𝛽̂ :The estimated change in the average value of the dependent
variable when the independent variable increases by one unit.
𝑌i gives average relationship between Y and X. i.e., 𝑌i is average value of Y given Xi.
3.3 Precision or Standard Errors of Least-Squares Estimates
It is evident that least-squares estimates are a function of the sample data. But since the data are likely
to change from sample to sample, the estimates will change ipso facto. Therefore, what is needed is
some measure of “reliability” or precision of the estimator‟s α and 𝛽. In statistics the precision of an
estimate is measured by its standard error (SE). The standard errors of the OLS estimates can be
obtained as follows:
(̂)
∑ ∑ ̅
̂
√∑
∑
̂
∑
∑
∑ ̅
∑
̂ ̂
√
∑ ̅
where var = variance and se = standard error and where σ2 is the constant or homoscedastic
variance of ui of Assumption 2.
All the quantities entering into the preceding equations except σ2 can be estimated from the data.
∑ 𝑢̂ 𝑅𝑆𝑆
𝜎̂
𝑛 𝑛
Where σ2 is the OLS estimator of the true but unknown σ2 and where the expression n − 2 is known as
the number of degrees of freedom (df), 𝑢𝑖2 being the sum of the residuals squared or the residual
sum of squares (RSS).
∑ 𝑢̂ 𝑅𝑆𝑆
𝜎̂ √𝜎̂ √ √
𝑛 𝑛
𝜎̂ 𝑖𝑠 𝑘𝑛𝑜𝑤𝑛 𝑠 𝑡
It is simply the standard deviation of the Y values about the estimated regression line and is often used as a
summary measure of the “goodness of the fit” of the estimated regression line.
Note the following features of the variances (and therefore the standard errors) of ̂ 𝑛 𝛽̂
the variance of 𝛽̂ is directly proportional to 𝜎
̂ but inversely proportional to ∑
̂ and ∑
the variance of ̂ is directly proportional to 𝜎 but inversely proportional to ∑ and
sample size(n)
Since ̂ 𝑛 𝛽̂ are estimators, they will not only vary from sample to sample but in a given sample they
are likely to be dependent on each other, this dependence being measured by the covariance between
them.
As noted earlier, given the assumptions of the classical linear regression model, the least-
squares estimates possess some ideal or optimum properties. These properties are contained in
the well-known Gauss–Markov theorem. To understand this theorem, we need to consider the
best linear unbiasedness property of an estimator. An estimator, say the OLS estimator 𝛽̂, is
Gauss–Markov Theorem: Given the assumptions of the classical linear regression model, the least-squares
estimators, in the class of unbiased linear estimators, have minimum variance, that is, they are BLUE.
Thus far we were concerned with the problem of estimating regression coefficients, their standard
errors, and some of their properties. We now consider the goodness of fit of the fitted regression line to
a set of data; that is, we shall find out how “well” the sample regression line fits the data. It is clear that
if all the observations were to lie on the regression line, we would obtain a “perfect” fit, but this is
rarely the case. Generally, there will be some positive 𝑢̂ and some negative 𝑢̂ . What we hope for is
that these residuals around the regression line are as small as possible.
The coefficient of determination 𝑟 (two-variable case) or 𝑅 (multiple regressions) is a summary
measure that tells how well the sample regression line fits the data. 𝑟 measures the proportion or
percentage of the total variation in Y explained by the independent variable X (the regression model).
To compute this 𝑟 , we proceed as follows: Recall that:
𝑌 𝑌̂
𝑢̂ or in the deviation form
𝑦 𝑦̂ 𝑢̂
Squaring on both sides and summarizing over the sample we obtain
∑𝑦 ∑ 𝑦̂ ∑ 𝑢̂ ∑ 𝑦̂ 𝑢̂
∑𝑦 ∑ 𝑦̂ ∑ 𝑢̂
……………………….3.1.18
∑𝑦 𝛽̂ ∑ 𝑥 ∑ 𝑢̂
since ∑ 𝑦̂ 𝑢̂ 𝑤 𝑦 and 𝑦̂ 𝛽̂ 𝑥
The various sums of squares appearing in equation (3.1.18) thus, known as measure of variation can
be described as follows:
∑𝑦 ∑ 𝑌 𝑌̅ 𝑇𝑆𝑆
The total variation of the actual Y values about their sample mean, which may be called total sum
square (TSS)
∑ 𝑌̂ ∑ 𝑌̂ 𝑌̅̂ ∑ 𝑌 𝑌̅ 𝛽̂ ∑ 𝑥 𝑆𝑆 𝑜𝑟 𝐸𝑆𝑆
The variation of the estimated Y values about their mean, which is called sum of squares due to
regression (i.e due to the explanatory variables or explained by regression or simply explained sum of
squares also called model sum of squares(ESS/MSS).
𝑅𝑆𝑆 ∑𝑢 𝛽̂ ∑ 𝑥
RSS= residual or unexplained variation of the Y values about the regression line, which is called
residual sum of squares.
𝑌𝑖 𝑌̅ 𝑇𝑜𝑡𝑎𝑙
𝑌̂𝑖 𝛼 𝛽𝑋𝐼
𝑌̅
𝑌̂𝑖 𝑌̅ 𝑑𝑢𝑒 𝑡𝑜 𝑟𝑒𝑔𝑟𝑒𝑠𝑠𝑖𝑜𝑛
X
0 Xi
∑ ̂ ̅ ∑
Now we divide by TSS both sides we get ∑ ̅ ∑ ̅
We now define 𝑟 as 𝑟 ∑ 𝑌̂ 𝑌̅ ∑ 𝑌 𝑌̅
Or alternatively, 𝑟 ∑ 𝑢̂ ∑ 𝑌 𝑌̅
𝑟 ∑ 𝑦̂ ∑𝑦
∑𝑥
𝑟 𝛽̂ ∑ 𝑥 ∑𝑦 𝛽̂
∑𝑦
If we divide the numerator and the denominator of (3.1.26) by the sample size n (or n− 1 if the sample
size is small), we obtain
𝑟 𝛽 ) -------------------------------------------------------------------------------3.1.27
∑
Where 𝑆 and 𝑆 are the sample variances of Y and X, respectively. Since 𝛽̂ ∑
the equation
[∑ 𝑌 𝑌̅ (𝑌̂ 𝑌̅)]
𝑟
∑ 𝑌 𝑌̅ ∑(𝑌̂ 𝑌̅)
that is
∑ 𝑦 𝑦̂
𝑟
∑ 𝑦 ∑ 𝑦̂
A Numerical Example
We illustrate the econometric theory developed so far by considering the Keynesian consumption
function discussed in the Introduction. Recall that Keynes stated that “The fundamental psychological
law is that men(women) are disposed, as a rule and on average, to increase their consumption as their
income increases, but not by as much as the increase in their income,” that is, the marginal propensity
to consume (MPC) is greater than zero but less than one. Although Keynes did not specify the exact
functional form of the relationship between consumption and income, for simplicity assume that the
relationship is linear as in (2.4.2). As a test of the Keynesian consumption function, we use the sample
data of Table 2.4, which for convenience is reproduced as Table 3.3. The raw data required to obtain
the estimates of the regression coefficients, their standard errors, etc., are given in Table 3.3. From
these raw data, the following calculations are obtained, and the reader is advised to check them.
Table: Hypothetical Data on Weekly Family Consumption Expenditure Y and Weekly Family Income X
Observati Y X 𝑦 𝑌 𝑌̅ 𝑥 𝑋 𝑋̅ 𝑥 𝑦 =(𝑌 𝑌̅) 𝑋 𝑋̅) 𝑥
on
1 70 80 -41 -90 3690 8100
2 65 100 -46 -70 3220 4900
3 90 120 -21 -50 1050 2500
4 95 140 -16 -30 480 900
5 110 160 -1 -10 10 100
6 115 180 4 10 40 100
7 120 200 9 30 270 900
8 140 220 29 50 1450 2500
9 155 240 44 70 3080 4900
10 90
150 260 39 3510 8100
Sum
1110 1700 16800 33000
Mean
111 170
Observati Y X 𝑌̂ 𝑢̂ 𝑌 𝑌̂ 𝑢̂ 𝑋
on
1 70 80 65.18181818 4.818181818 23.21487603 6400
2 65 100 75.36364545 -10.36364545 107.4051471 10000
3 90 120 85.54547 4.45453 19.84283752 14400
4 95 140 95.72729 -0.72729 0.528950744 19600
5 110 160 105.90911 4.09089 16.73538099 25600
6 115 180 116.09093 -1.09093 1.190128265 32400
7 120 200 126.27275 -6.27275 39.34739256 40000
8 140 220 136.45457 3.54543 12.57007388 48400
9 155 240 146.63639 8.36361 69.94997223 57600
10
150 260 156.81821 -6.81821 46.4879876 67600
sum
1110 1700 1110.000184 337.2727469 322000
mean
111 170 RSS
Observati Y X 𝑦̂ =𝑌̂ 𝑌̅ 𝑦̂ 𝑢̂ 𝑦
on
1 70 80 -45.81820022 2099.307471 23.21487603 1681
2 65 100 -35.63637295 1269.951077 107.4051471 2116
3 90 120 -25.4545484 647.9340342 19.84283752 441
4 95 140 -15.2727284 233.2562328 0.528950744 256
5 110 160 -5.0909084 25.91734834 16.73538099 1
6 115 180 5.0909116 25.91738092 1.190128265 16
7 120 200 15.2727316 233.2563305 39.34739256 81
8 140 220 25.4545516 647.9341972 12.57007388 841
9 155 240 35.6363716 1269.950981 69.94997223 1936
10
150 260 45.8181916 2099.306681 46.4879876 1521
Sum
1110 1700 8552.731734 337.2727469 8890
Mean ESS TSS
111 170 RSS
∑
𝛽̂ ∑
̂ 𝑌̅ 𝛽̂ 𝑋̅
𝜎 √𝜎 𝑟 √𝑟
∑
var(𝛽̂ ) ∑
, var( ̂) ∑
, SE(𝛽̂ ) 𝜎̂ √ 𝛽̂ , SE
( ̂) 𝜎̂ √ ̂
The estimated regression line therefore is 𝑌̂ 𝑋𝑖
The associated regression line are interpreted as follows: Each point on the regression line gives an
estimate of the expected or mean value of Y corresponding to the chosen X value; that is, 𝑌̂ is an
estimate of E(Y | Xi). The value of 𝛽̂ = 0.5091, which measures the slope of the line, shows that,
within the sample range of X between $80 and $260 per week, as X increases, say, by $1, the estimated
increase in the mean or average weekly consumption expenditure amounts to about 51 cents. The value
of ̂ = 24.4545, which is the intercept of the line, indicates the average level of weekly consumption
expenditure when weekly income is zero. However, this is a mechanical interpretation of the intercept
term. In regression analysis such literal interpretation of the intercept term may not be always
meaningful; although in the present example it can be argued that a family without any income
(because of unemployment, layoff, etc.) might maintain some minimum level of consumption
expenditure either by borrowing or dissaving. But in general one has to use common sense in
interpreting the intercept term, for very often the sample range of X values may not include zero as one
of the observed values.
Perhaps it is best to interpret the intercept term as the mean or average effect on Y of all the variables
omitted from the regression model. The value of 𝑟 of 0.9621 means that about 96 percent of the
variation in the weekly consumption expenditure is explained by income. Since 𝑟 can at most be 1, the
observed 𝑟 suggests that the sample regression line fits the data very well. The coefficient of
correlation, r of 0.9809 shows that the two variables, consumption expenditure and income are highly
positively correlated.
i) test of significance by using Standard error test
This test helps us decide whether the estimates are significantly different from zero, i.e. whether the
sample from which they have been estimated might have come from a population whose true
parameters are zero. .
Formally we test the null hypothesis Against the alternative hypothesis
The standard error test may be outlined as follows.
First: Compute standard error of the parameters.
Second: compare the standard errors with the numerical values of ̂ 𝑛 𝛽̂ .
Decision rule:
If 𝑆𝐸(𝛽̂ ) 𝛽̂ , accept the null hypothesis and reject the alternative hypothesis. We conclude
step 1: Test the significance of the slope parameter at 5% level of significance using the standard error
test.
𝑆𝐸(𝛽̂ ) 𝑛 𝛽̂ 𝑡 𝑟 𝑜𝑟 ⁄ 𝛽̂ =
level of significance.
Note: The standard error test is an approximated test (which is approximated from the z-test
and t-test) and implies a two tail test conducted at 5% level of significance.
Steps in hypothesis testing: the following procedure are involved in hypothesis testing
The three approach s for hypothesis testing: there are three most common approaches for hypothesis
testing.
1. Test- statistic approach(critical value approach)
2. Confidence interval approach
3. P-value approach
47 | Prepared by: Bishaw A. (MSc) Email Address: bish200821@[Link]
Debre Tabor University, CAES, Department of Agricultural Economics
̂
𝑡̂ ̂ With n-k degree of freedom.
̂
𝑡̂
𝑆𝐸 ̂
Where: SE = is standard error and k = number of parameters in the model.
Since we have two parameters in simple linear regression with intercept different from zero, our
degree of freedom is n-2. Like the standard error test we formally test the hypothesis: 𝛽
against the alternative 𝛽 for the slope parameter; and against the
alternative for the intercept
To undertake the above test we follow the following steps.
Step 1: Compute t*, which is called the computed value of t, by taking the value of 𝛽 in the
null hypothesis. In our case, then t* becomes:
̂ ̂
𝑡 ̂ (̂)
---------------------------------------------------3.1.34
Step 2: Choose level of significance. Level of significance is the probability of making „wrong‟
decision, i.e. the probability of rejecting the hypothesis when it is actually true or the probability
of committing a type I error. It is customary in econometric research to choose the 5% or the
1% level of significance. This means that in making our decision we allow (tolerate) five times
out of a hundred to be „wrong‟ i.e. reject the hypothesis when it is actually true.
Step 3: Check whether there is one tail test or two tail tests. If the inequality sign in the
alternative hypothesis is, then it implies a two tail test and divide the chosen level of
significance by two; decide the critical rejoin or critical value of t called tc. But if the inequality
sign is either > or < then it indicates one tail test and there is no need to divide the chosen level
of significance by two to obtain the critical value from the t-table.
Example: 𝛽
𝛽
Then this is a two tail test. If the level of significance is 5%, divide it by two to obtain critical
value of t from the t-table.
Step 4: Obtain critical value of t, called tc at and n-2 degree of freedom for two tail test.
Step 5 decision rule: Compare t* (the computed value of t) and tc (critical value of t)
48 | Prepared by: Bishaw A. (MSc) Email Address: bish200821@[Link]
Debre Tabor University, CAES, Department of Agricultural Economics
b. Since the alternative hypothesis (H1) is stated by inequality sign ( ) ,it is a two tail test,
hence we divide ⁄ to obtain the critical value of „t‟ at =0.025 and 18 degree
of freedom (df) i.e. (n-2=20-2). From the t-table „tc‟ at 0.025 level of significance and 18 df is
2.10.
c. Since t*=3.3 and tc=2.1, t*>tc. It implies that is statistically significant.
Right tail 𝛽 𝛽 𝛽 𝛽 𝑡 𝑡
Left tail 𝛽 𝛽 𝛽 𝛽 𝑡 𝑡
inevitable in all estimates, it is necessary to apply test of significance in order to measure the size of the
error and determine the degree of confidence in order to measure the validity of these estimates.
It is very important to know the following aspects of interval estimation:
Rejection of the null hypothesis doesn‟t mean that our estimate ̂ 𝑛 𝛽̂ is the correct estimate
of the true population parameter 𝑛 𝛽
It simply means that our estimate 𝛽 comes from a sample drawn from a population whose
parameter is different from zero.
In order to define how close the estimate to the true parameter, we must construct confidence
interval for the true parameter,
in other words we must establish limiting values around the estimate with in which the true
parameter is expected to lie within a certain “degree of confidence”.
In this respect we say that with a given probability the population parameter will be within the
defined confidence interval (confidence limits).
We choose a probability in advance and refer to it as confidence level (interval coefficient). It
is customarily in econometrics to choose the 95% confidence level.
The hypothetical consumption-income econometric model example
𝑌̂ 𝑋 --------------------------------------------------------3.1.33
Shows that the estimated marginal propensity to consume (MPC) β is 0.5091, which is a single (point)
estimate of the unknown population MPC β. This is a single (point) estimate of the unknown
population MPC β. Because of sampling fluctuations, a single estimate is likely to differ from the true
value, although in repeated sampling its mean value is expected to be equal to the true value[Note:
E(𝛽̂ = β].
To be more specific, assume that we want to find out how “close” is, say, 𝛽̂ 𝑡𝑜 𝛽. For
this purpose we try to find out two positive numbers δ and α, the latter lying between 0
and 1, such that the probability that the random interval (𝛽̂ − δ, β̂ + δ) contains the true
β2 is 1 − α. Symbolically,
Pr (̂
𝛽 −δ ≤β ≤ ̂
𝛽 + δ) = 1 – α………………………………………………………………..3.1.34
Such an interval, if it exists, is known as a confidence interval; 1 - α is known as the confidence
coefficient; and α (0 < α < 1) is known as the level of significance.2 The endpoints of the confidence
interval are known as the confidence limits (also known as critical values), βˆ - δ being the
lower confidence limit and βˆ + δ the upper confidence limit. In passing, note that in practice α and 1 -
α are often expressed in percentage forms as 100α and 100(1 - α) percent. Equation (3.1.34) shows that
an interval estimator, in contrast to a point estimator, is an interval constructed in such a manner that
it has a specified probability 1 - α of including within its limits the true value of the
parameter. For example, if α = 0.05, or 5 percent, (3.1.34) would read: The probability that the
(random) interval shown there includes the true β is 0.95, or 95 percent. The interval estimator thus
gives a range of values within which the true β may lie.
It is very important to know the following aspects of interval estimation:
1. Equation (3.1.34) does not say that the probability of β lying between the given limits is 1 − α.
Since β, although an unknown, is assumed to be some fixed number, either it lies in the interval
or it does not. What? (3.1.34) states is that, for the method described in this chapter, the
probability of constructing an interval that contains β is 1 − α.
2. The interval (3.1.34) is a random interval; that is, it will vary from one sample to the next
because it is based on βˆ, which is random. (Why?)
3. Since the confidence interval is random, the probability statements attached to it should be
understood in the long-run sense, that is, repeated sampling. More specifically, (3.1.34) means:
If in repeated sampling confidence intervals like it are constructed a great many times on the 1 −
α probability basis, then, in the long run, on the average, such intervals will enclose in 1 − α of
the cases the true value of the parameter.
4. As noted in 2, the interval (3.1.34) is random so long as βˆ is not known.
Confidence Interval for ̂ :with the normality assumption for ui, the OLS estimators ̂ ̂ are
themselves normally distributed with means and variances given therein. Therefore, for example, the
̂ ̂ √∑
variable. 𝑧 ̂
Is a standardized normal variable. It therefore seems that we can use the normal distribution to make
probabilistic statements about β provided the true population variance σ 2 is known. If σ 2 is known, an
important property of a normally distributed variable with mean µ and variance σ 2 is that the area
under the normal curve between µ ± σ is about 68 percent, that between the limits µ ± 2σ is about 95
But is rarely known, and in practice it is determined by the unbiased estimator ̂ If we replace σ by
̂ , (3.1.35) may be written as
Therefore, instead of using the normal distribution, we can use the t distribution to establish a
confidence interval for β as follows:
pr[ 𝑡 ⁄ 𝑡 𝑡 ⁄ ] -------------------------------------------------3.1.37
̂
pr[ 𝑡 ⁄ (̂)
𝑡 ⁄ ]
In the language of hypothesis testing, the 100(1 - α) % confidence interval established in (3.1.39) is
known as the region of acceptance (of the null hypothesis) and the region(s) outside the confidence
interval is (are) called the region(s) of rejection (of H0) or the critical region(s). As noted previously,
the confidence limits, the endpoints of the confidence interval, are also called critical values.
Equation (3.1.39) provides a 100(1 - α) percent confidence interval for β, which can be written more
compactly as 100(1 - α) % confidence interval for β:
𝛽̂ 𝑡 ⁄ 𝑆𝐸(𝛽̂ )----------------------------------------------------------------------------3.1.40
̂ 𝑡 ⁄ 𝑆𝐸 ̂ ----------------------------------------------------------------------------3.1.41
Notice an important feature of the confidence intervals given in (3.1.40) and (3.1.41): In both cases the
52 | Prepared by: Bishaw A. (MSc) Email Address: bish200821@[Link]
Debre Tabor University, CAES, Department of Agricultural Economics
width of the confidence interval is proportional to the standard error of the estimator. That is, the
larger the standard error, the larger is the width of the confidence interval. Put differently, the larger the
standard error of the estimator, the greater is the uncertainty of estimating the true value of the
unknown parameter. Thus, the standard error of an estimator is often described as a measure of the
precision of the estimator, i.e., how precisely the estimator measures the true population value.
Returning to our illustrative consumption–income example, in example we found that βˆ = 0.5091, se
(βˆ) = 0.0357, and df = 8. If we assume α = 5%, that is, 95% confidence coefficient, then the t table
shows that for 8 df the critical tα/2 = t0.025 = 2.306. Substituting these values in (3.1.39), the reader
should verify that the 95% confidence interval for β is as follows:
𝛽 Or 𝑡 𝑡 𝑖𝑠 0.5091 ± 0.0823
The interpretation of this confidence interval is: Given the confidence coefficient of 95%, in the
long run, in 95 out of 100 cases intervals like (0.4268, 0.5914) will contain the true β. But, as warned
earlier, we cannot say that the probability is 95 percent that the specific interval (0.4268 to 0.5914)
contains the true β because this interval is now fixed and no longer random; therefore, β either lies in it
or does not: The probability that the specified fixed interval includes the true β is therefore 1 or 0.
Decision Rule: Construct a 100(1 - α) % confidence interval for β. If the β under 𝐻 falls
within this confidence interval do not reject H0, but if it falls outside this interval, reject 𝐻
Is the observed 𝛽̂ compatible with H0? To answer this question, let us refer to the confidence interval
(3.1.39). We know that in the long run intervals like (0.4268, 0.5914) will contain the true β with 95
percent probability. Consequently, in the long run (i.e., repeated sampling) such intervals provide a
range or limits within which the true β may lie with a confidence coefficient of, say, 95%. Thus, the
confidence interval provides a set of plausible null hypotheses. Therefore, if β under H0 falls within the
100(1 − α) % confidence interval, we do not reject the null hypothesis; if it lies outside the interval, we
may reject it.
Following this rule, for our hypothetical example, H0: β = 0.3 clearly lies outside the 95% confidence
interval given in (3.1.39). Therefore, we can reject the hypothesis that the true MPC is 0.3, with 95%
confidence. If the null hypothesis were true, the probability of our obtaining a value of MPC of as
much as 0.5091 by sheer chance or fluke is at the most about 5 percent, a small probability.
In statistics, when we reject the null hypothesis, we say that our finding is statistically significant. On
the other hand, when we do not reject the null hypothesis, we say that our finding is not statistically
significant. Some authors use a phrase such as “highly statistically significant.” By this they usually
mean that when they reject the null hypothesis, the probability of committing a Type I error (i.e., α) is a
small number, usually 1 percent. But as our discussion of the p value approach will be show, it is better
to leave it to the researcher to decide whether a statistical finding is “significant,” “moderately
significant,” or “highly significant.”
One-Sided or One-Tail Test
Sometimes we have a strong a priori or theoretical expectation (or expectations based on some previous
empirical work) that the alternative hypothesis is one-sided or unidirectional rather than two-sided, as
just discussed. Thus, for our consumption–income example, one could postulate that
𝛽 and
𝛽
Perhaps economic theory or prior empirical work suggests that the marginal propensity to consume is
greater than 0.3. The actual mechanics are better explained in terms of the test-of-significance approach
Figure 3.4: the 95% confidence interval for βˆ under the hypothesis that β = 0.3
Numerical Example 2: Suppose we have estimated the following regression line from a sample of 20
observations.
𝑋
𝑃 Weak or no evidence
𝑃 Moderate evidence
𝑃 Strong evidence
Example: As just noted, the Achilles heel of the classical approach to hypothesis testing is its
arbitrariness in selecting α. Once a test statistic (e.g., the t statistic) is obtained in a given
example, why not simply go to the appropriate statistical table and find out the actual
probability of obtaining a value of the test statistic as much as or greater than that obtained in
the example? This probability is called the p value (i.e., probability value), also known as the
observed or exact level of significance or the exact probability of committing a Type I
error. More technically, the p value is defined as the lowest significance level at which a null
hypothesis can be rejected. To illustrate, let us return to our consumption–income example.
Given the null hypothesis that the true MPC is 0.3, we obtained a t value of 5.86. What is the p
57 | Prepared by: Bishaw A. (MSc) Email Address: bish200821@[Link]
Debre Tabor University, CAES, Department of Agricultural Economics
value of obtaining a t value of as much as or greater than 5.86? Looking up the t table given in
t-distribution, we observe that for 8 df the probability of obtaining such a t value must be much
smaller than 0.001 (one-tail) or 0.002 (two-tail). By using the computer, it can be shown that the
probability of obtaining a t value of 5.86 or greater (for 8 df) is about 0.000189.14 This is the p
value of the observed t statistic. This observed, or exact, level of significance of the t statistic is
much smaller than the conventionally, and arbitrarily, fixed level of significance, such as 1, 5,
or 10 percent. As a matter of fact, if we were to use the p value just computed, and reject the
null hypothesis that the true MPC is 0.3, the probability of our committing a Type I error is only
about 0.02 percent, that is, only about 2 in 10,000!
̂ ∑ ̂ ∑
F-statistic: -------------------------------3.1.42
̂ ∑̂ ̂
The F ratio of (3.1.42) provides a test of the null hypothesis H0: β = 0. Since all the quantities
entering into this equation can be obtained from the available sample, this F ratio provides a test
statistic to test the null hypothesis that true β is zero. All that needs to be done is to compute the
F ratio and compare it with the critical F value obtained from the F tables at the chosen level of
significance, or obtain the p value of the computed F statistic.
To illustrate, let us continue with our consumption–income example. The ANOVA table for this example is as
shown in Table below. The computed F value is seen to be 202.87. The p value of this F statistic corresponding
to 1 and 8 df cannot be obtained from the F table given in F-distribution, but by using electronic statistical tables
it can be shown that the p value is 0.0000001, an extremely small probability indeed. If you decide to choose
the level-of-significance approach to hypothesis testing and fix α at 0.01, or a 1 percent level, you can see that the
computed F of 202.87 is obviously significant at this level. Therefore, if we reject the null hypothesis that β2 = 0,
the probability of committing a Type I error is very small.
Review Question
1. A local restaurant advocacy group wants to study the relationship between a restaurant‟s average
weekly profits; the restaurant‟s seating capacity and average daily traffic that passes the
restaurant‟s location. The group took a sample of restaurants and recording their average weekly
profit (in $1000s), the seating restaurant‟s seating capacity, and the average number of cars (in
1000s) that passes the restaurant‟s location. The data is recorded in the following table:
a. Find the regression model to predict the average weekly profit from the other variables.
b. Interpret the coefficient for seating capacity.
c. Interpret the coefficient for traffic count.
Observation Seating Capacity Traffic Count (1000s) Weekly Net Profit ($1000s)
1 120 19 23.8
2 180 8 29.2
3 150 12 22
4 180 15 26.2
5 220 16 33.5
6 235 10 32
7 115 18 22.4
8 110 12 20.4
9 165 21 23.7
10 220 20 34.7
11 140 24 27.1
12 145 24 23.3
13 140 13 20.9
14 200 14 29.6
15 210 14 31.4
16 175 12 23.2
17 175 15 31.1
18 190 17 28.2
19 100 23 25.2
20 145 20 20.7
21 135 13 37.2
22 25 13 26.3
23 140 25 20
24 130 14 28.2
25 135 10 24.6
26 160 23 23.7
Sum
Average
d) Predict the average weekly profit for a restaurant with a seating capacity of 150 and a traffic
count of 25,000 cars.
e) Find the adjusted coefficient of determination
a) Find the regression model to predict GPA from the other variables.
b) Interpret the coefficient for the average number of hours spent studying each night.
c) Interpret the coefficient for the average number of nights a student goes out each week.
d) Predict the GPA for a student who spends an average of 4 hours a night studying and goes out
an average of 3 nights a week.
e) Find the adjusted coefficient of determination.
f) Interpret the adjusted coefficient of determination.
g) Find the standard error of the estimate.
h) Interpret the standard error of the estimate.
i) At the 1% significance level, test the validity of the model.
j) At the 1% significance level, test the coefficient of the average number of hours spent
studying each night.
k) At the 1% significance level, test the coefficient of the average number of nights a student
goes out each week.
1 3.72 5 1
2 3.88 3 1
3 3.67 2 1
4 3.87 3 4
5 2.49 1 4
6 1.29 1 2
7 1.01 2 4
8 2.12 1 1
9 1.9 1 5
10 3.42 3 2
11 1.33 1 4
12 1.07 0 2
13 2.75 3 1
14 3.82 4 1
15 3.91 5 0
16 2.25 2 3
17 2.06 1 5
18 2.92 3 2
19 3.06 3 1
20 3.65 2 2
21 3.69 4 1
3. A very large company wants to study the relationship between the salaries of employees in
management positions, their age, the number of years the employee spent in college, and the
number of years the employee has been with the company. A sample management employees is
taken and the data recorded below:
a. Find the regression model to predict salary from the other variables.
b. Interpret the coefficient for age.
c. Interpret the coefficient for years of college.
d. Interpret the coefficient for years with the company.
e. Predict the salary for a 47 year old management employee who spent 5 years in college
and has been with the company for 15 years.
f. Find the adjusted coefficient of determination.
g. Interpret the adjusted coefficient of determination.
h.