Assignment 1
a. Explain the nature of econometrics and economic data
b. Write down simple and multiple linear regression models
Explain the nature of econometrics and economic data:
What is econometrics?
First, let us see something about the origin of econometrics as a discipline. The term
econometrics is believed to have been crafted by Ragnar Frisch, co-winner of the first Nobel
Prize in Economic Sciences in 1969, along with fellow econometrician Jan Tinbergen. Both of
them were founders of the Econometric Society in 1933. In the constitution of this society, it is
stated that
“The Econometric Society is an international society for the advancement of economic theory in
its relation to statistics and mathematics. Its main object shall be to promote studies that aim at a
unification of the theoretical-quantitative and the empirical-quantitative approach to economic
problems and that are penetrated by constructive and rigorous thinking similar to that which has
come to dominate the natural sciences”
Today, we would also say that econometrics is the combined study of economic models,
mathematical statistics, and economic data. Within the field of econometrics, econometric theory
can be distinguished from applied econometrics. Econometric theory concerns the development
of tools and methods, and the study of the properties of econometric methods. Econometric
theory belongs to the field of statistics.
Applied econometrics is a term describing the development of quantitative economic models and
the application of econometric methods to these models using economic data. Applied
econometrics is mainly used in the field of applied economics.
Econometrics is the set of tools by which economists, and others in the social sciences, analyze
data. We can use econometrics to:
1. Estimate economic relationships;
2. Test economic theories;
3. Evaluate government and business policy
Econometrics focuses on problems inherent in analyzing data generated by individuals, firms,
and other entities acting strategically, and interacting with one another.
An empirical analysis uses data to test a theory, estimate an economic relationship, or determine
the effects of a policy or intervention. Econometrics allows us to analyze data using formal
statistical methods.
Why Study Econometrics
1. Important to be able to apply economic theory to real world data.
Page | 1
2. Theory may be ambiguous as to the effect of some policy change, and in any case theory
rarely tells us how large the effect might be.
3. Forecasting economic variables (inflation, interest rates, housing starts, and so on) is
important, too.
What are the goals of Econometrics? We are going to examine three:
1. Knowledge of the real economy: Econometric methods allow us to estimate economic
magnitudes such as the marginal propensity to consume or the elasticity of labor with
respect to output. These estimations are located in a determined time and space: for
example, in Spain in the last quarter of the 20th century. In addition to the estimation, in
which numerical values are obtained, econometric methods allow us to perform tests of
hypothesis; for example, in a production function, is the hypothesis of constant returns to
scale admissible?
2. Economic simulation policy: Econometrics methods can be used to simulate the effects
of alternative policies. For example, with an appropriate econometric model we could
see, in quantitative terms, how the different increases in tobacco tax affect the
consumption of tobacco
3. Prediction or forecasting: Very often econometric methods are used to predict values of
economic variables in the future. By making predictions we try to reduce our uncertainty
in the future of the economy. This is not an easy task, since in general the predictions are
only satisfactory when there are no drastic changes in the economy. Although it would be
useful to be able to predict these drastic changes accurately, both econometric and other
alternative methods tend to be imprecise.
Steps in developing an econometric model:
There are three main steps in developing an econometric model: specification, estimation and
validation.
While in a first approximation these stages follow a sequential order, in econometric analysis it is
generally necessary to go back more than once within this sequence. It is necessary to
continuously confront the model with the data and any other information source, in order to
obtain an econometric model compatible with the data. The model can be used to analyze reality,
offer better predictions or constitute a good basis for making decisions. Now we will describe the
steps listed above.
1. Specification
In this first step, the model or models used must be defined, as well as data to be used in the
estimation stage. In the specification step, we will refer to four elements: the economic model,
the econometric model, the statistical assumptions of the model and the data. In this section we
will refer to the first three elements; in the following section we will examine different types of
data used in econometric analysis. The first element we need is an economic model. In some
cases, a formal economic model is constructed entirely using economic theory. In other cases,
Page | 2
economic theory is used less formally in constructing an economic model. After we have an
economic model, we must convert it into an econometric model.
2. Estimation
In the estimation process we obtain numerical values of the coefficients of an econometric
model. To complete this stage, data are required on all observable variables that appear in the
specified econometric model, while it is also necessary to select the appropriate estimation
method, taking into account the implications of this choice on the statistical properties of
estimators of the coefficients. The distinction between estimator and estimate should be made
clear. An estimator is the result of applying an estimation method to an econometric
specification. On the other hand, an estimate consists of obtaining a numerical value of an
estimator for a given sample. For example, applying a very simple estimation method, called
ordinary least squares, to the specification of the consumption function (1-4) provides
expressions which determine the estimators β1and β2 . Substituting the sample data in these
expressions, two numbers are obtained: one for β1 and one for β2 which provide estimates of the
parameters β1 and β2. In general, it is possible to obtain analytical expressions of the estimators,
particularly in the case of estimating linear relationships. But in non-linear procedures of
estimation it is often difficult to establish their analytical expression.
3. Validation
The results are assessed in the validation stage, where we assess whether the estimates obtained
in the previous stage are acceptable, both theoretically and from the statistical point of view. On
the one hand, we analyze, whether estimates of model parameters have the expected signs and
magnitudes: that is to say, whether they satisfy the constraints established by economic theory.
From the statistical point of view, on the other hand, statistical tests are performed on the
significance of the parameters of the model, using the statistical assumptions made in the
specification step. In turn, it is important to test whether the statistical assumptions of the
econometric model are fulfilled, although it should be noted that not all assumptions are testable.
The violation of any of these assumptions implies, in general, the application of another
estimation method that allows us to obtain estimators whose statistical properties are as good as
possible.
One way to establish the ability of a model to make predictions is to use the model to forecast
outside the sample period, and then to compare the predicted values of the endogenous variable
with the values actually observed.
Economic data
As we have seen, an empirical analysis uses data to test a theory or to estimate a relationship. It
is important to stress that in Econometrics we use non-experimental data.
Non experimental or observational data are collected by observing the real world in a passive
way. In this case, data are not the outcome of controlled experiments.
Page | 3
Experimental data are often collected in laboratory environments in the same way as in natural
sciences. Now, we are going to see three types of data which can be used in the estimation of an
econometric model: time series, cross sectional data, and panel data.
1. Time Series Data
In time series, data are observations on a variable over time. For example: magnitudes from
national accounts such as consumption, imports, income, etc. The chronological ordering of
observations provides potentially important information. Consequently, ordering matters.
Time series data cannot be assumed to be independent across time. Most economic series are
related to their recent histories. Typical examples include macroeconomic aggregates such as
prices and interest rates. This type of data is characterized by serial dependence.
Given that most aggregated economic data are only available at a low frequency (annual,
quarterly or perhaps monthly), the sample size can be much smaller than in typical cross
sectional studies. The exception is financial data where data are available at a high frequency
(weekly, daily, hourly, etc.) and so sample sizes can be quite large
2. Cross Sectional Data
Cross sectional data sets have one observation per individual and data are referred to a
determined point in time. In most studies, the individuals surveyed are individuals (for example,
in the Labor Force Survey (EPA) more than 100000 individuals are interviewed every quarter),
households (for example, the Household Budget Survey), firms (for example, industrial firm
survey) or other economic agents. Surveys are a typical source for cross-sectional data. In many
contemporary econometric cross sectional studies the sample size is quite large.
In cross sectional data, observations must be obtained by random sampling. Thus, cross sectional
observations are mutually independent. The ordering of observations in cross sectional data does
not matter for econometric analysis. If the data are not obtained with a random sample, we have
a sample selection problem.
So far we have referred to micro data type, but there may also be cross sectional data relating to
aggregate units such as countries, regions, etc. Of course, data of this type are not obtained by
random sampling.
3. Pooled Cross Sections
Some data sets have both cross-sectional and time series features. For example, suppose that two
cross-sectional household surveys are taken in the United States, one in 1985 and one in 1990.
In 1985, a random sample of households is surveyed for variables such as income, savings,
family size, and so on. In 1990, a new random sample of households is taken using the same
survey questions. To increase our sample size, we can form a pooled cross section by combining
the two years.
Page | 4
Pooling cross sections from different years is often an effective way of analyzing the effects of a
new government policy. The idea is to collect data from the years before and after a key policy
change. As an example, consider the following data set on housing prices taken in 1993 and
1995, before and after a reduction in property taxes in 1994.
A pooled cross section is analyzed much like a standard cross section, except that we often need
to account for secular differences in the variables across the time. In fact, in addition to
increasing the sample size, the point of a pooled cross-sectional analysis is often to see how a
key relationship has changed over time.
4. Panel Data (or longitudinal data)
Panel data (or longitudinal data) are time series for each cross sectional member in a data set.
The key feature is that the same cross sectional units are followed over a given time period.
Panel data combines elements of cross sectional and time series data.
These data sets consist of a set of individuals (typically people, households, or corporations)
surveyed repeatedly over time. The common modeling assumption is that the individuals are
mutually independent of one another, but for a given individual, observations are mutually
dependent. Thus, the ordering in the cross section of a panel data set does not matter, but the
ordering in the time dimension matters a great deal. If we do not take into account the time in
panel data, we say that we are using pooled cross sectional data.
Write down simple and multiple linear regression models:
A linear regression model can be defined as the function approximation that represents a
continuous response variable as a function of one or more predictor variables. While building a
linear regression model, the goal is to identify a linear equation that best predicts or models the
relationship between the response or dependent variable and one or more predictor or
independent variables.
There are two different kinds of linear regression models. They are as follows:
Simple or Univariate linear regression models: These are linear regression models that are
used to build a linear relationship between one response or dependent variable and one predictor
or independent variable. The form of the equation that represents a simple linear regression
model is
Y=mX + b,
Where m is the coefficients of the predictor variable and b is bias. When considering the linear
regression line, m represents the slope and b represents the intercept.
In Simple Linear Regression (SLR), our goal is to predict the value of the dependent variable y
based on the independent variable x. We examine the relationship between these two variables
only. For example, we would utilize SLR to predict house prices based only on the square
footage of living. Though the relationship between x and y need not necessarily be linear, SLR
Page | 5
models include errors in the data, otherwise known as residuals. A residual is the difference
between the true value of y (blue dot) and the predicted value of y (red line). We can minimize
error by finding the “line of best fit”, otherwise called Ordinary Least Squares. To obtain this
line, we must take the sum of the difference between the true value of y and the predicted value
of y, square it, and take the minimum. We take the minimum because we are trying to minimize
the distance between the true and predicted value, measured by the black line in the visualization
below.
Multiple or Multi-variate linear regression models: These are linear regression models that are
used to build a linear relationship between one response or dependent variable and more than one
predictor or independent variable.
Regression models are used to describe relationships between variables by fitting a line to the
observed data. Regression allows you to estimate how a dependent variable changes as the
independent variable(s) change.
Multiple linear regression is used to estimate the relationship between two or more independent
variables and one dependent variable. You can use multiple linear regression when you want to
know:
1. How strong the relationship is between two or more independent variables and one
dependent variable (e.g. how rainfall, temperature, and amount of fertilizer added affect
crop growth).
2. The value of the dependent variable at a certain value of the independent variables (e.g.
the expected yield of a crop at certain levels of rainfall, temperature, and fertilizer
addition).
Assumptions of multiple linear regression
Multiple linear regression makes all of the same assumptions as simple linear regression:
Homogeneity of variance (homoscedasticity): the size of the error in our prediction
doesn’t change significantly across the values of the independent variable.
Page | 6
Independence of observations: the observations in the dataset were collected using
statistically valid methods, and there are no hidden relationships among variables.
In multiple linear regression, it is possible that some of the independent variables are
actually correlated with one another, so it is important to check these before developing
the regression model. If two independent variables are too highly correlated (r2 > ~0.6),
then only one of them should be used in the regression model.
Normality: The data follows a normal distribution.
Linearity: the line of best fit through the data points is a straight line, rather than a curve or some
sort of grouping factor.
Multiple linear regression formula
The formula for a multiple linear regression is:
Where,
y = the predicted value of the dependent variable
B0 = the y-intercept (value of y when all other parameters are set to 0)
B1X1 = the regression coefficient (B_1) of the first independent variable (X_1) (a.k.a. the effect
that increasing the value of the independent variable has on the predicted y value)
… = do the same for however many independent variables you are testing
BnXn = the regression coefficient of the last independent variable
Epsilon = model error (a.k.a. how much variation there is in our estimate of y)
Figure of multiple regression model:
Page | 7
To find the best-fit line for each independent variable, multiple linear regression calculates three
things:
1. The regression coefficients that lead to the smallest overall model error.
2. The t-statistic of the overall model.
3. The associated p-value (how likely it is that the t-statistic would have occurred by chance
if the null hypothesis of no relationship between the independent and dependent variables
was true).
It then calculates the t-statistic and p-value for each regression coefficient in the model.
References:
[Link]
[Link]
[Link]
%20Chapter%201%20%28Wooldridge%[Link]
[Link]
models-d2d5dbe9e704
[Link]
[Link]
different-results-7cf6c787766c
Page | 8