0% found this document useful (0 votes)
16 views11 pages

Understanding Econometrics Basics

Econometrics is the application of mathematical statistics to economic data, aimed at providing empirical support to economic theories and measuring economic relationships. It is a distinct discipline that combines economic theory, mathematical economics, and statistical inference, focusing on empirical verification and the estimation of economic models. The methodology of econometrics involves several steps, including hypothesis formulation, model specification, data collection, parameter estimation, hypothesis testing, and forecasting for policy-making.
Copyright
© All Rights Reserved
We take content rights seriously. If you suspect this is your content, claim it here.
Available Formats
Download as PDF, TXT or read online on Scribd
0% found this document useful (0 votes)
16 views11 pages

Understanding Econometrics Basics

Econometrics is the application of mathematical statistics to economic data, aimed at providing empirical support to economic theories and measuring economic relationships. It is a distinct discipline that combines economic theory, mathematical economics, and statistical inference, focusing on empirical verification and the estimation of economic models. The methodology of econometrics involves several steps, including hypothesis formulation, model specification, data collection, parameter estimation, hypothesis testing, and forecasting for policy-making.
Copyright
© All Rights Reserved
We take content rights seriously. If you suspect this is your content, claim it here.
Available Formats
Download as PDF, TXT or read online on Scribd

Chapter One

Introduction

1.1 What is Econometrics?

Literally interpreted, econometrics means “economic measurement or measurement in


economics.” Although measurement is an important part of econometrics, the scope of
econometrics is much broader.

Econometrics, the result of a certain outlook on the role of economics, consists of the application
of mathematical statistics to economic data to lend empirical support to the models constructed by
mathematical economics and to obtain numerical results. Econometrics may be defined as the
quantitative analysis of actual economic phenomena based on the concurrent development of
theory and observation, related by appropriate methods of inference. It may be defined as the
social science in which the tools of economic theory, mathematics, and statistical inference are
applied to the analysis of economic phenomena.

Econometrics is based upon the development of statistical methods for estimating economic
relationships, testing economic theories, and evaluating and implementing government and
business policy. Econometrics is the science that integrates economic theory, economic statistics,
and mathematical economics to investigate the empirical support of the general schematic law
established by economic theory. It is a special type of economic analysis and research in which the
general economic theories, formulated in mathematical terms, is combined with empirical
measurements of economic phenomena. Starting from the relationships of economic theory, we
express them in mathematical terms so that they can be measured. We then use specific methods,
called econometric methods in order to obtain numerical estimates of the coefficients of the
economic relationships.

1.2 Why a Separate Discipline?

Econometrics is a separate discipline. Why? Econometrics is an amalgam of economic theory,


mathematical economics, economic statistics, and mathematical statistics. Econometrics differs
from mathematical economics and statistics. Yet the subject deserves to be studied in its own right
for the following reasons.

Page | 1
Economic theory makes statements or hypotheses that are mostly qualitative in nature. For
example, microeconomic theory states that, other things remaining the same, a reduction in the
price of a commodity is expected to increase the quantity demanded of that commodity. Thus,
economic theory postulates a negative or inverse relationship between the price and quantity
demanded of a commodity. But the theory itself does not provide any numerical measure of the
relationship between the two; that is, it does not tell by how much the quantity will go up or down
as a result of a certain change in the price of the commodity. It is the job of the econometrician to
provide such numerical estimates. Stated differently, econometrics gives empirical content to
economic theories.

The main concern of mathematical economics is to express economic theory in mathematical form
(equations) without regard to measurability or empirical verification of the theory. Econometrics,
as noted previously, is mainly interested in the empirical verification of economic theory. The
econometrician often uses the mathematical equations proposed by the mathematical economist
but puts these equations in such a form that they lend themselves to empirical testing. And this
conversion of mathematical into econometric equations requires a great deal of ingenuity and
practical skill. Although econometrics presupposes the expression of economic relationships in
mathematical form, like mathematical economics, it does not assume that economic relationships
that are exact. On the contrary, econometrics assumes that economic relationships are not exact
but stochastic. Econometric methods are designed to take into account random disturbances which
create deviations from exact behavioural patterns suggested by economic theory and mathematical
economics. Econometric methods are designed in such a way that they take into account the
random disturbances.

Economic statistics is mainly concerned with collecting, processing, and presenting economic data
in the form of charts, graphs, and tables. These are the jobs of the economic statistician. It is he or
she who is primarily responsible for collecting data on gross national product (GNP), employment,
unemployment, prices, and so on. The data thus collected constitute the raw data for econometric
work. But the economic statistician does not go any further, not being concerned with using the
collected data to test economic theories. Of course, one who does that becomes an econometrician.

Page | 2
On the contrary, mathematical statistics deals with methods of measurement which are developed
on the basis of controlled experiments in laboratories. Statistical methods of measurement are not
appropriate for economic relationships, which cannot be measured on the basis of evidence
provided by controlled experiments, because such experiments cannot be designed for economic
phenomena.

1.3 Methodology of Econometrics

Broadly speaking, traditional econometric methodology proceeds along the following lines:

i) Statement of theory or hypothesis.


ii) Specification of the mathematical model of the theory.
iii) Specification of the statistical or econometric model.
iv) Obtaining the data.
v) Estimation of the parameters of the econometric model.
vi) Hypothesis testing.
vii) Forecasting or prediction.
viii) Using the model for control or policy purposes.

To illustrate the preceding steps, let us consider the well-known Keynesian theory of consumption.

i) Statement of Theory or Hypothesis

This step involves formulating an economic theory or hypothesis that describes how certain
economic variables are related. John Maynard Keynes stated: The fundamental psychological law
...is that men [women] are disposed, as a rule and on average, to increase their consumption as
their income increases, but not as much as the increase in their income. In short, Keynes postulated
that the marginal propensity to consume (MPC), the rate of change of consumption for a unit (say,
a dollar) change in income, is greater than zero but less than 1.

ii) Specification of the Mathematical Model

Although Keynes postulated a positive relationship between consumption and income, he did not
specify the precise form of the functional relationship between the two. For simplicity, a
mathematical economist might suggest the following form of the Keynesian consumption function:

Page | 3
where Y = consumption expenditure and X = income, and where β1 and β2, known as the
parameters of the model, are, respectively, the intercept and slope coefficients. The slope
coefficient β2 measures the MPC. The variable appearing on the left side of the equality sign is
called the dependent variable, and the variable(s) on the right side is called the independent, or
explanatory variable(s).

iii) Specification of the Econometric Model

The purely mathematical model assumes an exact or deterministic relationship between the
dependent variable and independent variable(s). But relationships between economic variables are
generally inexact.

To allow for the inexact relationships between variables, the econometrician would modify the
deterministic mathematical model (the consumption function in our case) as follows:

This is an example of an econometric model where u, known as the disturbance or error term, is a
random (stochastic) variable that has well-defined probabilistic properties. The disturbance term u
may well represent all those factors that affect the dependent variable but are not taken into account
explicitly.

In this step, the econometrician has to express the relationships between variables in mathematical
form. This step involves the determination of three important tasks:

a) the dependent and independent (explanatory) variables which will be included in the
model.
b) the a priori theoretical expectations about the size and sign of the parameters of the
function.
c) the mathematical form of the model (number of equations, specific form of the equations,
etc.)

The specification of the econometric model will be based on economic theory and on any available
information related to the phenomena under investigation. Thus, specification of the econometric

Page | 4
model presupposes knowledge of economic theory and familiarity with the particular phenomenon
being studied.

The specification of the model is the most important and the most difficult stage of any
econometric research. It is often the weakest point of most econometric applications. In this stage,
there exists enormous degree of likelihood of committing errors or incorrectly specifying the
model.

The most common errors of specification are:

• Omissions of some important variables from the function.


• The omissions of some equations (for example, in simultaneous equations model).
• The mistaken mathematical form of the functions.

iv) Obtaining Data

To estimate the econometric model, that is, to obtain the numerical values of β1 and β2, we need
data.

v) Estimation of the Econometric Model

Now that we have the data, our next task is to estimate the parameters of the model. This is purely
a technical stage which requires knowledge of the various econometric methods, their assumptions
and the economic implications for the estimates of the parameters.

This stage includes the following activities.

a) Gathering of the data on the variables included in the model.


b) Examination of the identification conditions of the function (especially for
simultaneous equations models).
c) Examination of the aggregation problems involved in the variables of the function.
d) Examination of the degree of correlation between the explanatory variables (i.e.,
examination of the problem of multicollinearity).
e) Choice of appropriate economic techniques for estimation, i.e., to decide a specific
econometric method to be applied in estimation, such as OLS, MLM, Logit, and Probit.

Page | 5
vi) Hypothesis Testing

Assuming that the model is a reasonably good approximation of reality, we have to develop
suitable criteria to find out whether the estimates obtained are in accord with the expectations of
the theory that is being tested.

vii) Forecasting or Prediction

Forecasting is one of the aims of econometric research. If the chosen model does not refute the
hypothesis or theory under consideration, we may use it to predict the future value(s) of the
dependent, or forecast variable Y based on the known or expected future value(s) of the
explanatory or predictor variable X.

Forecasting depends heavily on the quality of the estimated model. A model that is reliable, valid,
and statistically sound provides forecasts that are credible and useful for decision-making.
Conversely, a poorly specified or unstable model can lead to serious forecasting errors and wrong
policy conclusions. Before using an estimated econometric model for forecasting, we must ensure
the model is reliable, valid, and statistically sound. If the model is poorly specified or unstable, the
forecasts will be misleading.

viii) Use of the Model for Control or Policy Purposes

It involves applying the estimated model to guide decision-making, policy formulation, and
economic planning. Once an econometric model has been developed, estimated, and tested for
reliability, it can serve as a practical tool for control and policy analysis. Thus, an estimated model
may be used for control or policy purposes.

1.4 Significance of Stochastic Disturbance Term

The disturbance term is a surrogate for all those variables that are omitted from the model but that
collectively affect Y. The reasons for introducing the stochastic disturbance term 𝑈𝑖 are as follows:

i) Vagueness of theory: The theory, if any, determining the behavior of the dependent
variable (Y) may be, and often is, incomplete. We might know for certain that weekly
income X influences weekly consumption expenditure Y, but we might be ignorant or
unsure about the other variables affecting Y. Therefore, 𝑈𝑖 may be used as a substitute
for all the excluded or omitted independent variables from the model.
Page | 6
ii) Unavailability of data: Even if we know what some of the excluded independent
variables are, we may not have data for these variables. It is a common experience in
empirical analysis that the data we would ideally like to have often are not available.
For example, in principle, we could introduce family wealth as an explanatory variable
in addition to the income variable to explain family consumption expenditure. But
unfortunately, information on family wealth generally is not available. Therefore, we
may be forced to omit the wealth variable from our model despite its great theoretical
relevance in explaining consumption expenditure.
iii) Core variables versus peripheral variables: Assume in our consumption-income
example that besides income X1, the number of children per family X2, sex X3, religion
X4, education X5, and geographical region X6 also affect consumption expenditure. But
it is quite possible that the joint influence of all or some of these variables may be so
small and at best nonsystematic or random that as a practical matter and for cost
considerations it does not pay to introduce them into the model explicitly. One hopes
that their combined effect can be treated as a random variable 𝑈𝑖.
iv) Poor proxy variables: Although the classical regression model assumes that the
variables Y and X are measured accurately, in practice, the data may be plagued by
errors of measurement. Consider, for example, Keynes's well-known theory of the
Psychological law of consumption function regards consumption expenditure (Yp) as
a function of income (Xp). But since data on these variables are not directly observable,
in practice we use proxy variables, such as current consumption expenditure (Y) and
current income (X), which can be observable. Since the observed Y and X may not
equal Yp and Xp, there is the problem of errors of measurement. The disturbance term
U may in this case then also represent the errors of measurement. If there are such
errors of measurement, they can have serious implications for estimating the regression
coefficients.
v) Principle of parsimony: If we can explain the behavior of Y "substantially" with two
or three explanatory variables and if our theory is not strong enough to suggest what
other variables might be included, why introduce more variables? Let 𝑈𝑖 represent all
other variables. Of course, we should not exclude relevant and important variables just
to keep the regression model simple.

Page | 7
1.5 Goals of Econometrics

The three main goals of econometrics are as follows:

i) Analysis: Econometrics primarily aims at the verification or testing of economic


theories. In this case, we say that the purpose of the research is analysis. That is, the
economic models are formulated in an empirically testable form to decide how well
they explain the observed behavior of the economic units. Several econometric models
can be derived from an economic model. Such models differ due to different choices
of functional form, specification of stochastic structure of the variables, etc. So, a strong
analysis will be carried out by econometrics as a prime goal to verify any economic
theory and economic phenomena.
ii) Policy Making: The models are estimated on the basis of observed set of data and are
tested for their suitability. This is the part of statistical inference of the modeling.
Various estimation procedures are used to know the numerical values of the unknown
parameters of the model. Based on various formulations of statistical models, a suitable
and appropriate model is selected. The inference or the knowledge obtained from the
numerical value of the coefficients is important for decision-making as well as the
formulation of economic policies. It helps to compare the effects of alternative policy
decisions.
iii) Forecasting: The obtained models are used for forecasting and policy formulation
which is an essential part in any policy decision. Such forecasts help the policy makers
to judge the goodness of the fitted model and take necessary measures in order to re-
adjust the relevant economic variables.

1.6 Types of Data

The success of any econometric analysis ultimately depends on the availability of the appropriate
data. Three types of data may be available for empirical analysis: time series, cross-sectional, and
pooled (i.e., combination of time series and cross-sectional) data.

Page | 8
Cross-Sectional Data

Cross-sectional data are data on one or more variables collected at the same point in time.
Sometimes, the data on all units do not correspond to precisely the same time period. For example,
several families may be surveyed during different weeks within a year. In a pure cross-sectional
analysis, we would ignore any minor timing differences in collecting the data. If a set of families
was surveyed during different weeks of the same year, we would still view this as a cross-sectional
data set.

Examples:
• Income levels of 500 households in Addis Ababa in 2024.
• Agricultural productivity of 100 farms in a Woreda in 2022.
• Pollution levels across 50 Ethiopian cities in 2023.
• Wages of workers across different industries in Ethiopia in 2024.

Time Series Data

Time series data consist of observations of one variable or several variables collected for a single
entity (country, firm, household, company, market, etc.) over multiple time periods (e.g., monthly,
quarterly, yearly). A time series is a set of observations on the values that a variable or several
variables take at different times. Such data may be collected at regular time intervals, such as daily,
weekly, monthly, annually, or decennially, that is, every 10 years. Sometimes data are available
both quarterly as well as annually.

A time series data set consists of observations on a variable or several variables over time. Because
past events can influence future events and lags in behavior are prevalent in the social sciences,
time is an important dimension in a time series data set. Unlike the arrangement of cross-sectional
data, the chronological ordering of observations in a time series conveys potentially important
information.

Examples:
• Annual GDP of Ethiopia from 1990 – 2024.
• Monthly inflation rate in Ethiopia from 2010 – 2024.
• Daily exchange rate between Ethiopian Birr and USD from 2018 – 2024.
• Yearly grain crop production in Ethiopia from 1992 – 2024.

Page | 9
Pooled Cross-sectional Data

Some data sets have both cross-sectional and time series features. In typical pooled cross-sectional
data, the observations in each time period are different units (e.g., new households surveyed each
year). However, it is possible that some units reappear across periods — by coincidence or design
— but this does not make the data panel unless those units are systematically tracked over time.

For example, suppose that two cross-sectional household surveys are taken in Ethiopia, one in
2010 and one in 2015. In 2010, a random sample of households is surveyed for variables such as
income, savings, family size, and so on. In 2015, a new random sample of households is taken
using the same survey questions. To increase our sample size, we can form a pooled cross-section
by combining the two years.

Other examples
• Annual labor market survey (2010 – 2024) interviewing different workers each year to
analyze changes in unemployment or wage rates.
• Data on 10 different countries in 2000, another 15 in 2010, and another 20 in 2020 — not
necessarily the same countries each year.
• The Ministry of Trade collects data annually (2015 – 2024) from different manufacturing
firms each year to examine trends in investment or employment.
• Data on different sets of African countries each year (2010 – 2024) to study factors
affecting annual GDP growth or inflation trends.

Panel or Longitudinal Data

This is a special type of pooled data in which the same cross-sectional units are surveyed over
time. A panel data (longitudinal data) set consists of a time series for each cross-sectional member
in the data set. This is a special type of pooled data in which the same cross-sectional unit (say, a
family or a firm) is surveyed over time. The key feature of panel data that distinguishes them from
a pooled cross-section is that the same cross-sectional units (individuals, firms, or counties) are
followed over a given time period.

Page | 10
Example:

• Tracking the same 500 households from 2010 – 2024 to examine how changes in income
affect consumption.
• Annual data for the same 40 African countries from 1990 – 2024, including GDP, inflation,
trade balance, and exchange rates.
• Annual data from the same 200 manufacturing firms in Ethiopia from 2010 – 2024 to assess
how capital investment influences productivity growth.

Page | 11

Common questions

Powered by AI

In econometric analysis, the type of data used—time series, cross-sectional, or pooled—significantly impacts the analysis outcome. Time series data captures the temporal dynamics of a single variable or several variables over time, offering insights into trends and relationships subject to time lags . Cross-sectional data provides a snapshot of several units at a single point, useful for comparative analysis across subjects . Pooled data combines time series and cross-sectional data, enhancing sample size and variability, thus improving the estimative power of models . Understanding these data types helps in choosing the correct analytical approach and model specification, critical for accurate empirical research .

Economic theories often lack completeness, being unable to account for all factors influencing a dependent variable. The stochastic disturbance term 'u' compensates for this vagueness by capturing the effects of omitted variables that are not explicitly modeled but nonetheless affect the outcome . It acknowledges that economic behavior is influenced by numerous unforeseen factors, providing a more realistic representation of economic phenomena than would be possible with purely deterministic models. The inclusion of 'u' ensures that econometric models remain flexible and closer to empirical realities .

The primary challenges in model specification include determining the right variables to include, the appropriate functional form, and addressing potential errors like omitted variable bias or multicollinearity . These challenges impact analysis results significantly, as improper specification can lead to biased, inconsistent, or inefficient parameter estimates, ultimately affecting the model's predictive validity and the reliability of hypothesis testing . Ensuring that the model aligns with economic theory and empirical observations is crucial to mitigate these risks and produce credible econometric analysis .

The methodology of econometrics follows a structured framework that begins with the statement of the economic theory or hypothesis. This is followed by the specification of a mathematical model, which is translated into an econometric model that includes a stochastic disturbance term to account for inexact relationships . The next steps involve data collection, parameter estimation, hypothesis testing, forecasting, and application for policy control . This systematic approach ensures that economic theories are not only expressed in mathematical terms but are also empirically tested, allowing for the validation or refutation of theoretical predictions based on real-world data .

Selecting an appropriate econometric method, such as OLS, MLM, Logit, or Probit, is crucial as it influences the reliability and validity of the parameter estimates . Challenges include identifying the correct model specification, avoiding multicollinearity, ensuring data quality, and meeting the assumptions of the chosen method. Errors in model specification, such as omitting important variables or incorrect functional forms, can lead to biased and inconsistent estimates . The choice of method must also consider the nature of the data and the underlying economic theory .

The stochastic disturbance term, often denoted as 'u' in econometric models, accounts for the deviations between observed economic relationships and their theoretical counterparts, which are generally inexact. It serves as a substitute for all omitted variables that affect the dependent variable but are not explicitly included in the model . This term addresses the inherent variability in economic data that mathematical economics cannot capture due to its deterministic nature, and which are instead treated as stochastic in econometrics to allow for randomness and uncertainty .

When using an econometric model for forecasting, it must first be ensured that the model is properly specified, reliable, and statistically sound. Considerations include the model's ability to accurately represent the underlying data dynamics, account for stochastic variations, and avoid biases from omitted variables or incorrect assumptions . The quality of predictions heavily depends on the accuracy of the estimated relationships and the stability of the model over time. Inaccurate forecasts may arise from models that are poorly specified or unstable, leading to misleading conclusions and ineffective policy directions .

Econometric methods address multicollinearity, a situation where explanatory variables are highly correlated, through techniques such as ridge regression or principal component analysis, which aim to mitigate its effects . Multicollinearity inflates the variances of parameter estimates, potentially leading to statistically insignificant results despite strong relationships . By addressing this issue, econometricians can more accurately estimate the true impact of each predictor variable, thus enhancing the model's explanatory power and forecast accuracy. This is significant for making reliable inferences and policy recommendations based on the estimated model .

The specification of an econometric model involves extending a mathematical model by including a stochastic disturbance term to account for inexact relationships between variables. While a mathematical model assumes an exact, deterministic relationship, the econometric model allows for random deviations due to omitted variables or unobserved factors . This stage is challenging because it requires determining the appropriate dependent and independent variables, formulating a priori expectations, and selecting the right mathematical form, all of which heavily rely on sound economic theory and empirical insights. Errors in this stage, such as incorrect variable selection or functional form, can severely impact the model's reliability .

Econometrics plays a crucial role in policy-making by providing a framework for estimating and testing models based on observed data to inform economic policies. By deriving numerical values for model parameters, econometrics helps compare the effects of alternative policy decisions and provides empirical evidence to support policy formulation. This statistical inference is vital for decision-makers to assess the potential impact of policies, forecast economic trends, and adjust strategies as needed .

You might also like