0% found this document useful (0 votes)
5 views11 pages

Module 2 - Introduction Chapter 1

The document serves as an introduction to econometrics, covering key concepts such as regression analysis, types of data, and experimental design. It emphasizes the importance of causal relationships in economic studies and discusses the limitations of randomized control trials and other study designs. The document outlines various data types, including cross-sectional, pooled cross-sectional, time series, and panel data, as well as basic mathematical tools used in econometrics.
Copyright
© All Rights Reserved
We take content rights seriously. If you suspect this is your content, claim it here.
Available Formats
Download as PDF, TXT or read online on Scribd
0% found this document useful (0 votes)
5 views11 pages

Module 2 - Introduction Chapter 1

The document serves as an introduction to econometrics, covering key concepts such as regression analysis, types of data, and experimental design. It emphasizes the importance of causal relationships in economic studies and discusses the limitations of randomized control trials and other study designs. The document outlines various data types, including cross-sectional, pooled cross-sectional, time series, and panel data, as well as basic mathematical tools used in econometrics.
Copyright
© All Rights Reserved
We take content rights seriously. If you suspect this is your content, claim it here.
Available Formats
Download as PDF, TXT or read online on Scribd

Lauren Hoehn Velasco Introduction to Econometrics Econometrics

Introduction to Econometrics
Instructor Notes

Contents

Contents 1

I Introduction to Econometrics 2

II Introduction to Simple Regression Analysis 5


A Why Regression? . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 5
B What does regression do? . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 5
C Key Parameters . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 6

III Types of Data 8

IV Mathematical and Statistical Tools 10


A Summations . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 10
B Estimators . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 11

Based upon materials from Introductory Econometrics by Jeffrey Wooldridge


© 2022 Lauren Hoehn Velasco
Please do not post or share without permission.

Page 1 of 11
Lauren Hoehn Velasco Introduction to Econometrics Econometrics

I. Introduction to Econometrics

Economists often would ideally like to utilize the scientific method to test theoretical questions with
empirical methods:
• Does education improve earnings? Does it matter if you choose to go to an elite college versus a
state school?
• Does minimum wage affect unemployment?
• Does access to health care improve health outcomes? Does exercise make you healthier?
• etc... etc... etc...
In an ideal world, we would walk through the
scientific method just like a scientist (see diagram
to right)

What is the ideal experimental design?

• Assigning randomized treatment and


control groups, allows you to observe the
’scientific’ counterfactual

What happens to be different/limited/difficult in economics?

• Difficult to design experiment in the economic realm


• Cannot observe the same individuals in both the counterfactual and reality.... we need methods
to deal with this!

Ideally, we would see both the real world, and the counterfactual world where the person/firm/econ-
omy did not get exposed to condition x. This is not ever the case, instead we move to the next type of
experimental design – the ’gold standard’ or the randomized trial. Where we have a group of people
and we assign treatment randomly.

Introduction to Econometrics continued on next page. . . Page 2 of 11


Lauren Hoehn Velasco Introduction to Econometrics Econometrics

Definition 1: Causality

How does variable one change if variable two is changed but all other relevant factors are held
constant (ceteris paribus)?

Notes:
• Most economic questions are ceteris paribus questions – ’holding all fixed’ – which is critical for
policy analysis
• It is important to define which causal effect one is interested in
• It is useful to describe how an experiment would have to be designed to infer the causal effect in
question

Definition 2: Randomized Control Trial


Randomized Control Trial is a study in which:
1. There are two groups, one treatment group and one control group. The treatment group
receives the treatment under investigation, and the control group receives either no treat-
ment (placebo) or standard treatment.
2. Patients are randomly assigned to all groups.

Notes:

This is the "‘Gold Standard"’ but there are problems will external validity and there are problems with
the ethics of RCTS.

External Validity: RCTs usually are costly and thus performed on a small sample, thus it can be difficult
to generalize to the entire population unless the sample is very carefully chosen (often not possible in
economics)

Ethics: It is difficult in economics to assign people to certain treatments without ethical concerns, i.e.
cannot randomly encourage someone to smoke, or perform other risky activities, but economists might
want to know the harm associated with these activities.

Cost: Can be prohibitive!

Introduction to Econometrics continued on next page. . . Page 3 of 11


Lauren Hoehn Velasco Introduction to Econometrics Econometrics

Other possible options for studies include:


1. Case Control Study

2. Longitudinal Study:

Do you see problems with these types of studies?

• Correlation does not imply causation – just because there is a difference between the two groups,
does not mean the point of interest is a causal link
• Selection bias – people select themselves into treatment, and thus the groups are not comparable
• Other issues of not even controlling for observables.... Would like to get at causation, but cannot
separate if the groups are different. We would like a way to control for the difference between
groups...

What do we do?
• Regression analysis allows us to do control for observable differences to make better comparisons
between groups.
• Observables – income, age, sex, race, etc...

Page 4 of 11
Lauren Hoehn Velasco Introduction to Econometrics Econometrics

II. Introduction to Simple Regression Analysis

A. Why Regression?

Regression is the primary tool in an econometrician’s toolbox

• It is used to find relationships between one or more explanatory variables and a single outcome
variable
• Distinct from correlations, as we trying to explan y using and x variable
• Best attempt in social science to determine effect of one variable on another, ’holding all else
fixed’ (think back to randomized control trial!)
Goal – predict or explain differences in values of the outcome variable with information about values of the ex-
planatory variables

Notes:

Regression is still not perfect


• It does a better job of controlling for differing characteristics of individuals
• Allows you to control for observed differences between two groups is due to the outcome of in-
terest

B. What does regression do?

The questions addressed by regression analysis take the form of:

‘What explains variation in Y?’


Or more specifically
‘Does X explain variation in Y?’

Before you even start to use regression, you have to decide what your X is and what your Y

Notes:

For micro work examples of each are:


• X often would be income, education, gender, race
• Y can be a quality such as wages, or other measurable outcomes

Introduction to Simple Regression Analysis continued on next page. . . Page 5 of 11


Lauren Hoehn Velasco Introduction to Econometrics Econometrics

C. Key Parameters

From there, primarily interested in the isolating following issues:


1. The direction and strength of the relationships between y and x(’s)
2. Which explanatory variables are practically (and statistically) important and which are not
3. Predicting a value or set of values of the outcome variable for a given set of values of the explana-
tory variables
4. The form of the relationship among the outcome and explanatory variables, or what the equation
that represents the relationship looks like ( we will use linear & log )

Notes:

Scatter-plot and Best Fit Line Using Stata: Linear regression analysis creates a linear best fit line, for
the scatter plot of two variables:
1. Open the gpa1 dataset: bcuse gpa1
2. Create a scatterplot with colGPA and hsGPA:

scatter colGPA hsGPA

3. Now add a best fit line, which map out the linear relationship between increasing hsGPA and its
affect on colGPA:

twoway scatter colGPA hsGPA || l f it colGPA hsGPA

Notes: The lfit adds a best fit line, the two-way tells Stata we want to display both graphs

Notes:

• From this we want to isolate the slope and the intercept, we want estimates of both (call β1 versus
β0
• Thus we need to write down a way to understand this relationship.

Introduction to Simple Regression Analysis continued on next page. . . Page 6 of 11


Lauren Hoehn Velasco Introduction to Econometrics Econometrics

Example 1: Experimental Design

What do we hold fixed in the following? How would you run an experiment to capture the effect?
1. “By how much will the production of soybeans increase if one increases the amount of
fertilizer applied to the ground”
2. “If a person is chosen from the population and given another year of education, by how
much will his or her wage increase?”
3. “If a city is randomly chosen and given ten additional police officers, by how much would
its crime rate fall?”
4. “How much will unemployment increase if the minimum wage is increased?”

Notes:
Assumptions:
1. Implicit assumption: all other factors that influence crop yield such as quality of land, rainfall,
presence of parasites etc. are held fixed
2. Implicit assumption: all other factors that influence wages such as experience, family background,
intelligence etc. are held fixed
3. If two cities are the same in all respects, except that city A has ten more police officers than city
B, by how much would the two cities’ crime rates differ?
4. By how much (if at all) will unemployment increase if the minimum wage is increased by a certain
amount (holding other things fixed)?
Experiment:
1. Choose several one-acre plots of land; randomly assign different amounts of fertilizer to the dif-
ferent plots; compare yields. Experiment works because amount of fertilizer applied is unrelated
to other factors influencing crop yields
2. Choose a group of people; randomly assign different amounts of education to them (infeasable!);
compare wage outcomes. Problem without random assignment: amount of education is related
to other factors that influence wages (e.g. intelligence)
3. Randomly assign number of police officers to a large number of cities. In reality, number of police
officers will be determined by crime rate (simultaneous determination of crime and number of
police)
4. Government randomly chooses minimum wage each year and observes unemployment outcomes.
Experiment will work because level of minimum wage is unrelated to other factors determining
unemployment. In reality, the level of the minimum wage will depend on political and economic
factors that also influence unemployment

Page 7 of 11
Lauren Hoehn Velasco Introduction to Econometrics Econometrics

III. Types of Data

Concept 1: Types of Data

Different kinds of economic data sets include:

1. Cross-sectional data 3. Pooled cross sections


2. Time series data 4. Panel/Longitudinal data

Notes:
Econometric methods depend on the nature of the data used. Use of inappropriate methods may lead
to misleading results.

Definition 3: Cross-Sectional Data


A sample of many subjects– individuals, households, firms, cities, states, countries – taken with-
out regard to time

Notes:
• Cross-sectional observations are more or less independent
• For example, pure random sampling from a population
• Sometimes pure random sampling is violated, e.g. units refuse to respond in surveys, or if sam-
pling is characterized by clustering
• Cross-sectional data typically encountered in applied microeconomics

Individual Data
Obs Age Education Experience
1 20 12 2
2 30 12 18
3 32 16 10
4 41 13 20
5 55 8 35

Types of Data continued on next page. . . Page 8 of 11


Lauren Hoehn Velasco Introduction to Econometrics Econometrics

Definition 4: Pooled Cross-Sectional Data


Two or more cross sections are combined in one data set

Notes:
• Cross sections are drawn independently of each other
• Pooled cross sections often used to evaluate policy changes
• Example:
– Evaluate effect of change in property taxes on house prices
– Random sample of house prices for the year 1993
– A new random sample of house prices for the year 1995
– Compare before/after (1993: before reform, 1995: after reform)

Definition 5: Time Series Data


Observations of a variable or several variables over time

Notes:
• For example, stock prices, money supply, consumer price index, gross domestic product, annual
homicide rates, automobile sales.
• Time series observations are typically serially correlated
• Ordering of observations conveys important information
• Data frequency: daily, weekly, monthly, quarterly, annually.
• Typical features of time series: trends and seasonality
• Typical applications: applied macroeconomics and finance

Minimum Wage and Unemployment Data


Obs Year Minimum Wage Unemployment
1 1950 0.2 15.4
2 1951 0.21 16
3 1952 0.23 14.8

Types of Data continued on next page. . . Page 9 of 11


Lauren Hoehn Velasco Introduction to Econometrics Econometrics

Definition 6: Panel Data


The same cross-sectional units are followed over time.

Notes:
• Panel data have a cross-sectional and a time series dimension
• Panel data can be used to account for time-invariant unobservables
• Panel data can be used to model lagged responses
• Example:
– City crime statistics; each city is observed in two years
– Time-invariant unobserved city characteristics may be modeled
– Effect of police on crime rates may exhibit time lag
• Example with previous cross-section turned into panel:

Panel Data
Obs Year Age Education Experience
1 1 25 12 2
1 2 26 12 3
2 1 35 12 15
2 2 36 12 16
3 1 21 15 0
3 2 22 16 0

IV. Mathematical and Statistical Tools

A. Summations

Definition 7: Summation Operator

The summation operator is a useful shorthand for manipulating expressions involving sums of
many numbers:
n
∑ xi = x1 + x2 + · · · + xn−1 + xn
i=1

Notes:
More generally:
n
∑ ai = am + am+1 + am+2 + · · · + an−1 + an
i=m

Mathematical and Statistical Tools continued on next page. . . Page 10 of 11


Lauren Hoehn Velasco Introduction to Econometrics Econometrics

Definition 8: Properties of Summations

For any constants a, b, c:


1.
n
∑ c = nc
i=1
2.
n n
∑ cxi = c ∑ xi
i=1 i=1
3.
n n n
∑ (axi + byi ) = a ∑ xi + b ∑ yi
i=1 i=1 i=1

Notes:
We will use these properties in Chapter 2.

B. Estimators

Definition 9: Estimators
Given a random sample, drawn from the population distribution, that depends on an unknown
parameter θ , an estimator of θ is a rule that assigns each possible outcome of the sample a value
of θ .

Notes:
An example of an estimator is the sample mean, which estimates the population parameter µ.

1 n
X̄ = ∑ Xi
n i=1

For an actual set of data would get x̄ = 1n (x1 + ... + xn ).

Definition 10: Unbiased Estimator


An estimator W of θ is unbiased if
E(W ) = θ

for all values of θ .

Notes:
We can show that the sample mean is unbiased:

1 n 1 n
1 n
E(X̄) = E( ∑ Xi ) = E( ∑ Xi ) = ∑ E(Xi ) = n(1/n)µ = µ
n i=1 n i=1 n i=1

Page 11 of 11

You might also like