Chapter 1
Chapter 1
1 The Nature of
Econometrics and
Economic Data
Learning Objectives
1.1 Describe how econometric methods are used to conduct convincing data analysis in business,
economics, and other social sciences.
1.2 Identify, and describe the features of, the most common types of data sets used in economics and
related fields: cross-sectional data, time-series data, pooled (or repeated) cross sections, panel (or
longitudinal) data.
1.3 Apply the principle of causality using examples from policy analysis and business decision making.
1.4 List the challenges in trying to infer causality when one only has access to retrospective,
nonexperimental data.
1.5 Explain the ideas of “ceteris paribus” and counterfactual reasoning when considering problems of
causal inference in economics.
Chapter 1 discusses the scope of econometrics and raises general issues that arise in the application of econo-
metric methods. Section 1-1 provides a brief discussion about the purpose and scope of econometrics and how
it fits into economic analysis. Section 1-2 provides examples of how one can start with an economic theory
and build a model that can be estimated using data. Section 1-3 examines the kinds of data sets that are used in
business, economics, and other social sciences. Section 1-4 provides an intuitive discussion of the difficulties
associated with inferring causality in the social sciences.
Copyright 2025 Cengage Learning. All Rights Reserved. May not be copied, scanned, or duplicated, in whole or in part. Due to electronic rights, some third party content may be suppressed from the eBook and/or eChapter(s).
Editorial review has deemed that any suppressed content does not materially affect the overall learning experience. Cengage Learning reserves the right to remove additional content at any time if subsequent rights restrictions require it.
Econometrics is based upon the development of statistical methods for estimating economic relationships,
testing economic theories, and evaluating and implementing government and business policy. A common
application of econometrics is the forecasting of such important macroeconomic variables as interest rates,
inflation rates, and gross domestic product (GDP). Of course, forecasting is important in many other areas
of social policy, such as trying to predict outcomes of elections or rates of infection of a new virus. Whereas
forecasts of economic indicators are highly visible and often widely published, econometric methods can be
used in economic areas that have nothing to do with macroeconomic forecasting. For example, we will study
the effects of political campaign expenditures on voting outcomes. We will consider the effect of school
spending on student performance in the field of education. In addition, we will learn how to use econometric
methods for forecasting economic time series.
Econometrics has evolved as a separate discipline from mathematical statistics because the former focuses
on the problems inherent in collecting and analyzing nonexperimental economic data. Nonexperimental data
are not accumulated through controlled experiments on individuals, firms, or segments of the economy.
(Nonexperimental data are sometimes called observational data, or retrospective data, to emphasize the fact
that the researcher is a passive collector of the data.) Experimental data are often collected in laboratory envi-
ronments in the natural sciences, however, they are more difficult to obtain in the social sciences. Although
some social experiments can be devised, it is often impossible, prohibitively expensive, or morally repugnant
to conduct the kinds of controlled experiments that would be needed to address economic issues. We give
some specific examples of the differences between experimental and nonexperimental data in Section 1-4.
Naturally, econometricians have borrowed from mathematical statisticians whenever possible. The
method of multiple regression analysis is the mainstay in both fields, but its focus and interpretation can differ
markedly. In addition, economists have devised new techniques to deal with the complexities of economic
data and to test the predictions of economic theories.
Copyright 2025 Cengage Learning. All Rights Reserved. May not be copied, scanned, or duplicated, in whole or in part. Due to electronic rights, some third party content may be suppressed from the eBook and/or eChapter(s).
Editorial review has deemed that any suppressed content does not materially affect the overall learning experience. Cengage Learning reserves the right to remove additional content at any time if subsequent rights restrictions require it.
Formal economic modeling is sometimes the starting point for empirical analysis, but it is more common
to use economic theory less formally, or even to rely entirely on intuition. You may agree that the determinants
of criminal behavior appearing in equation (1.1) are reasonable based on common sense; we might arrive at
such an equation directly, without starting from utility maximization. This view has some merit, although
there are cases in which formal derivations provide insights that intuition can overlook.
Next is an example of an equation that we can derive through somewhat informal reasoning.
Copyright 2025 Cengage Learning. All Rights Reserved. May not be copied, scanned, or duplicated, in whole or in part. Due to electronic rights, some third party content may be suppressed from the eBook and/or eChapter(s).
Editorial review has deemed that any suppressed content does not materially affect the overall learning experience. Cengage Learning reserves the right to remove additional content at any time if subsequent rights restrictions require it.
where
wage = hourly wage,
educ = years of formal education,
exper = years of workforce experience, and
training = weeks spent in job training.
Again, other factors generally affect the wage rate, but equation (1.2) captures the essence of the problem.
After we specify an economic model, we need to turn it into what we call an econometric model. Because
we will deal with econometric models throughout this text, it is important to know how an econometric model
relates to an economic model. Take equation (1.1) as an example. The form of the function f (·) must be spec-
ified before we can undertake an econometric analysis. A second issue concerning (1.1) is how to deal with
variables that cannot reasonably be observed. For example, consider the wage that a person can earn in crim-
inal activity. In principle, such a quantity is well defined, but it would be difficult if not impossible to observe
this wage for a given individual. Even variables such as the probability of being arrested cannot realistically
be obtained for a given individual, but at least we can observe relevant arrest statistics and derive a variable
that approximates the probability of arrest. Many other factors affect criminal behavior that we cannot even
list, let alone observe. We must somehow account for them.
The ambiguities inherent in the economic model of crime are resolved by specifying a particular econo-
metric model:
crime = b 0 + b1 wage + b 2 othinc + b 3 freqarr + b 4 freqconv + b 5 avgsen + b 6 age + u, [1.3]
where
Copyright 2025 Cengage Learning. All Rights Reserved. May not be copied, scanned, or duplicated, in whole or in part. Due to electronic rights, some third party content may be suppressed from the eBook and/or eChapter(s).
Editorial review has deemed that any suppressed content does not materially affect the overall learning experience. Cengage Learning reserves the right to remove additional content at any time if subsequent rights restrictions require it.
merge any underlying economic theory into the econometric model specification. In the economic model of
crime example, we would start with an econometric model such as (1.3) and use economic reasoning and
common sense as guides for choosing the variables. Although this approach loses some of the richness of
economic analysis, it is commonly and effectively applied by careful researchers.
Once an econometric model such as (1.3) or (1.4) has been specified, various hypotheses of interest can be
stated in terms of the unknown parameters. For example, in equation (1.3), we might hypothesize that wage,
the wage that can be earned in legal employment, has no effect on criminal behavior. In the context of this
particular econometric model, the hypothesis is equivalent to b1 = 0.
An empirical analysis, by definition, requires data. After data on the relevant variables have been col-
lected, econometric methods are used to estimate the parameters in the econometric model and to formally
test hypotheses of interest. In some cases, the econometric model is used to make predictions in either the
testing of a theory or the study of a policy’s impact.
Because data collection is so important in empirical work, Section 1-3 will describe the kinds of data that
we are likely to encounter.
Copyright 2025 Cengage Learning. All Rights Reserved. May not be copied, scanned, or duplicated, in whole or in part. Due to electronic rights, some third party content may be suppressed from the eBook and/or eChapter(s).
Editorial review has deemed that any suppressed content does not materially affect the overall learning experience. Cengage Learning reserves the right to remove additional content at any time if subsequent rights restrictions require it.
first sample schools. In Chapter 14, we will discuss the special methods of statistical inference that arise in
the context of cluster sampling.
A more subtle issue arises when a policy that we are analyzing is determined at a higher level than the
unit we sample. For example, a policy may apply at the school level, and then all students within a school are
subjected to the same intervention (such as more spending at the school level). In these cases, even though
we may have a random sample of students from a population—we might even observe the entire population
of students—the policy variable will be correlated within a school. This is true even if the policy is assigned
differently to students within a school as long as the policy is initially determined at the school level. For
example, perhaps some schools are given extra teachers so that some class sizes can be made smaller. Within
a school, class size is made smaller for some students but not necessarily all students within the same school.
An introduction to inference methods appropriate to such situations is provided in Chapter 14. Fortunately,
all of the estimation methods we study up until that point can be applied in these more complicated settings.
A related problem is when using data on relatively large geographic units. For example, if we want to
explain new business activity across counties as a function of wage rates, energy prices, corporate and property
tax rates, services provided, quality of the workforce, and other county characteristics, we must acknowledge
the possibility that decisions on how to set, say, corporate tax rates will be correlated across nearby counties.
Moreover, a policy change in one county—perhaps changing a minimum wage, a property tax rate, and so
on—could have spillover effects on surrounding counties. Modeling such spillover effects is an important
topic in advanced econometrics. In this introductory text, we will largely ignore the intricacies that arise in
analyzing such situations until Chapter 14.
Cross-sectional data are widely used in economics and other social sciences. In economics, the analysis of
cross-sectional data is closely aligned with the applied microeconomics fields, such as labor economics, state
and local public finance, industrial organization, urban economics, demography, and health economics. Data
on individuals, households, firms, and cities at a given point in time are important for testing microeconomic
hypotheses and evaluating economic policies.
The cross-sectional data used for econometric analysis can be represented and stored in computers.
Table!1.1 contains, in abbreviated form, a cross-sectional data set on 526 working individuals for the year
1976. (This is a subset of the data in the file WAGE1.) The variables include wage (in dollars per hour), which
can be roughly turned into 2022 dollars by multiplying by four; educ (years of education); exper (years of
potential labor force experience); female (an indicator for gender); and married (marital status). These last
two variables are binary (zero-one) in nature and serve to indicate qualitative features of the individual. The
variable married indicates whether a person is married or not. In this data set, the variable female is equal to
one if the respondent identifies as female and zero otherwise. We will have much to say about binary variables
in Chapter 7 and beyond.
Table 1.1 A Cross-Sectional Data Set on Wages and Other Individual Characteristics
obsno wage educ exper female married
1 3.10 11 2 1 0
2 3.24 12 22 1 1
3 3.00 11 2 0 0
4 6.00 8 44 0 1
5 5.30 12 7 0 1
. . . . . .
. . . . . .
. . . . . .
525 11.56 16 5 0 1
526 3.50 14 5 1 0
Copyright 2025 Cengage Learning. All Rights Reserved. May not be copied, scanned, or duplicated, in whole or in part. Due to electronic rights, some third party content may be suppressed from the eBook and/or eChapter(s).
Editorial review has deemed that any suppressed content does not materially affect the overall learning experience. Cengage Learning reserves the right to remove additional content at any time if subsequent rights restrictions require it.
The variable obsno in Table 1.1 is the observation number assigned to each person in the sample. Unlike
the other variables, it is not a characteristic of the individual. All econometrics and statistics software pack-
ages assign an observation number to each data unit. Intuition should tell you that, for data such as that in
Table 1.1, it does not matter which person is labeled as observation 1, which person is labeled observation 2,
and so on. The fact that the ordering of the data does not matter for econometric analysis is a key feature of
cross-sectional data sets obtained from random sampling.
Different variables sometimes correspond to different time periods in cross-sectional data sets. For exam-
ple, to determine the effects of government policies on long-term economic growth, economists have studied
the relationship between growth in real per capita GDP over a certain period (say, 1960 to 1985) and variables
determined in part by government policy in 1960 (government consumption as a percentage of GDP and adult
secondary education rates). Such a data set might be represented as in Table 1.2, which constitutes part of the
data set used in the study of cross-country growth rates by De Long and Summers (1991).
The variable gpcrgdp represents average annual growth in real per capita GDP over the period 1960 to
1985. The fact that govcons60 (government consumption as a percentage of GDP) and second60 (percentage
of adult population with a secondary education) correspond to the year 1960, while gpcrgdp is the average
growth over the period from 1960 to 1985, does not lead to any special problems in treating this information
as a cross-sectional data set. The observations are listed alphabetically by country, but nothing about this
ordering affects any subsequent analysis.
Table 1.2 A Data Set on Economic Growth Rates and Country Characteristics
obsno country gpcrgdp govcons60 second60
1 Argentina 0.89 9 32
2 Austria 3.32 16 50
3 Belgium 2.56 13 69
4 Bolivia 1.24 18 12
. . . . .
. . . . .
. . . . .
61 Zimbabwe 2.30 17 6
Copyright 2025 Cengage Learning. All Rights Reserved. May not be copied, scanned, or duplicated, in whole or in part. Due to electronic rights, some third party content may be suppressed from the eBook and/or eChapter(s).
Editorial review has deemed that any suppressed content does not materially affect the overall learning experience. Cengage Learning reserves the right to remove additional content at any time if subsequent rights restrictions require it.
Another feature of time series data that can require special attention is the data frequency at which the
data are collected. In economics, the most common frequencies are daily, weekly, monthly, quarterly, and
annually. Stock prices are recorded at daily intervals (excluding Saturday and Sunday). The money supply in
the U.S. economy is reported weekly. Many macroeconomic series are tabulated monthly, including inflation
and unemployment rates. Other macro series are recorded less frequently, such as every three months (every
quarter). GDP is an important example of a quarterly series. Other time series, such as infant mortality rates
for states in the United States, are available only on an annual basis.
Many weekly, monthly, and quarterly economic time series display a strong seasonal pattern, which can
be an important factor in a time series analysis. For example, monthly data on housing starts differ across
the months simply due to changing weather conditions. We will learn how to deal with seasonal time series
in Chapter 10.
Table 1.3 contains a time series data set obtained from an article by Castillo-Freeman and Freeman (1992)
on minimum wage effects in Puerto Rico. The earliest year in the data set is the first observation, and the most
recent year available is the last observation. When econometric methods are used to analyze time series data,
the data should be stored in chronological order.
The variable avgmin refers to the average minimum wage for the year, avgcov is the average coverage rate
(the percentage of workers covered by the minimum wage law), prunemp is the unemployment rate, and prgnp
is the gross national product, in millions of 1954 dollars. We will use these data later in a time series analysis
of the effect of the minimum wage on employment. The subject of minimum wage effects on employment
remains a very important topic, and, with some effort, such older time series data sets could be updated to
produce a more current analysis.
Table 1.3 Minimum Wage, Unemployment, and Related Data for Puerto Rico
obsno year avgmin avgcov prunemp prgnp
1 1950 0.20 20.1 15.4 878.7
2 1951 0.21 20.7 16.0 925.0
3 1952 0.23 22.6 14.8 1015.9
. . . . . .
. . . . . .
. . . . . .
37 1986 3.35 58.1 18.9 4281.6
38 1987 3.35 58.2 16.8 4496.7
Copyright 2025 Cengage Learning. All Rights Reserved. May not be copied, scanned, or duplicated, in whole or in part. Due to electronic rights, some third party content may be suppressed from the eBook and/or eChapter(s).
Editorial review has deemed that any suppressed content does not materially affect the overall learning experience. Cengage Learning reserves the right to remove additional content at any time if subsequent rights restrictions require it.
A pooled cross section is analyzed much like a standard cross section, except that we often need to account
for secular differences in the variables across the time. In fact, in addition to increasing the sample size, the
point of a pooled cross-sectional analysis is often to see how a key relationship has changed over time.
Pooled (repeated) cross sections play an important role for policy analysis, where we can collect
information on two groups, one subjected to an intervention and the other not. We can obtain samples on
the two groups both before and after the intervention occurs, but we cannot follow the same units over
time. For example, we might obtain employment status, and other information, on individuals across two
states where one state increases its minimum wage. Often such survey sampling obtains random samples
in the separate time periods, in which case very few, if any, individuals would appear in the data set in
both years.
Copyright 2025 Cengage Learning. All Rights Reserved. May not be copied, scanned, or duplicated, in whole or in part. Due to electronic rights, some third party content may be suppressed from the eBook and/or eChapter(s).
Editorial review has deemed that any suppressed content does not materially affect the overall learning experience. Cengage Learning reserves the right to remove additional content at any time if subsequent rights restrictions require it.
same cities happen to show up in each year. But, as we will see in Chapters 13 and 14, we can also use the
panel structure to analyze questions that cannot be answered by simply viewing this as a pooled cross section.
In organizing the observations in Table 1.5, we place the two years of data for each city adjacent to one
another, with the first year coming before the second in all cases. For just about every practical purpose, this
is the preferred way for ordering panel data sets. Contrast this organization with the way the pooled cross
sections are stored in Table 1.4. In short, the reason for ordering panel data as in Table 1.5 is that we will need
to perform data transformations for each city across the two years.
Because panel data require replication of the same units over time, panel data sets, especially those on
individuals, households, and firms, are more difficult to obtain than pooled cross sections. Not surprisingly,
observing the same units over time leads to several advantages over cross-sectional data or even pooled
cross-sectional data. The benefit that we will focus on in this text is that having multiple observations on the
same units allows us to control for certain unobserved characteristics of individuals, firms, and so on. As we
will see, the use of more than one observation can facilitate causal inference in situations where inferring
causality would be very difficult if only a single cross section were available. A second advantage of panel
data is that they often allow us to study the importance of lags in behavior or the result of decision making.
This information can be significant because many economic policies can be expected to have an impact only
after some time has passed.
Most books at the undergraduate level do not contain a discussion of econometric methods for panel data.
However, economists now recognize that some questions are difficult, if not impossible, to answer satisfacto-
rily without panel data. As you will see, we can make considerable progress with simple panel data analysis,
a method that is not much more difficult than dealing with a standard cross-sectional data set.
Copyright 2025 Cengage Learning. All Rights Reserved. May not be copied, scanned, or duplicated, in whole or in part. Due to electronic rights, some third party content may be suppressed from the eBook and/or eChapter(s).
Editorial review has deemed that any suppressed content does not materially affect the overall learning experience. Cengage Learning reserves the right to remove additional content at any time if subsequent rights restrictions require it.
In Part 3, we will treat pooled cross sections and panel data explicitly. The analysis of independently
pooled cross sections and simple panel data analysis are fairly straightforward extensions of pure cross-
sectional analysis. Nevertheless, we will wait until Chapter 13 to deal with these topics.
Copyright 2025 Cengage Learning. All Rights Reserved. May not be copied, scanned, or duplicated, in whole or in part. Due to electronic rights, some third party content may be suppressed from the eBook and/or eChapter(s).
Editorial review has deemed that any suppressed content does not materially affect the overall learning experience. Cengage Learning reserves the right to remove additional content at any time if subsequent rights restrictions require it.
We rely, for now, on your intuitive understanding of such terms as random, independence, and correla-
tion, all of which should be familiar from an introductory probability and statistics course. (These concepts
are reviewed in Math Refresher B.) We begin with an example that illustrates some of these important issues.
The next example is more representative of the difficulties that arise when inferring causality in applied
economics.
Copyright 2025 Cengage Learning. All Rights Reserved. May not be copied, scanned, or duplicated, in whole or in part. Due to electronic rights, some third party content may be suppressed from the eBook and/or eChapter(s).
Editorial review has deemed that any suppressed content does not materially affect the overall learning experience. Cengage Learning reserves the right to remove additional content at any time if subsequent rights restrictions require it.
Unlike the fertilizer-yield example, the experiment described in Example 1.4 is unfeasible. The ethical
issues, not to mention the economic costs, associated with randomly determining education levels for a group
of individuals are obvious. As a logistical matter, we could not give someone only an eighth-grade education
if he or she already has a college degree.
Even though experimental data cannot be obtained for measuring the return to education, we can certainly
collect nonexperimental data on education levels and wages for a large group by sampling randomly from the
population of working people. Such data are available from a variety of surveys used in labor economics, but
these data sets have a feature that makes it difficult to estimate the ceteris paribus return to education. People
choose their own levels of education; therefore, education levels are probably not determined independently
of all other factors affecting wage. This problem is a feature shared by most nonexperimental data sets.
One factor that affects wage is experience in the workforce. Because pursuing more education generally
requires postponing entering the workforce, those with more education usually have less experience. Thus,
in a nonexperimental data set on wages and education, education is likely to be negatively associated with
a key variable that also affects wage. It is also believed that people with more innate ability often choose
higher levels of education. Because higher ability leads to higher wages, we again have a correlation between
education and a critical factor that affects wage.
The omitted factors of experience and ability in the wage example have analogs in the fertilizer example.
Experience is generally easy to measure and therefore is similar to a variable such as rainfall. Ability, on
the other hand, is nebulous and difficult to quantify; it is similar to land quality in the fertilizer example. As
we will see throughout this text, accounting for other observed factors, such as experience, when estimating
the!ceteris paribus effect of another variable, such as education, is relatively straightforward. We will also find
that accounting for inherently unobservable factors, such as ability, is much more problematic. It is fair to say
that many of the advances in econometric methods have tried to deal with unobserved factors in econometric
models.
One final parallel can be drawn between Examples 1.3 and 1.4. Suppose that in the fertilizer example, the
fertilizer amounts were not entirely determined at random. Instead, the assistant who chose the fertilizer levels
thought it would be better to put more fertilizer on the higher-quality plots of land. (Agricultural researchers
should have a rough idea about which plots of land are of better quality, even though they may not be able to
fully quantify the differences.) This situation is completely analogous to the level of schooling being related
to unobserved ability in Example 1.4. Because better land leads to higher yields, and more fertilizer was used
on the better plots, any observed relationship between yield and fertilizer might be spurious.
Difficulty in inferring causality can also arise when studying data at fairly high levels of aggregation, as
the next example on city crime rates shows.
Copyright 2025 Cengage Learning. All Rights Reserved. May not be copied, scanned, or duplicated, in whole or in part. Due to electronic rights, some third party content may be suppressed from the eBook and/or eChapter(s).
Editorial review has deemed that any suppressed content does not materially affect the overall learning experience. Cengage Learning reserves the right to remove additional content at any time if subsequent rights restrictions require it.
Although policies can be used to affect the size of police forces, we clearly cannot tell each city how
many police officers it can hire. If, as is likely, a city’s decision on how many police officers to hire is
correlated with other city factors that affect crime, then the data must be viewed as nonexperimental. In fact,
one way to view this problem is to see that a city’s choice of police force size and the amount of crime are
simultaneously determined. We will explicitly address such problems in Chapter 16.
The first three examples we have discussed have dealt with cross-sectional data at various levels of
aggregation (e.g., at the individual or city levels). The same hurdles arise when inferring causality in time
series problems.
Even when economic theories are not most naturally described in terms of causality, they often have
predictions that can be tested using econometric methods. The following example demonstrates this approach.
Copyright 2025 Cengage Learning. All Rights Reserved. May not be copied, scanned, or duplicated, in whole or in part. Due to electronic rights, some third party content may be suppressed from the eBook and/or eChapter(s).
Editorial review has deemed that any suppressed content does not materially affect the overall learning experience. Cengage Learning reserves the right to remove additional content at any time if subsequent rights restrictions require it.
can sell it for in three months. Therefore, there is uncertainty in this investment for someone who has a
three-month investment horizon.
The actual returns on these two investments will usually be different. According to the expectations
hypothesis, the expected return from the second investment, given all information at the time of investment,
should equal the return from purchasing a three-month T-bill. This theory turns out to be fairly easy to test,
as we will see in Chapter 11.
Summary
In this introductory chapter, we have discussed the purpose and scope of econometric analysis. Econometrics is
used in all applied economics fields to test economic theories, to inform government and private policy makers,
and to predict economic time series. Sometimes, an econometric model is derived from a formal economic
model, but in other cases, econometric models are based on informal economic reasoning and intuition. The
goals of any econometric analysis are to estimate the parameters in the model and to test hypotheses about
these parameters; the values and signs of the parameters determine the validity of an economic theory and
the effects of certain policies.
Cross-sectional, time series, pooled cross-sectional, and panel data are the most common types of data
structures that are used in applied econometrics. Data sets involving a time dimension, such as time series
and panel data, require special treatment because of the correlation across time of most economic time series.
Other issues, such as trends and seasonality, arise in the analysis of time series data but not cross-sectional data.
In Section 1-4, we discussed the notions of causality, ceteris paribus, and counterfactuals. In most cases,
hypotheses in the social sciences are ceteris paribus in nature: all other relevant factors must be fixed when
studying the relationship between two variables. As we discussed, one way to think of the ceteris paribus
requirement is to undertake a thought experiment where the same economic unit operates in different states
of the world, such as different policy regimes. Because of the nonexperimental nature of most data collected
in the social sciences, uncovering causal relationships is very challenging.
Key Terms
Causal Effect Economic Model Pooled Cross Section
Ceteris Paribus Empirical Analysis Potential Outcomes
Counterfactual Outcomes Experimental Data Random Sampling
Counterfactual Reasoning Longitudinal Data Repeated Cross Section
Cross-Sectional Data Set Nonexperimental Data Retrospective Data
Data Frequency Observational Data Time Series Data
Econometric Model Panel Data
Problems
1 Suppose that you are asked to conduct a study to determine whether smaller class sizes lead to
improved student performance of fourth graders.
(i) If you could conduct any experiment you want, what would you do? Be specific.
(ii) More realistically, suppose you can collect observational data on several thousand fourth
graders in a given state. You can obtain the size of their fourth-grade class and a standardized
test score taken at the end of fourth grade. Why might you expect a negative correlation
between class size and test score?
(iii) Would a negative correlation necessarily show that smaller class sizes cause better
performance? Explain.
Copyright 2025 Cengage Learning. All Rights Reserved. May not be copied, scanned, or duplicated, in whole or in part. Due to electronic rights, some third party content may be suppressed from the eBook and/or eChapter(s).
Editorial review has deemed that any suppressed content does not materially affect the overall learning experience. Cengage Learning reserves the right to remove additional content at any time if subsequent rights restrictions require it.
2 A justification for job training programs is that they improve worker productivity. Suppose that you
are asked to evaluate whether more job training makes workers more productive. However, rather
than having data on individual workers, you have access to data on manufacturing firms in Ohio. In
particular, for each firm, you have information on hours of job training per worker (training) and
number of nondefective items produced per worker hour (output).
(i) Carefully state the ceteris paribus thought experiment underlying this policy question.
(ii) Does it seem likely that a firm’s decision to train its workers will be independent of
worker characteristics? What are some of those measurable and unmeasurable worker
characteristics?
(iii) Name a factor other than worker characteristics that can affect worker productivity.
(iv) If you find a positive correlation between output and training, would you have convincingly
established that job training makes workers more productive? Explain.
3 Suppose at your university you are asked to find the relationship between weekly hours spent study-
ing (study) and weekly hours spent working (work). Does it make sense to characterize the problem
as inferring whether study “causes” work or work “causes” study? Explain.
4 States (and provinces) that have control over taxation sometimes reduce taxes in an attempt to spur
economic growth. Suppose that you are hired by a state to estimate the effect of corporate tax rates
on, say, the growth in per capita gross state product (GSP).
(i) What kind of data would you need to collect to undertake a statistical analysis?
(ii) Is it feasible to do a controlled experiment? What would be required?
(iii) Is a correlation analysis between GSP growth and tax rates likely to be convincing? Explain.
Computer Exercises
C1 Use the data in WAGE1 for this exercise.
(i) Find the average education level in the sample. What are the lowest and highest years of
education?
(ii) Find the average hourly wage in the sample. Does it seem high or low?
(iii) The wage data are reported in 1976 dollars. Using the Internet or a printed source, find the
Consumer Price Index (CPI) for the years 1976 and 2022.
(iv) Use the CPI values from part (iii) to find the average hourly wage in 2022 dollars. Now does
the average hourly wage seem reasonable?
(v) How many women are in the sample? How many men?
C2 Use the data in BWGHT to answer this question.
(i) How many women are in the sample, and how many report smoking during pregnancy?
(ii) What is the average number of cigarettes smoked per day? Is the average a good measure of
the “typical” woman in this case? Explain.
(iii) Among women who smoked during pregnancy, what is the average number of cigarettes
smoked per day? How does this compare with your answer from part (ii), and why?
(iv) Find the average of fatheduc in the sample. Why are only 1,192 observations used to
compute this average?
(v) Report the average family income and its standard deviation in dollars.
C3 The data in MEAP01 are for the state of Michigan in the year 2001. Use these data to answer the
following questions.
(i) Find the largest and smallest values of math4. Does the range make sense? Explain.
(ii) How many schools have a perfect pass rate on the math test? What percentage is this of the
total sample?
(iii) How many schools have math pass rates of exactly 50%?
(iv) Compare the average pass rates for the math and reading scores. Which test is harder to pass?
Copyright 2025 Cengage Learning. All Rights Reserved. May not be copied, scanned, or duplicated, in whole or in part. Due to electronic rights, some third party content may be suppressed from the eBook and/or eChapter(s).
Editorial review has deemed that any suppressed content does not materially affect the overall learning experience. Cengage Learning reserves the right to remove additional content at any time if subsequent rights restrictions require it.
(v) Find the correlation between math4 and read4. What do you conclude?
(vi) The variable exppp is expenditure per pupil. Find the average of exppp along with its
standard deviation. Would you say there is wide variation in per pupil spending?
(vii) Suppose School A spends $6,000 per student and School B spends $5,500 per student.
By what percentage does School A’s spending exceed School B’s? Compare this to 100 ·
[log(6,000) " log(5,500)], which is the approximation percentage difference based on the
difference in the natural logs. (See Section A.4 in Math Refresher A.)
C4 The data in JTRAIN2 come from a job training experiment conducted for low-income men during
1976–1977; see Lalonde (1986).
(i) Use the indicator variable train to determine the fraction of men receiving job training.
(ii) The variable re78 is earnings from 1978, measured in thousands of 1982 dollars. Find the
averages of re78 for the sample of men receiving job training and the sample not receiving
job training. Is the difference economically large?
(iii) The variable unem78 is an indicator of whether a man was unemployed or not in 1978. What
fraction of the men who received job training were unemployed? What about for men who
did not receive job training? Comment on the difference.
(iv) From parts (ii) and (iii), does it appear that the job training program was effective? What
would make our conclusions more convincing?
C5 The data in FERTIL2 were collected on women living in the Republic of Botswana in 1988. The
variable children refers to the number of living children. The variable electric is a binary indicator
equal to one if the woman’s home has electricity, and zero if not.
(i) Find the smallest and largest values of children in the sample. What is the average of
children?
(ii) What percentage of women have electricity in the home?
(iii) Compute the average of children for those without electricity and do the same for those with
electricity. Comment on what you find.
(iv) From part (iii), can you infer that having electricity “causes” women to have fewer children?
Explain.
C6 Use the data in COUNTYMURDERS to answer this question. Use only the year 1996. The variable
murders is the number of murders reported in the county. The variable execs is the number of exe-
cutions that took place of people sentenced to death in the given county. Most states in the United
States have the death penalty, but several do not.
(i) How many counties are there in the data set? Of these, how many have zero murders? What
percentage of counties have zero executions? (Remember, use only the 1996 data.)
(ii) What is the largest number of murders? What is the largest number of executions? Compute
the average number of executions and explain why it is so small.
(iii) Compute the correlation coefficient between murders and execs and describe what you find.
(iv) You should have computed a positive correlation in part (iii). Do you think that more
executions cause more murders to occur? What might explain the positive correlation?
C7 The data set in ALCOHOL contains information on a sample of men in the United States. Two key
variables are self-reported employment status and alcohol abuse (along with many other variables).
The variables employ and abuse are both binary, or indicator, variables: they take on only the values
zero and one.
(i) What percentage of the men in the sample report abusing alcohol? What is the employment
rate?
(ii) Consider the group of men who abuse alcohol. What is the employment rate?
(iii) What is the employment rate for the group of men who do not abuse alcohol?
(iv) Discuss the difference in your answers to parts (ii) and (iii). Does this allow you to conclude
that alcohol abuse causes unemployment?
Copyright 2025 Cengage Learning. All Rights Reserved. May not be copied, scanned, or duplicated, in whole or in part. Due to electronic rights, some third party content may be suppressed from the eBook and/or eChapter(s).
Editorial review has deemed that any suppressed content does not materially affect the overall learning experience. Cengage Learning reserves the right to remove additional content at any time if subsequent rights restrictions require it.
C8 The data in ECONMATH were obtained on students from a large university course in introductory
microeconomics. For this problem, we are interested in two variables: score, which is the final course
score, and econhs, which is a binary variable indicating whether a student took an economics course
in high school.
(i) How many students are in the sample? How many students report taking an economics
course in high school?
(ii) Find the average of score for those students who did take a high school economics class.
How does it compare with the average of score for those who did not?
(iii) Do the findings in part (ii) necessarily tell you anything about the causal effect of taking high
school economics on college course performance? Explain.
(iv) If you want to obtain a good causal estimate of the effect of taking a high school economics
course using the difference in averages, what experiment would you run?
C9 The data set in LOANAPP contains information on a sample of mortgage applications in the Boston,
Massachusetts area; the data were originally compiled by the Federal Reserve Bank of Boston. The
race variable, denoted white, is recorded as binary, equaling one if the mortgage applicant is classified
as White and zero otherwise. The variable approve is also binary, equaling one if the loan application
was approved, and zero otherwise.
(i) What proportion of the applicants in the sample are recorded as being White?
(ii) For the group of non-White applicants, what is the proportion of loans approved?
(iii) What is the loan approval rate for White applicants?
(iv) Report the difference between (iii) and (ii) in percentage point terms. Does this analysis
allow you to convincingly conclude that there is racial bias in mortgage application
approvals? Explain.
Copyright 2025 Cengage Learning. All Rights Reserved. May not be copied, scanned, or duplicated, in whole or in part. Due to electronic rights, some third party content may be suppressed from the eBook and/or eChapter(s).
Editorial review has deemed that any suppressed content does not materially affect the overall learning experience. Cengage Learning reserves the right to remove additional content at any time if subsequent rights restrictions require it.