CHAPTER THREE
Chapter Two
Introduction to Basic Regression Analysis with Time Series Data
2.1. Nature of Time Series Data
Recall that one of the important types of data used in empirical analysis is time series data. A
time series data set consists of observations on a variable or several variables over time. These
are data that can be collected over time, for instances, weekly, monthly, quarterly, semiannually,
annually, etc. Examples of time series data include stock prices, money supply, consumer price
index, gross domestic product, and automobile sales figures. Because past events can influence
future events and lags in behavior are prevalent in the social sciences, time is an important
dimension in time series data set. Unlike the arrangement of cross-sectional data, the
chronological ordering of observations in a time series conveys potentially important
information. In this and the following sections we take a closer look at such data not only
because of the frequency with which they are used in practice but also because they pose several
challenges to econometricians and practitioners.
2.2. Stochastic Processes
From a theoretical point of view, a time series is a collection of random variables (Xt). Such a
collection of random variables ordered in time is called a stochastic process. Loosely speaking, a
random or stochastic process is a collection of random variables ordered in time. The word
stochastic has a Greek origin and means "pertaining to chance." If we let Y denote a random
variable, and if it is continuous, we denote it as Y(t), but if it is discrete, we denoted it as Yt. We
distinguish two types of stochastic process: stationary and non-stationary stochastic processes.
2.2.1. Stationary Stochastic Processes
A type of stochastic process that has received a great deal of attention by time series analysts is
the so-called stationary stochastic process. A stochastic process is said to be stationary if its
mean and variance are constant over time and the value of the covariance between the two time
periods depends only on the distance or gap or lag between the two time periods and not the
actual time at which the covariance is computed. In the time series literature, such a stochastic
1|Page
process is known as a weakly stationary, or covariance stationary, or second-order stationary, or
wide sense, stochastic process. To explain stationarity, let Yt be a stochastic time series with
these properties:
Mean: 𝐸(𝑌𝑡 ) = 𝜇 (2.1)
Variance: 𝑉𝑎𝑟(𝑌𝑡 ) = 𝐸(𝑌𝑡 − 𝜇)2= 𝜎 2 (2.2)
Covariance: 𝛶𝑘 =𝐸(𝑌𝑡 − 𝜇) (𝑌𝑡+𝑘 − 𝜇) (2.3)
where 𝛶𝑘 , the covariance (or autocovariance) at lag k, is the covariance between the values of 𝑌𝑡
and 𝑌𝑡+𝑘 , that is, between two Y values k periods apart. If k = 0, we obtain 𝛶0 , which is simply the
variance of Y (=𝜎 2 ); if k = 1, 𝛶1 is the covariance between two adjacent values of Y (recall the
first-order autoregressive scheme).
Suppose we shift the origin of Y from 𝑌𝑡 to 𝑌𝑡+𝑚 (for instance, from the first quarter of 1970 to
the first quarter of 1975 say for GDP data). Now if 𝑌𝑡 is to be stationary, the mean, variance, and
autocovariances of 𝑌𝑡+𝑚 must be the same as those of 𝑌𝑡 . In short, if a time series is stationary,
its mean, variance, and autocovariance (at various lags) remain the same no matter at what
point we measure them; that is, they are time invariant. Such a time series will tend to return to
its mean (called mean reversion) and fluctuations around this mean (measured by its variance)
will remain constant.
Why are stationary time series so important? Because if a time series is non-stationary, we can
study its behavior only for the time period under consideration. Each set of time series data will
therefore be for a particular time period. As a consequence, it is not possible to generalize it to
other time periods. Therefore, for the purpose of forecasting, such (non-stationary) time series
may be of little practical value. We call a stochastic process purely random (white noise,
process) if it has zero mean, constant variance 𝜎 2 , and is serially uncorrelated. Loosely speaking,
the error term 𝑢𝑡 is assumed to be a white noise process if 𝑢𝑡 ̴ IIDN (0,𝜎 2 ); that is, 𝑢𝑡 is
independently and identically distributed as a normal distribution with zero mean and constant
variance.
2|Page
2.2.2. Nonstationary Stochastic Processes
Although our interest is in stationary time series, one often encounters non-stationary time series,
the classic example being the random walk model (RWM). That is, a non-stationary time series
will have a time varying mean or a time-varying variance or both. It is often said that asset
prices, such as stock prices or exchange rates, follow a random walk; that is, they are non-
stationary. We distinguish two types of random walks: (1) random walk without drift (i.e., no
constant or intercept term) and (2) random walk with drift (i.e., a constant term is present).
A) Random Walk without Drift. Suppose 𝑢𝑡 is a white noise error term with mean 0 and
variance σ2. Then the series Yt is said to be a random walk if
𝑌𝑡 = 𝑌𝑡−1 + 𝑢𝑡 (2.4)
In the random walk model, as (2.4) shows, the value of Y at time t is equal to its value at time (t −
1) plus a random shock; thus it is an AR(1) model. We can think of (2.4) as a regression of Y at
time t on its value lagged one period. Believers in the efficient capital market hypothesis argue
that stock prices are essentially random and therefore there is no scope for profitable speculation
in the stock market: If one could predict tomorrow’s price on the basis of today’s price, we
would all be millionaires.
Now from (2.4) we can write
𝑌1 = 𝑌0 + 𝑢1
𝑌2 = 𝑌1 + 𝑢2 = 𝑌0 + 𝑢1 + 𝑢2
𝑌3 = 𝑌2 + 𝑢3 = 𝑌0 + 𝑢1 + 𝑢2 + 𝑢3
In general, if the process started at some time 0 with a value of 𝑌0 , we have
𝑌𝑡 = 𝑌0 +∑ 𝑢𝑡 (2.5)
Therefore,
𝐸(𝑌𝑡 ) =𝐸(𝑌0 +∑ 𝑢𝑡 ) = 𝑌0 (why?) (2.6)
In like fashion, it can be shown that
𝑉𝑎𝑟(𝑌𝑡 ) = 𝑡𝜎 2 (2.7)
3|Page
As the preceding expression shows, the mean of Y is equal to its initial, or starting, value, which
is constant, but as t increases, its variance increases indefinitely, thus violating a condition of
stationarity. In short, the RWM without drift is a non-stationary stochastic process. In practice 𝑌0
is often set at zero, in which case 𝐸(𝑌𝑡 ) =𝑌0 .
An interesting feature of RWM is the persistence of random shocks (i.e., random errors), which
is clear from (2.5): 𝑌𝑡 is the sum of initial 𝑌0 plus the sum of random shocks. As a result, the
impact of a particular shock does not die away. For example, if 𝑢2 = 2 rather than 𝑢2 = 0, then all
𝑌𝑡 ’s from 𝑌2 onward will be 2 units higher and the effect of this shock never dies out so that
random walk is said to have an infinite memory. As Kerry Patterson notes, random walk
remembers the shock forever; that is, it has infinite memory.
Interestingly, if you write (2.4) as
𝑌𝑡 - 𝑌𝑡−1 =∆𝑌𝑡 = 𝑢𝑡 (2.8)
where ∆ is the first difference operator. It is easy to show that, while 𝑌𝑡 is non-stationary, its first
difference is stationary. In other words, the first differences of a random walk time series are
stationary.
B) Random Walk with Drift. Let us modify (2.4) as follows:
𝑌𝑡 = 𝛿 + 𝑌𝑡−1 + 𝑢𝑡 (2.9)
where δ is known as the drift parameter. The name drift comes from the fact that if we write
the preceding equation as
𝑌𝑡 - 𝑌𝑡−1 =∆𝑌𝑡 = 𝛿 + 𝑢𝑡 (2.10)
it shows that 𝑌𝑡 drifts (moves) upward or downward, depending on δ being positive or negative.
Note that model (2.9) is also an AR(1) model. Following the procedure discussed for random
walk without drift, it can be shown that for the random walk with drift model (2.9),
𝐸(𝑌𝑡 ) = 𝑌0 + t. 𝛿 (2.11)
𝑉𝑎𝑟(𝑌𝑡 ) = 𝑡𝜎 2 (2.12)
4|Page
Hence, for RWM with drift the mean as well as the variance increases over time, again violating
the conditions of stationarity. In short, RWM, with or without drift, is a non-stationary stochastic
process.
2.3. Unit Root Stochastic Process
Let us write the RWM (2.4) as:
𝑌𝑡 = 𝜌𝑌𝑡−1 + 𝑢𝑡 −1 ≤ 𝜌 ≤ 1 (2.13)
This model resembles the Markov first-order autoregressive model that we discussed in the case
of autocorrelation. If ρ = 1, (2.13) becomes a RWM (without drift). If ρ is in fact 1, we face what
is known as the unit root problem, that is, a situation of non-stationarity; we already know that
in this case the variance of 𝑌𝑡 is not stationary. The name unit root is due to the fact that ρ = 11.
Thus, the terms non-stationarity, random walk, and unit root can be treated as synonymous.
If, however, |ρ| ≤ 1, that is if the absolute value of ρ is less than one, then it can be shown that the
time series 𝑌𝑡 is stationary in the sense we have defined it2. In practice, then, it is important to
find out if a time series possesses a unit root. In the next section we will discuss tests of unit root,
that is, tests of stationarity.
2.4. Trend Stationary (TS) and Difference Stationary (DS) Stochastic Processes
The distinction between stationary and non-stationary stochastic processes (or time series) has a
crucial bearing on whether the trend is deterministic or stochastic. Broadly speaking, if the
trend in a time series is completely predictable and not variable, we call it a deterministic trend,
whereas if it is not predictable, we call it a stochastic trend. To make the definition more formal,
consider the following model of the time series 𝑌𝑡 .
𝑌𝑡 = β1 + β2t + β3𝑌𝑡−1 + 𝑢𝑡 (2.14)
1
A technical point: If ρ = 1, we can write (2.13) as Yt -Yt−1 = ut . Now using the lag operator L so that LYt = Yt−1 ,
L2Yt = Yt−2 , and so on, we can write (2.13) as (1 − L) Yt = ut . The term unit root refers to the root of the polynomial
in the lag operator. If you set (1 − L) = 0, we obtain, L = 1, hence the name unit root.
2
If in (2.13) it is assumed that the initial value of Y (=𝑌0 ) is zero, |ρ| ≤ 1, and 𝑢𝑡 is white noise and distributed
normally with zero mean and unit variance, then it follows that 𝐸(𝑌𝑡 ) = 0 and 𝑉𝑎𝑟(𝑌𝑡 ) = 1/(1 − ρ2). Since both
these are constants, by the definition of stationarity, 𝑌𝑡 is stationary. On the other hand, as we saw before, if 𝜌 =1,
𝑌𝑡 is a random walk or nonstationary.
5|Page
Where, 𝑢𝑡 is a white noise error term and where t is time measured chronologically. Now we
have the following possibilities:
❖ Pure random walk: If in (2.14) β1 = 0, β2 = 0, β3 = 1, we get;
𝑌𝑡 = 𝑌𝑡−1 + 𝑢𝑡 (2.15)
Which is nothing but a RWM without drift and is therefore non-stationary. But note that, if we
write (2.15) as
∆𝑌𝑡 = 𝑌𝑡 − 𝑌𝑡−1 = 𝑢𝑡 (2.8)
It becomes stationary, as noted before. Hence, a RWM without drift is a difference stationary
process (DSP).
❖ Random walk with drift: If in (2.14) β1 ≠ 0, β2 = 0, β3 = 1, we get;
𝑌𝑡 = β1+𝑌𝑡−1 + 𝑢𝑡 (2.16a)
Which is a random walk with drift and is therefore non-stationary. If we write it as
𝑌𝑡 − 𝑌𝑡−1 = ∆𝑌𝑡 = β1 + 𝑢𝑡 (2.16b)
This means 𝑌𝑡 will exhibit a positive (β1 > 0) or negative (β1 < 0) trend. Such a trend is called a
stochastic trend. Equation (2.16b) is a DSP process because the non-stationarity in 𝑌𝑡 can be
eliminated by taking first differences of the time series.
❖ Deterministic trend: If in (2.14), β1 ≠ 0, β2 ≠ 0, β3 = 0, we obtain;
𝑌𝑡 = β1 + β1t + 𝑢𝑡 (2.17)
which is called a trend stationary process (TSP). Although the mean of 𝑌𝑡 is β1 + β1t, which is
not constant, its variance (= σ2) is. Once the values of β1 and β2 are known, the mean can be
forecast perfectly. Therefore, if we subtract the mean of 𝑌𝑡 from 𝑌𝑡 , the resulting series will be
stationary, hence the name trend stationary. This procedure of removing the (deterministic)
trend is called detrending.
6|Page
❖ Random walk with drift and deterministic trend: If in (21.5.1), β1 ≠ 0, β2 ≠ 0, β3 = 1,
we obtain;
𝑌𝑡 = β1 + β2t +𝑌𝑡−1 + 𝑢𝑡 (2.18a)
we have a random walk with drift and a deterministic trend, which can be seen if we write this
equation as
∆𝑌𝑡 = β1 + β1t + 𝑢𝑡 (2.18b)
❖ Deterministic trend with stationary AR(1) component: If in (21.5.1) β1 ≠ 0, β2 ≠ 0, β3 <
1, then we get;
𝑌𝑡 = β1 + β2t +β3𝑌𝑡−1 + 𝑢𝑡 (2.19)
which is stationary around the deterministic trend.
2.5. Integrated Stochastic Processes
The random walk model is a specific case of a more general class of stochastic processes known
as integrated processes. Recall that the RWM without drift is non-stationary, but its first
difference, as shown in (2.8), is stationary. Therefore, we call the RWM without drift integrated
of order 1, denoted as 𝐼(1). Similarly, if a time series has to be differenced twice (i.e., take the
first difference of the first differences) to make it stationary, we call such a time series
integrated of order 23. In general, if a (nonstationary) time series has to be differenced d times
to make it stationary, that time series is said to be integrated of order d. A time series 𝑌𝑡
integrated of order d is denoted as 𝑌𝑡 ̴ 𝐼(𝑑). If a time series 𝑌𝑡 is stationary to begin with (i.e., it
does not require any differencing), it is said to be integrated of order zero, denoted by 𝑌𝑡 ̴ 𝐼(0).
Thus, we will use the terms “stationary time series” and “time series integrated of order zero” to
mean the same thing. Most economic time series are generally 𝐼(1); that is, they generally
become stationary only after taking their first differences.
3
For example if 𝑌𝑡 is I(2), then ∆∆𝑌𝑡 = ∆(𝑌𝑡 −𝑌𝑡−1 ) = ∆𝑌𝑡 −∆𝑌𝑡−1 = 𝑌𝑡 −2𝑌𝑡−1 + 𝑌𝑡−2 will become stationary. But
note that ∆∆𝑌𝑡 = ∆𝟐 𝑌𝑡 ≠ 𝑌𝑡 −𝑌𝑡−2 .
7|Page
Properties of Integrated Series
The following properties of integrated time series may be noted: Let Xt , Yt ,and Zt be three time
series.
1) If 𝑋𝑡 ̴ 𝐼(0) and 𝑌𝑡 ̴ 𝐼(1), then 𝑍𝑡 = 𝑋𝑡 + 𝑌𝑡 = I(1); that is, a linear combination or sum of
stationary and non-stationary time series is non-stationary.
2) If 𝑋𝑡 ̴ 𝐼(𝑑), then 𝑍𝑡 = (𝑎 + 𝑏𝑋𝑡 ) = 𝐼(𝑑), where a and b are constants. That is, a linear
combination of a 𝐼(𝑑)series is also𝐼(𝑑). Thus, if 𝑋𝑡 ̴ 𝐼(0), then 𝑍𝑡 = (𝑎 + 𝑏𝑋𝑡 ) ̴ 𝐼(0).
3) If 𝑋𝑡 ̴ 𝐼(𝑑1 ) and 𝑌𝑡 ̴ 𝐼(𝑑2 ), then 𝑍𝑡 = (𝑎𝑋𝑡 + 𝑏𝑌𝑡 ) ̴ 𝐼(𝑑2 ), where 𝑑1 < 𝑑2 .
4) If 𝑋𝑡 ̴ 𝐼(𝑑) and 𝑌𝑡 ̴ 𝐼(𝑑), then 𝑍𝑡 = (𝑎𝑋𝑡 + 𝑏𝑌𝑡 ) ̴ I(d*); d* is generally equal to d, but in some
cases d* < d. As you can see from the preceding statements, one has to pay careful attention
in combining two or more-time series that are integrated of different order.
2.6. Spurious Regression
To see why stationary time series are so important, consider the following two random walk
models:
𝑌𝑡 = 𝑌𝑡−1 + 𝑢𝑡 (2.20)
𝑋𝑡 = 𝑋𝑡−1 + 𝑣𝑡 (2.21)
where we generated 500 observations of 𝑢𝑡 from 𝑢𝑡 ̴ N(0, 1) and 500 observations of 𝑣𝑡 from 𝑣𝑡 ̴
N(0, 1) and assumed that the initial values of both Y and X were zero. We also assumed that 𝑢𝑡
and 𝑣𝑡 are serially uncorrelated as well as mutually uncorrelated. As you know by now, both
these time series are nonstationary; that is, they are I(1) or exhibit stochastic trends.
Suppose we regress 𝑌𝑡 on 𝑋𝑡 . Since 𝑌𝑡 and 𝑋𝑡 are uncorrelated I(1) processes, the R2 from the
regression of Y on X should tend to zero; that is, there should not be any relationship between the
two variables. Consider the regression results:
____________________________________________________________________
Variable Coefficient Std. error t statistic
------------------------------------------------------------------------------------------------------
_cons -13.2556 0.6203 -21.36856
X 0.3376 0.0443 7.61223
R2 = 0.1044 d = 0.0121
---------------------------------------------------------------------------------------------------------
8|Page
As you can see, the coefficient of X is highly statistically significant, and, although the R2 value
is low, it is statistically significantly different from zero. From these results, you may be tempted
to conclude that there is a significant statistical relationship between Y and X, whereas a priori
there should be none. This is in a nutshell the phenomenon of spurious or nonsense regression,
first discovered by Yule. Yule showed that (spurious) correlation could persist in non-stationary
time series even if the sample is very large.
The regression result is characterized by the extremely low Durbin–Watson d value, which
suggests very strong first-order autocorrelation. According to Granger and Newbold, an R2 > d is
a good rule of thumb to suspect that the estimated regression is spurious. That the regression
results presented above are meaningless can be easily seen from regressing the first differences
of 𝑌𝑡 (=∆𝑌𝑡 ) on the first differences of 𝑋𝑡 (=∆𝑋𝑡 ); remember that although 𝑌𝑡 and 𝑋𝑡 are non-
stationary, their first differences are stationary. In such a regression you will find that R2 is
practically zero, as it should be, and the Durbin–Watson d is about 2.
2.7. Tests of Stationarity
By now you have a good idea about the nature of stationary stochastic processes and their
importance. In practice we face two important questions:
1) How do we find out if a given time series is stationary?
2) If we find that a given time series is not stationary, is there a way that it can be made
stationary?
2.7.1. The Unit Root Test
A test of stationarity (or non-stationarity) that has become widely popular over the past several
years is the unit root test. We will first explain it, and then illustrate it. To start with, we
consider equation (2.13).
𝑌𝑡 = 𝜌𝑌𝑡−1 + 𝑢𝑡 −1 ≤ 𝜌 ≤ 1 (2.13)
Where, 𝑢𝑡 is a white noise error term.
9|Page
We know that if ρ = 1, that is, in the case of the unit root, (2.13) becomes a random walk model
without drift, which we know is a non-stationary stochastic process. Therefore, to test whether
series Yt stationary or not, we regress 𝑌𝑡 on its (one period) lagged value 𝑌𝑡−1 and find out if the
estimated ρ is statistically equal to 1? If it is, then 𝑌𝑡 is non-stationary. This is the general idea
behind the unit root test of stationarity. For theoretical reasons, we manipulate (2.13) as follows:
Subtract 𝑌𝑡−1 from both sides of (2.13) to obtain:
𝑌𝑡 − 𝑌𝑡−1 = 𝜌𝑌𝑡−1 − 𝑌𝑡−1 + 𝑢𝑡 = (𝜌 − 1)𝑌𝑡−1 + 𝑢𝑡 (2.22)
which can be alternatively written as:
∆𝑌𝑡 =𝛿𝑌𝑡−1 + 𝑢𝑡 (2.23)
where δ = (ρ − 1) and ∆, as usual, is the first-difference operator.
In practice, therefore, instead of estimating (2.13), we estimate (2.23) and test the (null)
hypothesis that δ = 0. If δ = 0, then ρ = 1, that is we have a unit root, meaning the time series
under consideration is non-stationary. Before we proceed to estimate (2.23), it may be noted that
if δ = 0, (2.23) will become
∆𝑌𝑡 =𝑌𝑡 − 𝑌𝑡−1 =𝑢𝑡 (2.24)
Since 𝑢𝑡 is a white noise error term, it is stationary, which means that the first differences of a
random walk time series are stationary, a point we have already made before.
Now let us turn to the estimation of (2.23). We take the first differences of 𝑌𝑡 and regress them
on 𝑌𝑡−1 and see if the estimated slope coefficient in this regression (=𝛿̂ ) is zero or not. If it is
zero, we conclude that 𝑌𝑡 is non-stationary. But if it is negative, we conclude that 𝑌𝑡 is stationary.
The only question is which test we use to find out if the estimated coefficient of 𝑌𝑡−1 in (2.23) is
zero or not. You might be tempted to say, why not use the usual t test? Unfortunately, under the
null hypothesis that δ = 0 (i.e., ρ = 1), the t value of the estimated coefficient of 𝑌𝑡−1 does not
follow the t distribution even in large samples; that is, it does not have an asymptotic normal
distribution. This suggests that t-test is not applicable and instead we use widely used test called
Dickey-Fuller test.
10 | P a g e
Dickey and Fuller have shown that under the null hypothesis that δ = 0, the estimated t value of
the coefficient of 𝑌𝑡−1 in (2.23) follows the τ (tau) statistic. A sample of these critical values is
given D-F table. In the literature the tau statistic or test is known as the Dickey–Fuller (DF)
test, in honor of its discoverers.
The actual procedure of implementing the DF test involves several decisions. In discussing the
nature of the unit root process in sections 2.3 and 2.4, we noted that a random walk process may
have no drift, or it may have drift or it may have both deterministic and stochastic trends. To
allow for the various possibilities, the DF test is estimated in three different forms, that is, under
three different null hypotheses.
𝑌𝑡 is a random walk: ∆𝑌𝑡 =𝛿𝑌𝑡−1 + 𝑢𝑡 (2.23)
𝑌𝑡 is a random walk with drift: ∆𝑌𝑡 =β1 + 𝛿𝑌𝑡−1 + 𝑢𝑡 (2.25)
𝑌𝑡 is a random walk with drift
around a stochastic trend: ∆𝑌𝑡 =β1+ β2t + 𝛿𝑌𝑡−1 + 𝑢𝑡 (2.26)
where t is the time or trend variable. In each case, the null hypothesis is that δ = 0; that is, there is
a unit root-the time series is non-stationary. The alternative hypothesis is that δ is less than zero;
that is, the time series is stationary. If the null hypothesis is rejected, it means that 𝑌𝑡 is a
stationary time series with zero mean in the case of (2.23), that 𝑌𝑡 is stationary with a non-zero
mean [= β1/(1 − ρ)] in the case of (2.25), and that 𝑌𝑡 is stationary around a deterministic trend in
(2.26).
It is extremely important to note that the critical values of the tau test to test the hypothesis that δ
= 0, are different for each of the preceding three specifications of the DF test, which can be seen
tau table. Moreover, if, say, specification (2.25) is correct, but we estimate (2.23), we will be
committing a specification error, whose consequences we already know from Chapter 4 of
Econometrics I.
11 | P a g e
The actual estimation procedure is as follows: Estimate (2.23), or (2.24), or (2.25) by OLS;
divide the estimated coefficient of 𝑌𝑡−1 in each case by its standard error to compute the (τ) tau
statistic; and refer to the DF tables (or any statistical package). If the computed absolute value of
the tau-statistic (|τ |) exceeds the DF or MacKinnon critical tau values, we reject the hypothesis
that δ = 0, in which case the time series is stationary. On the other hand, if the computed |τ | does
not exceed the critical tau value, we do not reject the null hypothesis, in which case the time
series is non-stationary. Make sure that you use the appropriate critical τ values.
Example: The results of the regressions is as follows: The dependent variable in is ∆𝑌𝑡 = ∆𝐺𝐷𝑃𝑡
̂ 𝑡 = 28.2054 − 0.00136𝐺𝐷𝑃𝑡−1
∆𝐺𝐷𝑃 (2.27)
t = (1.1576) (−0.2191) R2 = 0.00056 d = 1.35
Our primary interest here is in the t (= τ) value of the 𝐺𝐷𝑃𝑡−1 coefficient. The critical 1, 5, and
10 percent τ values are −3.5064, −2.8947, and −2.5842. The estimated δ coefficient is negative,
implying that the estimated ρ is less than 1. The estimated τ value is −0.2191, which in absolute
value is below even the 10 percent critical value of −2.5842. Since, in absolute terms, the former
is smaller than the latter, our conclusion is that the GDP time series is not stationary.
2.7.2. Transforming Nonstationary Time Series
Now that we know the problems associated with non-stationary time series, the practical
question is what to do. To avoid the spurious regression problem that may arise from regressing
a non-stationary time series on one or more non-stationary time series, we have to transform non-
stationary time series to make them stationary. The transformation method depends on whether
the time series are difference stationary (DSP) or trend stationary (TSP). We consider each of
these methods in turn.
Difference-Stationary Processes
If a time series has a unit root, the first differences of such time series are stationary4. Therefore,
the solution here is to take the first differences of the time series. Returning to our U.S. GDP
4
If a time series is I(2), it will contain two unit roots, in which case we will have to difference it twice. If it is I(d), it has
to be differenced d times, where d is any integer.
12 | P a g e
time series, we have already seen that it has a unit root. Let us now see what happens if we take
the first differences of the GDP series.
Let ∆𝐺𝐷𝑃𝑡 = (𝐺𝐷𝑃𝑡 − 𝐺𝐷𝑃𝑡−1 ). For convenience, let 𝐷𝑡 =𝐺𝐷𝑃𝑡 . Now consider the following
regression:
̂𝑡 = 16.0049 − 0.06827𝐷𝑡−1
𝐷
t = (3.6402) (−6.6303) (2.28)
R2 = 0.3435 d = 2.0344
The 1 percent critical DF τ value is −3.5073. Since the computed τ (= t) is more negative than the
critical value, we conclude that the first-differenced GDP is stationary; that is, it is I(0).
Trend-Stationary Process
The simplest way to make Trend-Stationary Process time series stationary is to regress it on time
and the residuals from this regression will then be stationary. In other words, run the following
regression:
𝑌𝑡 = β + β2t + 𝑢𝑡 (2.29)
Where, 𝑌𝑡 is the time series under study and where t is the trend variable measured
chronologically. Now
𝑢 ̂1 -𝛽
̂𝑡 = (𝑌𝑡 -𝛽 ̂1 𝒕) (2.30)
will be stationary. 𝑢
̂𝑡 is known as a (linearly) detrended time series.
It should be pointed out that if a time series is DSP but we treat it as TSP, this is called
underdifferencing. On the other hand, if a time series is TSP but we treat it as DSP, this is
called overdifferencing. The consequences of these types of specification errors can be serious,
depending on how one handles the serial correlation properties of the resulting error terms.
In passing it may be noted that most macroeconomic time series are DSP rather than TSP.
13 | P a g e