INDIAN INSTITUTE OF TECHNOLOGY ROORKEE
Time Series Analysis
HYT-501 Data Analysis and Numerical Modelling
(Spring Semester, Session 2024-25)
Module 1: Data Analysis
Part 5 – Time Series Modeling
Dr. Pankaj Dey
DST INSPIRE Faculty
Department of Hydrology
Indian Institute of Technology, Roorkee
Time Series
A time series is an ordered sequence of observations of the same phenomenon, which order
is defined by the time when each data point is obtained.
The data points are typically measured at equally spaced successive instants of time.
The times are usually spaced at uniform time intervals, so the data might be denoted by
Univariate and MultivariateTime Series
RH: Relative Humidity, WS: Windspeed, Ta: Air Temperature
Streamflow time series
doi:10.1002/2016WR020216
10.5194/hess-25-2187-2021
A time series containing records of a single variable is termed as univariate, but if records of more than one
variable are considered then it is termed as multivariate.
Time Series Plots
The linear trend has a
constant positive slope
with random, year-to-
year variation.
The U.S. annual production of blue and gorgonzola cheeses.
Source: USDA-NASS.
Time Series Plots
The plot reveals overall
increasing trend, with a
distinct cyclic pattern that
is repeated within each year.
The U.S. beverage manufacturer monthly product shipments.
unadjusted. (Source: C.S. Census Bureau.)
Time Series Plots
The plot of the annual mean
anomaly in global surface
air temperature shows an
increasing trend since
1880.
Global mean surface air temperature annual anomaly. (Source:
NASA-GISS.)
Time Series Plots
• The cycle lengths were 13, 10,
12, 12, 11, and 9 years,
respectively.
• The cycle lengths seem to vary
randomly, and there does not
appear to be an “adjustment to
a fixed cycle length”.
SILSO data/image, Royal Observatory of Belgium, Brussels.
Constituents of Time Series
In general, a time series is affected by four components, i.e. trend, seasonal, cyclical
and irregular(random) components.
Trend: The general tendency of a time series to increase, decrease or stagnate over
a long period of time.
Seasonal variation: This component explains fluctuations within a year during the
season, usually caused by climate and weather conditions, customs, traditional
habits, etc.
Cyclical variation: This component describes the medium-term changes caused by
circumstances, which repeat in cycles.
Irregular variation: Irregular or random variations in a time series are caused by
unpredictable influences, which are not regular and also do not repeat in a particular
pattern.
Combinations of four constituents
Considering the effects of these four components, two different types of models are generally used
for a time series.
Additive Model
Assumption: These four components are independent of each other.
Multiplicative Model
Assumption: These four components of a time series are not necessarily independent and they can affect
one another.
White Noise Process
A simple time series could be a collection of uncorrelated random variables, 𝒘𝒕 , with zero mean
𝜇 = 0 and finite variance 𝜎𝑤2 , denoted as 𝒘𝒕 ~𝒘𝒏 𝟎; 𝝈𝟐𝒘 .
Gaussian White Noise
A particular useful white noise is Gaussian white noise, wherein the 𝒘𝒕 are independent normal
random variables (with mean 0 and variance 𝜎𝑤2 ), denoted as 𝒘𝒕 ~𝒊𝒊𝒅 𝓝 𝟎; 𝝈𝟐𝒘
Properties of White Noise Process:
1. Uncorrelated.
2. Unpredictable.
3. Independent and identically distributed (iid)
Stochastic Processes
A time series can be defined as a collection of random variables indexed according to
the order they are obtained in time, X1; X2; X3; …, Xt will typically be discrete and vary
over the integers t = 0; ±1; ±2; …
The collection of random variables 𝑋𝑡 is referred to as a stochastic process, while the
observed values are referred to as a realization of the stochastic process.
Joint distribution function
A complete description of a time series observed as a collection of n random variables at arbitrary
time points t1; t2; : : : ; tn, for any positive integer n, is provided by the joint distribution function,
evaluated as the probability that the values of the series are jointly less than the n constants, c1; c2;
: : : ; cn; i.e.,
Unfortunately, these multidimensional distribution functions cannot usually be written easily.
Therefore some informative descriptive measures can be useful, such as mean function and more.
Descriptive measures of JDF
Mean Function or Expectation
Descriptive measures of JDF
Variance
The variance function 𝜎𝑡2 is defined for all t by
+∞
𝜎𝑡2 = න 𝑋 𝑡 −𝜇 𝑡 2𝑓
𝑡 𝑥 𝑑𝑥
−∞
Descriptive measures of JDF
Autocovariance
Assuming the variance of Xt is finite, the autocovariance function is defined as the second-
moment product
Autocovariance
The autocovariance measures the linear dependence between two points on the
same series observed at different times.
Scatter diagram of pharmaceutical product sales at lag k = 1 Scatter diagram of chemical viscosity readings at lag k = 1
The plotted pairs of adjacent observations yt, yt+1 seem to be The pairs of adjacent observations yt, yt+1 are positively correlated.
uncorrelated. That is, the value of y in the current period does not That is, a small value of y tends to be followed in the next time period
provide any helpful information about the value of y that will be by another small value of y, and a large value of y tends to be
observed in the next period followed immediately by another large value of y.
Autocorrelation Function (ACF)
Autocorrelation indicates the
memory of a stochastic
process
Sample Autocorrelation Estimates
….. Sample estimate of auto covariance
Slide credit: Stochastic Hydrology course, P P Mujumdar
Sample Autocorrelation Estimates
𝟏
𝒓𝒌 ~𝑵 𝟎,
𝒏
Slide credit: Stochastic Hydrology course, P P Mujumdar
Sample Autocorrelation Estimates
Slide credit: Stochastic Hydrology course, P P Mujumdar
Example
Obtain autocorrelation for k=1
σ 𝑥𝑡 − 𝑥ഥ𝑡 𝑥𝑡 − 𝑥𝑡+𝑘
𝑟𝑘 =
2 2
σ 𝑥𝑡 − 𝑥ഥ𝑡 σ 𝑥𝑡 − 𝑥𝑡+𝑘
Example
Example
Correlogram
A completely random series Short-term correlated series
Correlogram
Alternating series Nonstationary series
Stationarity of Stochastic Process
Forecasting is difficult as time series is non-deterministic in nature, i.e. we cannot predict
with certainty what will occur in the future.
But the problem could be a little bit easier if the time series is stationary: you simply predict
its statistical properties will be the same in the future as they have been in the past!
A stationary time series is one whose statistical properties such as mean, variance,
autocorrelation, etc. are all constant over time.
Most statistical forecasting methods are based on the assumption that the time series
can be rendered approximately stationary after mathematical transformations.
Strict Stationarity
A stochastic process 𝑋𝑡 is said to be Strictly Stationary when the joint distribution of
𝑋𝑡1 , 𝑋𝑡2 , … , 𝑋𝑡𝑘 is the same as that of 𝑋𝑡1+𝑠 , 𝑋𝑡2+𝑠 , … , 𝑋𝑡𝑘+𝑠 , for any s, k and t1; · · · tk, ie, the
probability properties of the sequence do not change over time.
However in most applications this stationary condition is too strong.
Weak Stationarity
𝕫: 𝑠𝑒𝑡 𝑜𝑓 𝑖𝑛𝑡𝑒𝑔𝑒𝑟𝑠
In other words, a weakly stationary time series 𝑋𝑡 must have three features:
1. finite variation,
2. constant first moment, and
3. the second moment 𝛾𝑋 𝑠, 𝑡 only depends on 𝑡 − 𝑠 and not depends on s or t.
Usually the term stationary means weakly stationary, and when people want to emphasize a
process is stationary in the strict sense, they will use strictly stationary.
Remarks on Stationarity
Strict stationarity does not assume finite variance thus strictly stationary does
NOT necessarily imply weakly stationary.
A nonlinear function of a strictly stationary time series is still strictly stationary, but
this is not true for weakly stationary.
Weak stationarity usually does not imply strict stationarity as higher moments of
the process may depend on time t.
If time series 𝑋𝑡 is Gaussian (i.e. the distribution functions of 𝑋𝑡 are all
multivariate Gaussian), then weakly stationary also implies strictly stationary. This
is because a multivariate Gaussian distribution is fully characterized by its first two
moments.