A PROJECT REPORT ON
“FORECASTING OF RAINFALL IN INDIA”
A PROJECT REPORT
SUBMITTED FOR THE PARTIAL FULFILMENT OF THE [Link].
IN STATISTICS 2016-17
UNDER SUPERVISION OF: SUBMITTED BY:
Prof. Nirpeksh Kumar AVINASH KUMAR
Department of Statistics [Link]. (Statistics)
Institute of Science Roll No. – 15420STA006
Enrollment No. – 342459
DEPARTMENT OF STATISTICS
BANARAS HINDU UNIVERSITY
VARANASI - 221005
ACKNOWLEDGMENT
I have put efforts and dedication in this Project Work. However, it
would not have been possible without the kind support and help of
many individuals and organizations. I would like to extend my sincere
thanks to all of them.
It gives me satisfaction and great pleasure to express my sincere
respect gratitude and indebtedness to the honourable teacher and
supervisor Prof. Nirpeksh Kumar, Department of Statistics, Institute
of Science, Banaras Hindu University, for his enthusiastic &
persistent help, constant encouragement and priceless supervision
without which it would not have been possible for me to complete this
project work. The guidance and valuable opinion that I received from
him, during the entire period of this work, has been a great help in the
completion of this work.
I also want to express deep respect to honourable teachers, Dr. B. P
Singh, Dr. S. K. Upadhyay, Dr. Umesh Singh, Dr. K. K. Singh, Dr.
Rajesh Singh Dr. S. Kumar, Dr. Sushma Tripathi, Dr. usha Srivastava
Department of Statistics, for their valuable suggestions and hope they
will continue their guidance in future also. I also want thanks to
research scholars who co-operated me in Project Work Analysis
I am very thankful to my parents for their invaluable motivational and
emotional help in this project, without them it would have not been
easier. I also express my sincere thanks to my friends for their healthy
co-operation.
At last but not the least I thank the Department of Statistics, [Link]
which gave me the opportunity for this project work and for which I
shall ever remain grateful.
AVINASH KUMAR
Contents
1. Abstract
2. Introduction
3. Time Series
4. Components of Time Series
5. Estimation & Elimination of the
Trend and Seasonal components
6. Forecasting & its methodology
7. Testing of Result & Hypothesis
8. Model Identification
9. Result and Conclusion
10. Appendix
1. ABSTRACT
Rainfall prediction is one of the most important and challenging task
in the modern world. In this report we asses and forecast the data of
rainfall in INDIA. The whole project is based on the data available
on the government website. The data taken for the analysis work is
from 1991 – 2015. In this report we try to investigate the trend and
seasonality of the rainfall. The sole purpose of the project is to
forecast the amount of monthly rainfall of INDIA. The method of
forecasting is based on the principles of Holtz-Winter forecasting
process. We also try to fit the given data and find out the trend line.
Further the process of Hypothesis testing is carried out for the
obtained results. We also try to fit the data to a specific Time Series
Model on the basis of ACF & PACF. The result obtained through
this project work can be used as a help in planning the future
activities related to agriculture and other works as well as it could
guide a foreign tourist about the climatic conditions during the
various time periods throughout the year.
The given historical data given helped us to obtain the rainfall
forecasting of next three years. The monthly forecasted data shows a
similar pattern to the historical data. The three months July, August
& September receives the high amount of rainfall as these months
belong to the monsoon period.
Model identification was by visual inspection of both the sam
ple ACF and sample PACF to postulate many possible models and the
n use the model selection criterion of Akaike’s Information Criterion
[AICc] to choose the best model. The model with least value of AICc
is considered to be the best model. The chosen model was Seasonal
ARIMA (2, 0, 1) (0, 0, 2) [12] with non-zero mean with AICc score
of 3039.83.
2. INTRODUCTION
Rainfall is the primary source of water and is of great importance for
India's Economy, specially its agriculture industry. It is highly
variable over space and time, leads to flood and drought every year
on one or the other part of the country. Rainfall statistics is therefore
required by the Policy Makers, Planners, Design Structure Engineers,
Agriculturists, Hydrologists, Research Scholars and many more for
proper management of water resources.
Accurate measurement and prediction of rainfall is very important
and helpful for human life & its development.
Average annual rainfall of India is about 1190 mm. Accurate
information on rainfall is essential for the planning and
management of water resources and also crucial for reservoir
operation and flooding prevention.
This Project study is carried out using the Time Series
analysis and its methodologies. Time series analysis comprises
methods for analysing time series data in order to extract
meaningful statistics and other characteristics of the data. Time
series forecasting is the use of a model to predict future values
based on previously observed values. The data we analyse is the
weighted monthly rainfall of India from Jan 1991 to Dec 2015. The
data source is Indian ‘Government Website’. The objective of this
Project is to make the prediction of rainfall more accurate in the
recent future. The report has been constructed with different
chapters highlighting about the Time series and its components,
smoothing, forecasting, models & also about the methodology of
Hypothesis testing. And lastly the findings & results of the Project
are embedded.
3. TIME SERIES
Data obtained from observations collected sequentially over time are
extremely common. In business, we observe weekly interest rates,
daily closing stock prices, monthly price indices, yearly sales figures,
and so forth. In meteorology, we observe daily high and low
temperatures, annual precipitation and drought indices, and hourly
wind speeds. In agriculture, we record annual figures for crop and
livestock production, soil erosion, and export sales. In the biological
sciences, we observe the electrical activity of the heart at millisecond
intervals. In ecology, we record the abundance of an animal species .
These sequentially collected data over time are generally referred as
Time Series Data.
A time series is a series of data points indexed (or listed or
graphed) in time order. Most commonly, a time series is a sequence
taken at successive equally spaced points in time. Thus it is a
sequence of discrete-time data.
Time series are very frequently plotted via line charts. Time
series are used in statistics, economics, mathematical finance ,
weather forecasting, control engineering, astronomy, communications
engineering, and largely in any domain of applied science and
engineering which involves temporal measurements.
Time series analysis comprises methods for analysing time
series data in order to extract meaningful statistics and other
characteristics of the data. Time series forecasting is the use of
a model to predict future values based on previously observed values.
Time series analysis can be applied to real-valued, continuous
data, discrete numeric data, or discrete symbolic data i.e. sequences of
characters, such as letters and words in the English language.
There are a number of things which are of interest in time series
analysis. The most important of these are:
0 Smoothing 0 Modelling 0 Forecasting 0 Control
3.1 Stationarity and Non-Stationarity
A key idea in time series is that of stationarity. Roughly speaking, a
time series is stationary if its behaviour does not change over time.
This means, for example, that the values always tend to vary about the
same level and that their variability is constant over time. Stationary
series have a rich theory and their behaviour is well understood. This
means that they play a fundamental role in the study of time series.
Obviously, not all time series that we encounter are stationary. Indeed,
non-stationary series tend to be the rule rather than the exception.
However, many time series are related in simple ways to series which
are stationary
3.2 Some Examples
1. Annual Rainfall
The annual amount of rainfall in a country for the years from
1949 to 2000. The general pattern of rainfall looks similar
throughout the record, so this series could be regarded as being
stationary.
2. Random Walk
Examples of non-stationary processes are random walk with or
without a drift (a slow steady change) and deterministic trends.
A random walk is a mathematical object, known as a stochastic
or process, that describes a path that consists of a succession
of random steps on some mathematical space such as the
integers. The example of non-stationary processes is: -
The monthly number of housing starts in the Unites States (in
Thousands). Housing starts are a leading economic indicator.
This means that an increase in the number of housing starts
indicates that economic growth is likely to follow and a decline
in housing starts indicates that a recession may be on the way.
NOTE:- (FOR THIS PROJECT)
This project is based on the Rainfall data of India over the period of
1991 - 2015 and the Time Series data selected in the Project is a
Stationary data as general pattern of rainfall looks similar throughout
each year. We will be analysing and forecasting the data on R-studio.
The Data obtained are as follows:
DATA
Year Jan Feb Mar Apr May Jun Jul Aug Sep Oct Nov Dec
1991 14.3 28.1 27.8 51.7 68.9 184.7 279.2 268.1 140.7 61.8 30.2 14.7
1992 16 16.5 24.8 26.1 59.3 139.7 262.5 274 171.7 64.7 41.6 5.6
1993 18.2 25.6 41.6 27 71.3 172.1 305.4 203.2 208.5 87.9 30.5 16.5
1994 25 27.9 25.2 45.9 53.1 205.7 350 282.2 149.4 82.8 25.5 22.6
1995 31.3 29.4 28.3 32.4 82.4 143.3 323.4 269 179 78 36.8 9.2
1996 22.9 23.2 32.1 31.4 56 185.7 262.1 292.4 146.1 100.5 13.6 16.9
1997 14.3 10.4 30.3 46 48.6 171.7 281.5 261.9 151.4 61.1 57.6 48.3
1998 16.4 28.2 39.1 36.3 49.2 163.9 278.4 243.8 196.5 107.4 39.3 10.3
1999 13.7 11.2 8.8 19.3 94.9 169.9 261.7 213.2 183 117.2 20 3.7
2000 18.4 28.2 17.9 34.7 71.6 179 263.5 221.1 134.5 41.9 14.6 10
2001 7.3 8.8 18.8 43.7 67.2 219.1 279.7 209.4 114.2 107.5 22.5 7.1
2002 15.7 20.3 21.5 38.3 63.3 180.1 146.2 262.5 151.2 59.5 18.2 5.1
2003 7.6 45.6 33.3 35.4 39.1 184.5 316.7 255.3 191.4 100.6 15.5 18.6
2004 25.7 8.8 11.4 59 88.9 158.7 242.1 248.7 124.6 92.2 15.8 4.6
2005 28.1 41.8 42.5 37.7 46.1 143.2 334.1 190.1 206.9 99.3 27.2 11.2
2006 17.7 11.9 35.6 32.7 75 141.8 287.6 281.3 178.6 51.8 34.6 13.1
2007 1.7 36.7 35.2 30.6 46.8 192.5 286.2 257.4 206.8 55.7 14.4 15.3
2008 18.4 19.3 41.2 29.5 43.7 202 245 265.8 165.1 51.6 25.5 11
2009 12 12 14.2 25.1 56 85.7 280.7 192.5 139.4 71.4 53.7 11.1
2010 7.5 17 14 39 73.8 138.1 300.7 274.7 197.7 69 61.4 22.7
2011 6.8 25.8 22.4 41.1 53.1 183.5 246 284.9 186.9 38.1 20.1 7.6
2012 26.5 12.7 11.3 47.5 31.7 117.8 250.2 262.4 193.5 58.7 30.7 11.7
2013 11.3 40.1 15.7 30.4 57.8 219.8 310 254.7 152.7 129.4 14 6.7
2014 19.2 27.4 36.1 22.2 72.9 95.4 261.2 237.5 188 60.2 14.4 10.7
2015 19.2 22.2 30.9 38.3 62.3 163.6 289.2 261.3 173.4 80.9 29.7 16.6
Plot:-
4. COMPONENTS OF TIME SERIES
The factors that are responsible to bring about changes in a time
series, also called the components of time series, are as follows:
1. Secular Trend
2. Seasonal Movements
3. Cyclical Movements
4. Irregular Fluctuations
Secular Trend:
The secular trend is the main component of a time series which results
from long term effect of socio-economic and political factors. This
trend may show the growth or decline in a time series over a long
period. This is the type of tendency which continues to persist for a
very long period. Prices, export and imports data, for example, reflect
obviously increasing tendencies over time.
Seasonal Trend:
These are short term movements occurring in a data due to seasonal
factors. The short term is generally considered as a period in which
changes occur in a time series with variations in weather or festivities.
For example, it is commonly observed that the consumption of ice-
cream during summer us generally high and hence sales of an ice-
cream dealer would be higher in some months of the year while
relatively lower during winter months. Employment, output, export
etc. are subjected to change due to variation in weather. Similarly
sales of garments, umbrella, greeting cards and fire-work are
subjected to large variation during festivals like Republic Day, Eid,
Christmas, New Year etc. These types of variation in a time series are
isolated only when the series is provided biannually, quarterly or
monthly.
Cyclic Movements:
These are long term oscillation occurring in a time series. These
oscillations are mostly observed in economics data and the periods of
such oscillations are generally extended from five to twelve years or
more. These oscillations are associated to the well-known business
cycles. These cyclic movements can be studied provided a long series
of measurements, free from irregular fluctuations is available.
Irregular Fluctuations:
These are sudden changes occurring in a time series which are
unlikely to be repeated, it is that component of a time series which
cannot be explained by trend, seasonal or cyclic movements. It is
because of this fact these variations some-times called residual or
random component. These variations though accidental in nature, can
cause a continual change in the trend, seasonal and cyclical
oscillations during the forthcoming period. Floods, fires, earthquakes,
revolutions, epidemics and strikes etc. are the root cause of such
irregularities.
Graphs:-
5. Estimation & Elimination of the trend and
Seasonal component
Consider the “classical decomposition” model
Xt = mt + st + Yt ,
Where
mt is a slowly changing function (the “trend component”);
st is a function with known period d (the “seasonal component”);
Yt is a stationary time series.
Our aim is to estimate and extract the deterministic components mt
and st in hope that the residual component Y t will turn out to be a
stationary time series.
5.1 No Seasonal Component
Assume that
Xt = mt + Yt, t = 1, . . . , n
Where, without loss of generality, EYt = 0.
Method 1 (Least Squares estimation of mt)
If we assume that mt = a0 + a1t + a2t2 we choose a to minimize
Method 2 (Smoothing by means of a moving average)
Let q be a non-negative integer and consider
Provided
q is so small that mt is approximately linear over [t − q, t + q]
and q is so large that
For t < q and t > n − q some modification is necessary, e.g.
The two requirements on q may be difficult to fulfil in the same time.
Let us therefore consider a linear filter
Where ∑aj = 1 and aj = a−j . Such a filter will allow a linear trend to
pass without distortion since
In the above example we have
It is possible to choose the weights {aj} so that a larger class of trend
functions pass without distortion. Its example is the Spencer-15point
moving average where
5.2 Trend and Seasonality
Let us go back to
Xt = mt + st + Yt,
Where EYt = 0, St+d = st and ∑ =0. For simplicity we assume that n/d
is an integer.
Typical values of d are:
• 24 for period: day and time-unit: hours;
• 7 for period: week and time-unit: days;
• 12 for period: year and time-unit: months.
Sometimes it is convenient to index the data by period and time-unit
xj,k = xk+d(j−1), k = 1,...,d, j = 1,...,n/d.
i.e. xj,k is the observation at the k:th time-unit of the j:th period.
Method S1 (Small trends)
It is natural to regard the trend as constant during each period, which
means that we consider the model
Method S2 (Moving average estimation)
First we apply a moving average in order to eliminate the seasonal
variation and to dampen the noise. For d even we use q = d/2 and the
estimate
and for d odd we use q = (d − 1)/2 and the estimate
In order to estimate s we first form the “natural” estimates
(Note that it is “only” the end-effects that force us to this formally
complicated estimate. What we really are doing is to take the average
of those xk+jd: s where the mk+jd: s can be estimated.)
NOTE: -
A seasonal time series consists of a trend component, a seasonal
component and an irregular component. Decomposing the time series
means separating the time series into these three components: that is,
estimating these three components.
To estimate the trend component and seasonal component of a seasonal
time series that can be described using an additive model, we can use
the “decompose ()” function in R. This function estimates the trend,
seasonal, and irregular components of a time series that can be
described using an additive model.
Decomposition is an important technique for all types of time series
analysis, especially for seasonal adjustment. It seeks to construct,
from an observed time series, a number of component series such as
Trend component, seasonal component and Random component.
PLOT: - DECOMPOSITION OF TIME SERIES
6. Forecasting & its methodology
A common goal of time series analysis is extrapolating past behaviour
into the future. This process is known as Forecasting. Forecasting is
the process of making predictions of the future based on past and
present data and most commonly by analysis of trends. A
commonplace example might be estimation of some variable of
interest at some specified future date. For example: - Prediction of
rainfall or Prediction of temperature.
Mostly, we assume that the model is exactly known, including the
specific values for all parameters. Although this is never true in
practice, for large sample size, the use of estimated parameters doesn’t
seriously affect forecast.
Type of Forecasts
Forecasts can be obtained in different ways:-
Qualitative
These approaches are based on judgments and opinions. Here are four
examples.
1. Ask the guy in contact with the customers. Compile the results level
by level.
2. Ask a panel of people from a variety of positions. Derive the
forecasts and submit again. An example is the Delphi method used by
the Rand Corporation in the 1950s.
3. Perform a customer survey (questionnaire or phone calls).
4. Look how similar products were sold. For example, the demand for
CDs and for CD players are correlated. Washing machine and dryers
are also related.
Time series analysis
What we will analyse in details. The idea is that the evolution in the
past will continue into the future.
Time series:
Stationary
Trend-based
Seasonal
Different time series will be considered: stationary, trend-based and
seasonal. They differ by the shape of the line which best fits the
observed data.
Methods:
Moving average
Regression
Exponential smoothing
The methods which can be used are (linear) regressions, moving
averages and exponential smoothing. They differ by the importance
they give to the data and by their complexity.
Moving Average
If a time series is generated by a constant process subject to random
error, then mean is a useful statistic and can be used as a forecast for
the next period.
Averaging methods are suitable for stationary time series data where
the series is in equilibrium around a constant value (the underlying
mean) with a constant variance over time
This method uses the average of all the historical data as the forecast
1 t
Ft 1 yi
t i 1
When new data becomes available, the forecast for time t+2 is the
new mean including the previously observed data plus this new
observation.
1 t 1
Ft 2 yi
t 1 i 1
This method is appropriate when there is no noticeable trend or
seasonality.
Forecasting with regression
Forecasts from a simple linear model are easily obtained using the
equation
Where x is the value of the predictor for which we require a forecast.
That is, if we input a value of x in the equation we obtain a
corresponding forecast ŷ.
When this calculation is done using an observed value of x from the
data, we call the resulting value of ŷ a “fitted value”. This is not a
genuine forecast as the actual value of y for that predictor value was
used in estimating the model, and so the value of ŷ is affected by the
true value of y. When the values of x is a new value (i.e., not part of
the data that were used to estimate the model), the resulting value of ŷ
is a genuine forecast.
Assuming that the regression errors are normally distributed, an
approximate 95% forecast interval (also called a prediction interval)
associated with this forecast is given by
Where N is the total number of observations, ẍ is the mean of the
observed x values, sx is the standard deviation of the observed x
values and se is given by
Forecasting with Exponential Smoothing
Simple Exponential Smoothing
Simple exponential smoothing is used when data do not have
underlying trend and seasonality and are mostly stationery. The
formula for exponential smoothing methods is
Ft+1 = α yt + (1-α) Ft
Where, Ft old forecast for period t & Ft+1 is forecast for the next
period. yt is observed value of series in period t and α is the
Smoothing constant, whose value can be from 0 to 1.
The smoothed estimate is equal to old estimate plus some adjustment.
The forecast Ft+1 is based on weighting the most recent observation yt
with a weight and weighting the most recent forecast F t with a
weight of 1-
The exponential smoothing equation can be rewritten in the following
form to elucidate the role of weighting factor
Ft+1 = Ft + α (yt - Ft)
Holt’s Exponential smoothing
Holt’s two parameter exponential smoothing method is an extension
of simple exponential smoothing.
It adds a growth factor (or trend factor) to the smoothing equation as a
way of adjusting for the trend. Three equations and two smoothing
constants are used in the model.
The forecasting equation for this model is given by
Xn+h = an+bnh
Where an is estimate of level and bn is estimate of trend slope.
The smoothing equations are given by
an+1 = αxn+1 + (1-α)(an+bn)
bn+1 = β(an+1 – an) + (1-β)bn
and the assumed initial conditions are
a2=x2
b2=x2-x1
NOTE:- α & β are independent and lie between 0 & 1.
Holt-Winter seasonal Exponential smoothing
Winter’s exponential smoothing model is the second extension of the
basic Exponential smoothing model. It is used for data that exhibit
both trend and seasonality with period ‘d’. It is a three parameter
model that is an extension of Holt’s method. An additional equation
adjusts the model for the seasonal component.
The forecasting equation for this model is given by
Xn+h = an+bnh+cn+h
The smoothing equations are given by
an+1 = α (xn+1 - cn+1-d)+(1-α)(an+bn)
bn+1 = β (an+1 – an)+(1-β)bn
cn+1 = γ (xn+1- an+1)+(1-γ)cn+1-d
and the assumed initial conditions are
ad+1 = xd+1
bd+1 = (xd+1 - x1)/d
ci = xi – (x1+bd+1*x(i-1))
Where , & are constants which lies between 0 & 1.
= smoothing constant for the data.
= smoothing constant for trend estimate.
= smoothing constant for seasonality estimate
, & are chosen such that measure of forecast error is minimized.
The smoothing constant smoothes the data to eliminate randomness.
The smoothing constant smoothes the trend in the data set.
The smoothing constant smoothes the seasonality in the data.
The Values of , & in this Project has been obtained using the
Rstudio.
The obtained values are as follows
: 0.05925109
: 0.02174705
: 0.105429
These values are used in the smoothing equations to obtain the
forecasting values.
In this Project we have forecasted the value of rainfall of upcoming 36
months i.e. from Jan 2016- Dec 2018.
The obtained forecasted values are as follows:-
Month Rainfall
Jan 2016 18.17898
Feb 2016 25.42079
Mar 2016 27.71837
Apr 2016 36.80213
May 2016 61.63054
Jun 2016 159.45962
Jul 2016 278.30266
Aug 2016 255.08750
Sep 2016 173.38477
Oct 2016 75.90044
Nov 2016 29.65518
Dec 2016 14.94399
Jan 2017 18.69667
Feb 2017 25.93848
Mar 2017 28.23607
Apr 2017 37.31982
May 2017 62.14823
Jun 2017 159.97732
Jul 2017 278.82035
Aug 2017 255.60519
Sep 2017 173.90246
Oct 2017 76.41814
Nov 2017 30.17288
Dec 2017 15.46169
Jan 2018 19.21436
Feb 2018 26.45618
Mar 2018 28.75376
Apr 2018 37.83752
May 2018 62.66593
Jun 2018 160.49501
Jul 2018 279.33805
Aug 2018 256.12289
Sep 2018 174.42016
Oct 2018 76.93583
Nov 2018 30.69057
Dec 2018 15.97938
Graph:- Blue line shows the forecasted data & shaded region shows
the upper and lower limits of forecasting at 95% level of significance.
7. Testing of Result & Hypothesis
In this section we will be testing the forecasted results and some
hypothesises. The main idea of testing is to validate the results
obtained and Project Work findings.
Testing of Result
In testing of result we wish to know that the forecasted result is
accurate or not. For this purpose we will omit the last 2 year data from
the main data and we will then forecast those 2 year data. Then we
will compare the actual data and forecasted data to know if the
forecasting is accurate or not.
The process of omission and forecasting of data from 2014-2015 is
carried out in the R-studio.
The obtained results are as follows:-
Month Actual Forecasted
Jan-14 19.2 11.5
Feb-14 27.4 16.8
Mar-14 36.1 20.1
Apr-14 22.2 32.7
May-14 72.9 52.6
Jun-14 95.4 152.3
Jul-14 261.2 267.5
Aug-14 237.5 249.6
Sep-14 188 169
Oct-14 60.2 63
Nov-14 14.4 26.6
Dec-14 10.7 8.5
Jan-15 19.2 10.8
Feb-15 22.2 16.1
Mar-15 30.9 19.5
May-15 62.3 51.9
Jun-15 163.6 151.7
Jul-15 289.2 266.9
Aug-15 261.3 248.9
Sep-15 173.4 168.4
Oct-15 80.9 62.4
Nov-15 29.7 25.9
Dec-15 16.6 7.9
Total 2232.8 2132.6
From the above table we can see that the difference between actual
data and forecasted data is very minute. For further research we check
the graphs and the R2 values.
Graph:-
In the above graph we can see that both actual and forecasted are
almost same as they are both overlapping and thus we can say that the
forecasting is accurate.
Testing of Hypothesis
Hypothesis is a statement or assumption about population. It is usually
concerned with the parameter of the population about which the
statement is made. A statistical hypothesis is a hypothesis that is
testable on the basis of observing a process that is modelled via a set
of random variables. Statistical hypothesis is composed of two types:
Null hypothesis( Ho): It is the particular hypothesis under test,
and it is the hypothesis of “no difference”
Alternative hypothesis (HA): which disagree with the null
hypothesis
There are five steps in hypothesis testing:
1. Making assumptions
2. Stating the research and null hypotheses and selecting (setting)
alpha
3. Selecting the sampling distribution and specifying the test
statistic
4. Computing the test statistic
5. Making a decision and interpreting the results
Test statistic is a mathematical expression of sample values which
provides a basis for testing a statistical hypothesis.
The result of this test determines whether we will accept the null
hypothesis and so the H A will be rejected, or we reject the null
hypothesis and so the HA will be accepted.
Now Hypothesis testing for this Project:-
To find out the suitable model for the derived data, first we have to
ensure that the data at hand is stationary or not. For this purpose we
wish to test the hypothesis for stationarity of data.
For testing purpose we use Augmented Dickey–Fuller test.
An Augmented Dickey–Fuller test (ADF) is a test for a unit root in a
time series sample. It is an augmented version of the Dickey–Fuller
test for a larger and more complicated set of time series models. The
augmented Dickey–Fuller (ADF) statistic, used in the test, is a
negative number. The more negative it is, the stronger the rejection of
the hypothesis that there is a unit root at some level of confidence .
The unit root test is then carried out under the
Null hypothesis, H0: ϒ=0 against
Alternative hypothesis H1: ϒ<0
In more common way we can say about hypotheses for the test:
The null hypothesis for this test is that there is a unit root.
The alternate hypothesis differs slightly according to which
equation you’re using. The basic alternate is that the time series is
stationary.
Test statistic:
DFτ = ϒ(hat)/SE(ϒ (hat))
Further calculations are carried using R-studio.
Augmented Dickey-Fuller Test
Data: desesnlizd_detrndd_dt
Dickey-Fuller = -5.3657, Lag order = 6, p-value = 0.01
Result:-
Alternative hypothesis: stationary
We can see from above that our null hypothesis is rejected and
alternate one is selected. Thus we can state that the time series
data is Stationary.
8. Model Identification
Now, we proceed to check the model of data to which it belong.
The process of time series modeling begins with the selection of the
preliminary models interpreted from the characteristics of ACF and
PACF functions.
The ACF and PACF functions are as follows:-
The figures of ACF and PACF show that rainfalls are periodic in
nature. These functions behave similarly in their period cycles
involving seasonal variations. Therefore given time series is periodic
and involve seasonal variations. So we examined the ACF and PACF
of the 12 th differences (seasonal differencing).
Based on the features portrayed by the plots, we expect a Seasonal
ARIMA Model of the form: - SEASONAL ARIMA (p, d, q) (P, D, Q) [12]
The estimated order of the model parameters p, q, P and Q are
identified by visual inspection of ACF and PACF of the stationary
process of rainfall. The models that have the lowest AICc values tend
to give slightly better results than the other models . Thus the best
model selected have the lowest AICc value.
Further the model identification process is carried out with the help of
R-studio.
The results obtained are as follows:-
Series: raintimeseries
ARIMA (2, 0, 1) (0, 0, 2) [12] with non-zero mean with AICc score
of 3039.83
The model parameters are:-
ar1 ar2 ma1 sma1 sma2
1.5010 -0.7514 -0.9038 0.4071 0.4284
s.e. 0.0406 0.0415 0.0309 0.0591 0.0546
Where,
ar1 & ar2 represent Auto Regressive models.
ma1 & ma2 represent Moving Average models.
sma1 & sma2 represent Moving Average seasonal model.
Residuals Plots for SARIMA (2, 0, 1) (0, 0, 2) [12]
9. Results and Conclusions
The data for this Project has been obtained from the website of
Indian Meteorological Department. The key findings and results of
this project are as follows:-
1 The observed data is periodic in nature with d=12. Total of 300
observations were obtained i.e. 25 years of data is obtained.
2 The data observed is stationary in nature.
3 Forecasting of data is carried out using the Exponential forecasting
method.
4 The data contains seasonality with no trend hence we apply Holt-Winter
seasonal Exponential smoothing for forecasting.
5 The smoothing constants α, β & are selected in a way such that it
minimizes the error.
The obtained values are as follows
α: 0.05925109
β : 0.02174705
: 0.105429
6 Augmented Dickey–Fuller test verifies our hypothesis that data is
stationary.
7 Forecasting of rainfall is obtained for next 36 months.
8 The known average total rainfall of India is 1190 mm and for the
9 forecasted years total rainfall of India was obtained as follows:-
2016 - 1156.485mm
2017 - 1162.697mm
2018 - 1168.910mm
10 The given data set follows ARIMA (2, 0, 1) (0, 0, 2) [12] with non-zero
mean.
11 On comparing the actual data with predicted data we find predicted data
to be almost accurate.
The findings and results lead us to conclude that seasonal ARIMA model can be
used. The identified model is ARIMA (2, 0, 1) (0, 0, 2) [12].We also conclude
that the rainfall can be predicted well before in advance which can be helpful for
farmers, seasonal businesses and also to governmental bodies for planning and
management purposes.
10. Appendix
Programme for Forecasting of Rainfall
getwd()
rain=[Link]("[Link]")
rain
raintimeseries <- ts(rain, frequency=12, start=c(1991,1))
raintimeseries
[Link](raintimeseries)
raintimeseriescomponents <- decompose(raintimeseries)
raintimeseriescomponents$seasonal
raintimeseriescomponents$trend
raintimeseriescomponents$random
plot(raintimeseriescomponents)
library(forecast)
raintimeseries <- raintimeseries
raintimeseriesforecasts <- HoltWinters(raintimeseries)
raintimeseriesforecasts
plot(raintimeseriesforecasts)
raintimeseriesforecasts2 <-
[Link](raintimeseriesforecasts, h=36)
raintimeseriesforecasts2
[Link](raintimeseriesforecasts2)
Programme for ACF & PACF (With & Without Seasonal
Differencing) and Model Identification
getwd()
rain=[Link]("[Link]")
rain
raintimeseries <- ts(rain, frequency=12, start=c(1991,1))
raintimeseries
[Link](raintimeseries)
acf(raintimeseries, [Link] = 20)
acf(raintimeseries, [Link] = 20, plot=FALSE)
pacf(raintimeseries, [Link]=20)
pacf(raintimeseries, [Link]=20, plot=FALSE)
seasonal<-diff(raintimeseries, lag=12, differences=1)
seasonal
plot(seasonal)
acf(seasonal, [Link] = 20)
acf(seasonal, [Link] = 20, plot=FALSE)
pacf(seasonal, [Link]=20)
pacf(seasonal, [Link]=20, plot=FALSE)
library(forecast)
[Link](raintimeseries, stepwise = FALSE, approximation =
FALSE)