Time Series Analysis Course Project
Course: Time Series Analysis
Instructor: Prof. Dr. Md. Rezaul Karim
Department: Statistics and Data Science, Jahangirnagar University
Submitted By: Ishrat Ashique
ID; 2024-1-83-038
Semester: Fall 2025
Objectives of the Study
The primary objective of this assignment is to apply core time series analysis techniques to real-world
economic and business data in order to understand their underlying statistical properties and forecasting
behavior. Specifically, this study aims to:
1. Examine the temporal behavior of multiple economic and business indicators through graphical and
statistical exploration.
2. Identify the presence of trend, seasonality, and irregular components in each time series.
3. Assess the stationarity of the series using both visual inspection and formal unit root tests.
4. Develop appropriate ARIMA and SARIMA models for selected series based on their statistical
characteristics.
5. Evaluate and compare forecasting performance across different modeling approaches, including
exponential smoothing methods.
6. Interpret results critically and provide economically meaningful insights rather than reporting
numerical outputs alone.
Dataset Description
The dataset used in this project consists of monthly observations from January 2019 to December 2023,
resulting in a total of 60 observations per series. The dataset contains four distinct time series representing
key economic and business indicators:
1. Retail Sales – Monthly retail sales measured in million USD.
2. Energy Consumption – Monthly energy usage measured in gigawatt-hours (GWh).
3. Stock Price Index – Monthly values of a stock market index.
4. Unemployment Rate – Monthly unemployment rate expressed as a percentage.
All series are observed at a monthly frequency, making them suitable for seasonal time series analysis with
an annual seasonal period.
Methodological Framework
The analysis follows the exact structure outlined in the assignment instructions and is divided into five main
parts:
1. Exploratory Data Analysis (EDA)
2. Unit Root and Stationarity Testing
3. ARIMA Modeling
4. SARIMA Modeling
5. Exponential Smoothing Models
Each part builds on the results of the previous section, ensuring methodological consistency and logical
progression.
PART 1: EXPLORATORY DATA ANALYSIS
Time Series Plots and Pattern Identification
The first step of the analysis involves plotting each time series to visually inspect its behavior over time. This
allows for preliminary identification of trends, seasonal patterns, cyclical movements, and irregular
fluctuations
Figure 1: Monthly Sales (Retail, Energy,Stock Price Index, Unemployment Rate ) (2019–2023)
The Retail Sales series exhibits a clear upward long-term trend, increasing steadily over the five-year
period. This indicates sustained growth in retail activity rather than short-term volatility alone. In addition
to the trend, the series displays regular seasonal fluctuations, with peaks and troughs recurring at
approximately 12-month intervals.
Notably, the magnitude of seasonal fluctuations increases over time, suggesting that seasonal effects scale
with the level of the series. This behavior implies potential multiplicative seasonality, which has important
implications for model selection in later sections.
Temporary declines are observed at several points; however, these are followed by rapid recoveries,
indicating that shocks to retail sales are largely transitory.
Energy Consumption shows very strong and highly regular seasonal behavior, making it the most seasonal
series in the dataset. Each year exhibits a nearly identical pattern, with clear peaks and troughs occurring at
consistent times.
Alongside seasonality, there is a moderate upward trend, reflecting gradual growth in baseline energy
demand over time. The seasonal component dominates the series, with seasonal variation far exceeding
irregular fluctuations. The regularity and magnitude of seasonal effects suggest that ignoring seasonality
would lead to severe model misspecification.
The Stock Price Index demonstrates a persistent upward movement throughout the sample period, with no
evidence of mean reversion. Unlike Retail Sales and Energy Consumption, the series does not exhibit a
consistent seasonal pattern.
Short-term fluctuations appear irregular and non-repeating, and shocks to the series tend to persist. This
behavior is characteristic of a non-stationary stochastic process, commonly observed in financial time
series.
The Unemployment Rate fluctuates within a relatively narrow range, with no sustained upward or
downward trend across the entire period. A noticeable spike is observed around 2021–2022, after which
the series gradually returns toward its earlier levels.
Seasonal effects are weak and inconsistent, and the overall variability of the series remains stable over time.
Compared to the other series, the Unemployment Rate appears substantially more stable.
Descriptive Statistics
To complement the visual analysis, descriptive statistics were computed for each series, including the mean,
standard deviation, minimum, and maximum values.
Table 1: Descriptive Statistics of the Time Series
The results indicate that:
● Retail Sales and Energy Consumption exhibit higher variability due to trend and seasonality.
● The Stock Price Index shows large dispersion driven by its trending nature.
● The Unemployment Rate has the lowest variability, reflecting relative stability over time.
Autocorrelation and Partial Autocorrelation Analysis
Autocorrelation Function (ACF) and Partial Autocorrelation Function (PACF) plots were generated up to 24
lags for each series.
Figures 3-6: ACF and PACF Plots
● Retail Sales and Energy Consumption show slow decay in ACF, indicating non-stationarity.
● Significant spikes at seasonal lags (12 and 24) are clearly visible for Energy Consumption and Retail
Sales.
● The Stock Price Index exhibits strong persistence across many lags.
● The Unemployment Rate shows faster decay in autocorrelation, consistent with near-stationarity.
Time Series Decomposition
Seasonal-Trend decomposition using Loess (STL) was applied to each series to separate trend, seasonal, and
residual components.
Figures 7–10: STL Decomposition Results
● Retail Sales and Energy Consumption exhibit strong and stable seasonal components.
● The Stock Price Index is dominated by a smooth upward trend.
● The Unemployment Rate shows a weak trend and minimal seasonality.
The residual components across all series are relatively small, indicating that the decomposition successfully
captures the main systematic patterns.
Stationarity and Seasonality Assessment
Based on the graphical analysis:
● Non-stationary series: Retail Sales, Energy Consumption, Stock Price Index
● Approximately stationary series: Unemployment Rate
● Strong seasonal patterns: Energy Consumption (strongest), Retail Sales
These observations motivate the formal stationarity testing conducted in Part 2.
Part 2: Unit Root Testing (15 Points)
The purpose of this section is to formally assess the stationarity properties of each time series using
statistical unit root tests. While visual inspection in Part 1 suggested that most series are non-stationary,
formal testing is necessary to confirm these observations and determine the appropriate level of
differencing required prior to model estimation.
Two complementary tests are employed:
● Augmented Dickey–Fuller (ADF) Test, where the null hypothesis is the presence of a unit root
(non-stationarity).
● KPSS Test, where the null hypothesis is stationarity.
Using both tests provides robustness, as they rely on opposing null hypotheses.
Unit Root Tests on First-Differenced Series
Based on the visual evidence from Part 1, first differencing was applied to all series that appeared
non-stationary. The ADF and KPSS tests were then re-applied to the differenced series.
Series ADF p-value (Diff 1) KPSS p-value (Diff 1)
Retail Sales 9.60 × 10⁻¹¹ 0.100
Energy Consumption 1.04 × 10⁻⁸ 0.100
Stock Price Index 5.39 × 10⁻¹⁵ 0.083
Unemployment Rate 2.77 × 10⁻¹⁴ 0.100
Table 2: ADF and KPSS Test Results on First-Differenced Series
Interpretation
For all four series, the ADF p-values after first differencing are far below the 5% significance level, leading
to rejection of the null hypothesis of a unit root. This indicates that the differenced series are stationary.
Simultaneously, the KPSS p-values exceed 0.05, meaning that the null hypothesis of stationarity cannot be
rejected. The agreement between the ADF and KPSS tests provides strong evidence that first differencing
is sufficient to achieve stationarity for all series.
Conclusion of Unit Root Testing
● All series become stationary after first differencing.
● The appropriate order of differencing for ARIMA modeling is d = 1.
● The results confirm the visual impressions obtained in Part 1 and justify proceeding with ARIMA
model estimation.
Part 3: ARIMA Modeling (25 Points)
Following the assignment instructions, two series were selected for ARIMA modeling. Based on the
outputs provided, the selected series are:
1. Stock Price Index
2. Energy Consumption
These series were chosen because they exhibit distinct dynamics—one dominated by stochastic trend and
the other by strong seasonality—allowing for meaningful comparison.
Model Identification and Candidate Selection
After confirming stationarity through first differencing, candidate ARIMA models were specified based on
standard low-order combinations:
● ARIMA(1,1,1)
● ARIMA(2,1,1)
● ARIMA(1,1,2)
These models were estimated using a training sample, and their performance was evaluated using:
● Akaike Information Criterion (AIC)
● Bayesian Information Criterion (BIC)
● Out-of-sample Root Mean Squared Error (RMSE)
ARIMA Model Comparison: Stock Price Index
Order (p,d,q) AIC BIC RMSE (Test)
(1,1,2) 226.96 234.10 5.80
(2,1,1) 229.04 236.27 5.86
(1,1,1) 230.07 235.49 5.69
Table 3: ARIMA Model Comparison for Stock Price Index
Model Selection
The ARIMA(1,1,2) model yields the lowest AIC and BIC values, indicating superior in-sample fit compared
to alternative specifications. Although RMSE differences across models are small, the information criteria
clearly favor ARIMA(1,1,2). Therefore, ARIMA(1,1,2) is selected as the preferred model for the Stock Price
Index.
Diagnostic Analysis: Stock Price Index
Figure 11: Residuals over Time
The residual plot shows that residuals fluctuate randomly around zero with no visible trend or systematic
pattern. Apart from an initial adjustment spike at the beginning of the sample, residuals remain stable
throughout the period, suggesting that the model captures the underlying data structure effectively.
Figure 12: Residual ACF
The residual ACF shows that all autocorrelations lie within the 95% confidence bounds, indicating the
absence of remaining serial correlation.
Ljung–Box Test
The Ljung–Box test statistic at lag 12 yields a high p-value (p > 0.05), implying that the null hypothesis of no
autocorrelation cannot be rejected. This confirms that the residuals behave like white noise.
Conclusion for Stock Price Index
● ARIMA(1,1,2) adequately captures the dynamics of the series.
● Residual diagnostics support model adequacy.
● The model is suitable for forecasting purposes.
ARIMA Model Comparison: Energy Consumption
Order (p,d,q) AIC BIC RMSE (Test)
(1,1,2) 427.40 434.53 145.86
(1,1,1) 445.40 450.82 148.03
(2,1,1) 448.85 456.07 134.72
Table 4: ARIMA Model Comparison for Energy Consumption
Model Selection
The ARIMA(1,1,2) model achieves the lowest AIC and BIC, indicating better overall fit despite the slightly
lower RMSE observed for ARIMA(2,1,1). Given the assignment’s emphasis on information criteria and model
parsimony, ARIMA(1,1,2) is selected as the preferred non-seasonal model for Energy Consumption.
Diagnostic Analysis: Energy Consumption
Figure 13: Residuals over Time
The residuals fluctuate around zero with no visible trend. However, the magnitude of residual variation is
larger compared to the Stock Price Index, reflecting the strong seasonal structure that a non-seasonal
ARIMA model cannot fully capture.
Figure 14: Residual ACF
The residual ACF shows no significant autocorrelation at non-seasonal lags. Nevertheless, the overall
residual variance remains relatively high, suggesting potential model misspecification due to omitted
seasonal components.
Forecast Behavior
Figure 15: 12-Month Forecast – Energy Consumption (ARIMA)
The forecast converges to a relatively flat path with rapidly widening confidence intervals. This behavior
indicates that the non-seasonal ARIMA model struggles to extrapolate the strong seasonal structure
observed in the historical data.
Conclusion for Energy Consumption (ARIMA)
● Although ARIMA(1,1,2) fits the differenced data reasonably well, it fails to adequately represent
seasonality.
● Forecast uncertainty increases substantially over time.
● These findings strongly motivate the use of SARIMA models, which explicitly incorporate seasonal
dynamics.
Exponential Smoothing Models
The objective of this part is to apply exponential smoothing (ES) techniques to the time series exhibiting the
strongest seasonal behavior and to compare their forecasting performance with the SARIMA model
developed in Part 4.
Based on the exploratory analysis and modeling results from previous sections, Energy Consumption is
selected for exponential smoothing analysis due to its strong and stable seasonal structure.
Exponential Smoothing Model Specification
The following four exponential smoothing models were fitted to the Energy Consumption series:
1. Simple Exponential Smoothing (SES)
– Captures level only
– No trend or seasonality
2. Holt’s Linear Trend Model
– Captures level and trend
– No seasonality
3. Holt–Winters Additive Model
– Captures level, trend, and additive seasonality
– Assumes constant seasonal amplitude
4. Holt–Winters Multiplicative Model
– Captures level, trend, and multiplicative seasonality
– Assumes seasonal effects scale with the series level
These models were evaluated using in-sample goodness-of-fit and information criteria
Model Evaluation and Selection
The models were compared using:
● Akaike Information Criterion (AIC)
● Mean Squared Error (MSE)
● Mean Absolute Error (MAE)
● Mean Absolute Percentage Error (MAPE)
Table 5: Exponential Smoothing Model Comparison for Energy Consumption
Interpretation
The results indicate that:
● Simple Exponential Smoothing performs poorly due to the absence of both trend and seasonal
components.
● Holt’s Linear Trend Model improves performance by accounting for trend but still fails to capture
strong seasonal fluctuations.
● Both Holt–Winters models significantly outperform SES and Holt’s model.
● Among them, the Holt–Winters Additive model achieves the lowest information criteria and error
measures.
Therefore, the Holt–Winters Additive model is selected as the best exponential smoothing specification for
Energy Consumption.
Forecasting Using the Best Exponential Smoothing Model
Using the Holt–Winters Additive model, a 12-month ahead forecast was generated.
The forecast reproduces the historical seasonal pattern effectively, with peaks and troughs occurring at
expected intervals. The seasonal amplitude remains stable, consistent with the additive seasonal
assumption.
Figure 22: 12-Month Forecast – Energy Consumption (Holt–Winters Additive)
Comparison Between Exponential Smoothing and SARIMA Forecasts
To assess relative forecasting performance, the Holt–Winters Additive forecast is compared with the best
SARIMA forecast developed in Part 4.
Figure 23: Forecast Comparison – Energy Consumption (ES vs SARIMA)
Interpretation of Forecast Comparison – Energy Consumption (ES vs SARIMA)
Several important observations emerge from the comparison:
1. Seasonal Structure
Both the Holt–Winters Additive and SARIMA forecasts successfully preserve the strong annual
seasonality observed in the historical data. Seasonal peaks and troughs align closely across the two
methods.
2. Forecast Level
The SARIMA forecast exhibits slightly higher peak values compared to the exponential smoothing
forecast, particularly at the seasonal maximum. This suggests that SARIMA extrapolates recent
upward movements more aggressively.
3. Smoothness vs Flexibility
The exponential smoothing forecast is smoother and more stable, reflecting its reliance on
weighted averages of past observations. In contrast, SARIMA produces sharper seasonal turning
points, reflecting its parametric structure.
4. Uncertainty Representation
While both models generate plausible forecasts, SARIMA explicitly provides confidence intervals
grounded in statistical theory. Exponential smoothing, although effective, does not offer the same
level of inferential support.
Model Preference and Forecast Reliability
From a forecasting perspective:
● SARIMA is preferred when:
○ Statistical inference is required
○ Confidence intervals are important
○ Seasonality is stable and well-defined
● Exponential Smoothing (Holt–Winters) is preferred when:
○ Simplicity and interpretability are prioritized
○ Rapid operational forecasting is needed
○ The goal is accurate short-term prediction rather than inference
In this study, both models perform well; however, SARIMA is considered more reliable overall due to its
stronger theoretical foundation and diagnostic validation.
Does the Best Model Match Expectations?
Yes, the results align with expectations. Given the strong and regular seasonal behavior of Energy
Consumption observed in Part 1, models explicitly incorporating seasonality—SARIMA and Holt–Winters,
naturally outperform non-seasonal approaches.
Conclusion
This project systematically applied time series analysis techniques to four economic and business indicators.
The analysis demonstrates that:
● Most real-world economic time series are non-stationary and require differencing.
● Seasonality plays a dominant role in certain series, particularly Energy Consumption and Retail
Sales.
● ARIMA models perform well for non-seasonal series such as the Stock Price Index.
● SARIMA and Holt–Winters models significantly improve forecasting accuracy for strongly seasonal
data.
● Model diagnostics and careful interpretation are essential for reliable forecasting.
Limitations and Recommendations
● The analysis is based on a relatively short time span (60 observations).
● Structural breaks and external regressors were not explicitly modeled.
● Future studies could extend this framework using multivariate models or longer datasets.
Remarks
This assignment highlights the importance of matching modeling techniques to the statistical properties of
the data. Through a structured, step-by-step approach, the study demonstrates how rigorous time series
analysis can yield meaningful insights and reliable forecasts for real-world economic indicators.