Demand Forecasting Rsearch Paper
Demand Forecasting Rsearch Paper
B. Research Gap
I. INTRODUCTION
Even with these improvements, there are still some big
A. Background and Motivation problems that need to be fixed:
In today's business world, supply chain management is a Limitations of single models: Individual models,
moving target; it is becoming more complicated, unstable, whether they are statistical or neural, work well in
and unpredictable [1], [2]. It's clear that demand is the most some situations but not in others[13].
important part of good inventory management, production Error systematicity: Deep learning models often
planning, and logistics management[3]. Real-world demand make systematic mistakes when predicting things that
is messy, though. It doesn't always stay the same; it has simple ensembling can't fix [14].
several seasonal cycles; promotional activity can create extra Sensitivity to volatility: Most forecasting methods
lift; and it can change quickly in response to outside don't work as well on demand streams with a lot of
shocks[4]. volatility, which is where accurate predictions are most
important [15].
Static ensembling: Traditional ensemble methods use piecewise-linear dynamics. More complex nonlinear
fixed weighting schemes that don't change when the interactions can make them less useful[20].
accuracy of the forecast changes over time [16].
C. Research Contribution B. Deep Learning for Time Series
This study tackles these issues by suggesting a new The application of deep learning for time series prediction
ensemble architecture that is aware of errors and uses the has evolved over the years with the development of new
best parts of attention-based deep learning and gradient architectures:
boosting. At its core, PatchTST has two parts: it finds hidden Recurrent Neural Networks: Standard RNNs had a
patterns, and a residual learning module controlled by problem with vanishing gradients, but LSTM networks [9]
XGBoost is in charge of finding and fixing systematic errors. and Gated Recurrent Units [21] solved it. This made it
This is different from traditional ensemble methods that just possible to learn long-range relationships. DeepAR [22]
average predictions. We use a validation-adaptive weighting improved LSTMs by adding probabilistic forecasting and
mechanism that changes the relative weight of each autoregressive likelihood modeling.
component in real time to keep our system working well
even when the market changes. We test the proposed method
Convolutional Architectures: Temporal Convolutional
on 127 real-world product-location time series that cover a
Networks [23] used dilated causal convolutions to get large
wide range of volatility conditions. The final solution is built
receptive fields while still being fast to compute. WaveNet
to be able to handle the amount of computation needed to be
[24] demonstrated the efficacy of stacked dilated
used in a real-world supply chain forecasting system.
convolutions in generating sequential data.
The rest paper is structured as follows: Section II reviews
related work in time series forecasting and ensemble Models Based on Attention: The Transformer architecture
methods. Section III presents the proposed error-adaptive [11] changed how sequences are modeled by using self-
ensemble framework, including mathematical formulations attention processes that find global dependencies in a
and algorithmic details. Section IV describes the consistent amount of time. Temporal Fusion Transformers
experimental methodology, dataset characteristics, and [25] changed transformers so that they could be used for
evaluation protocols. Section V presents comprehensive multi-horizon forecasting by adding attention patterns that
results and analysis. Section VI discusses theoretical can be understood. ProbSparse self-attention was created by
insights, practical implications, and limitations. Section VII Informer [26] to make it easier for computers to predict long
concludes with future research directions. sequences.
II. RELATED WORK
Right now, PatchTST [12] is the best patch-based
A. Classical Time Series Forecasting transformer. It splits time series into pieces that can be used
For a long time, the best way to predict time series has as input tokens for transformer encoders. This method
been to use classic statistical methods. The Box-Jenkins shortens the length of the sequence, lets you learn from other
methodology[5] produced the ARIMA models, which, in tasks, and improves both accuracy and speed compared to
theory, divide a series into three parts: autoregressive, pointwise transformer inputs.
integrated, and moving average. After that, SARIMA[17]
came along, which combined seasonal patterns by using C. Ensemble Methods in Forecasting
seasonal differencing and seasonal AR/MA terms.
Using variety and aggregation, ensemble learning
combines different models to get better results [20].
Exponential Smoothing uses a different way to weight
Ensemble methods have consistently surpassed alternative
past observations to make predictions. There are different
techniques in time series forecasting
types of smoothing, such as simple, double, and triple
smoothing, depending on whether the series has level, trend,
and seasonality[6][18]. The state-space perspective Simple Averaging for putting equal weight on many
illustrated exponential smoothing within a cohesive forecasts has been shown to lower prediction variance and
statistical framework. make predictions more reliable [28], [29]. Weighted
Prophet[8] recently came up with an additive decomposition Ensembles for systems that weight models based on how
that has piecewise trends, multiple seasonalities, and holiday well they work give more weight to models that are more
effects. It is made for business time series that have a lot of accurate [30]. It is still hard to figure out the best weight, and
seasonality and change regimes often. there are many ways to do it, such as inverse-error weighting
and regression-based algorithms [31]. Stacked generalization
[32] teaches meta-models how to use advanced aggregation
These classical methods are simple to understand and
functions to combine base predictions. Recent studies have
quick to calculate, but they only work well for linear or
employed neural networks as meta-learners [14].
|xvali | = 0.15 · Ti, and
Puzzles for Merging Predictions of empirical research has |xtesti | = 0.15 · Ti.
demonstrated scenarios where simple averaging outperforms
optimal weighted combinations [33], highlighting the The forecasting objective is to learn a function f : RL → RH
importance of robustness over theoretical optimality. that maps a lookback window of length L to a prediction
D. Gradient Boosting for Time Series horizon of length H, minimizing expected forecasting error:
Gradient boosting machines [34] have been very good at
using sequential ensemble learning to predict structured data.
f∗¿ argmi n {f ∈ F } E ( 2)
XGBoost [35] and LightGBM [36] made scaled versions that { (x , y ) D} [ L (f (x { t−L :t} ) , y {t +1:t +H } ) ]
[ ]
values and temporal encodings:
T
r t −1 ,
Here, WkQ,WkK,WkV ∈ Rd×dk are learnable projection
r t−7 ,
matrices, and K is the number of attention heads. Position- r t −14 ,
Wise Feed-Forward Network: f t= mean ( r t−1 : t−1 ) , ( 15 )
mean ( r t−7 : t−1 ) ,
FFN ( z )=max ( 0 , z W 1 +b 1 ) W 2+ b2 ( 9 )
dow ( t ) ,
Each sub-layer employs residual connections and layer month ( t )
normalization:
where dow(t) and month(t) represent day-of-week and
Z =LayerNorm ( Z + MultiHead ( Z ) ) ( 10 )
m m −1 m−1
month encodings, respectively.
3) XGBoost Training: We train an XGBoost regressor
[6] to predict residuals:
Z =LayerNorm ( Z + FFN ( Z ) ) (11 )
m m m
K
r^ t =Σ k=1 f k ( f t ) ( 16 )
4) Forecasting Head: The final transformer output is
flattened and passed through a linear projection to generate where fk represents the k-th regression tree in the
forecasts: ensemble, learned through gradient boosting:
LPatchTST = ( N·H1 ) Σ N
i=1 Σ Hh=1 ( y i ,h − ŷ i ,h )2 ( 13 )
1 T 2
Ω ( f )=γT + λ Σ j=1 w j ( 18 )
2
We use the Adam optimizer [22] with a learning rate of η =
0.001 and train for E = 30 epochs, stopping early if the
validation loss goes down. where T is the number of leaves in tree f, wj is the score
assigned to leaf j, and γ,λ are regularization hyperparameters.
C. Residual XGBoost Correction We use the following XGBoost configuration:
The main new thing about our framework is that it uses • Number of estimators: K = 100
gradient boosting to model PatchTST's prediction errors
• Maximum tree depth: dmax = 3
directly. This residual learning method is based on the fact
that deep learning models often make the same mistakes • Learning rate: ηXGB = 0.05
over and over again, even when the time period changes. • Subsample ratio: ρ = 0.8
1) Residual Calculation: We calculate residuals on the • Column subsample ratio: ρcol = 0.8
same training data after training PatchTST on x train.:
D. Adaptive Weighting Mechanism
PatchTST Our framework uses a validation-driven adaptive
r t =x t – x^ t ( 14 )
weighting scheme that dynamically balances the
contributions of PatchTST and the combined
where xˆPatchTSTt is the PatchTST prediction at time t during PatchTST+XGBoost prediction. This is different from
one-step-ahead forecasting. traditional ensemble methods that use fixed weights.
1) Component Performance Evaluation: We calculate
forecasting errors for two configurations on the validation
set xval:
Configuration 1 - PatchTST Only:
patch=RMSE ( y i , ŷi
ε (i) ) ( 19 )
val PatchTST
ensemble =RMSE ( y i , ŷ i
ε (i) ) ( 21 )
val ensemble
(i)
(i) w patch
α patch = ( 23 )
( w(i)patch +w(i)ensemble)
(i)
(i) w ensemble
α ensemble = ( 24 )
( w(i)patch + w(iensemble
)
)
4) Final Prediction: The final forecast combines both
components using adaptive weights: The computational complexity of our framework consists
of three main components:
PatchTST Training: The transformer encoder has
( 27 )
xt = Hyperparameters match the residual XGBoost component
( max ( x train )−min ( x train ) ) configuration.
D. Evaluation Metrics
Traditional statistical models (ARIMA, Prophet) and We assess forecasting performance using four
XGBoost operate on the original scale without complementary metrics:
normalization.
1) Mean Absolute Error (MAE):
B. Volatility Regime Classification
To analyze model robustness across different demand 1
patterns, we categorize each time series into volatility
(i)
MA E = Σ t ∈T | y(i)t − ŷ (i)t |( 32 )
|T |
test
test
regimes based on the coefficient of variation (CV):
2) Root Mean Squared Error (RMSE):
σ ( x train
i )
C V (i)= (28 )
μ ( xi )
√[
train
(i) 1
RMS E =
|T test|
Volatility Categories:
• Low Volatility: CV < 0.20 (stable, predictable demand)
3) Mean Absolute Percentage Error (MAPE):
• Medium Volatility: 0.20 ≤ CV ≤ 0.35 (moderate
variation)
| |
i i
(| | )
• High Volatility: CV > 0.35 (highly erratic, spiky i 100 yt − ŷt
demand) MAP E = test
Σ {t ∈T test
} i
( 34 )
C. Baseline Models
T yt + ε
We compare the proposed error-adaptive ensemble against where ϵ = 0.1 prevents division by
four established forecasting approaches: zero.
1) ARIMA: We implement seasonal ARIMA with
automatic order selection: 4) Coefficient of Determination (R2):
2
Φ P ( Bs ) φ p ( B ) ∇ sD ∇d x t =ΘQ ( B s ) θq ( B ) ε t ( 29 ) 2 (i )
Σ t ∈T test ( y it − ŷ it )
R =1− 2
( 35 )
( y it − ȳ i )
(AIC) with p,d,q ∈ {0,1,2} and seasonal order (P,D,Q,s)
Order selection uses the Akaike Information Criterion Σ t ∈T test
Table II shows how well different prediction horizons 2) Business Value: The 28.5% MAPE reduction translates
work for forecasting and how strong the proposed ensemble to substantial operational benefits:
approach is. The ensemble consistently gets lower MAPE Inventory Cost Savings: Lower forecast errors reduce
values than the other models at all time horizons. This shows
[37], a 28.5% error reduction enables ∼20% reduction in
safety stock requirements. Using the newsvendor model
that it has a clear and lasting advantage in both short- and
long-term forecasting situations. As the horizon lengthens, safety stock while maintaining service levels.
all methods make more mistakes when trying to predict the Quantitative Example: For a mid-sized retailer with
future. This is a common problem in time series prediction, $100M annual inventory carrying costs, a 20% safety stock
but the ensemble shows a more gradual decline in reduction yields $20M annual savings.
performance. This means that the suggested method is better
at keeping useful time-based information and dealing with C. Limitations and Future Work
uncertainty over longer forecasting windows. This makes it The results are very good, but there are some clear
especially useful for real-world supply chain planning, problems that need to be fixed. The model is a simple one
that only looks at past demand. It doesn't take into account
other things that could make it work better, like changes in management. By synergistically combining PatchTST for
price, sales, the weather, or the economy. Second, the model trend modeling with XGBoost-based residual learning for
has negative R² values even when the conditions are very systematic error correction, and employing validation-driven
unstable. This shows that even the best machine learning adaptive weighting, the proposed method achieves
algorithms can fail when the conditions are very unclear. substantial performance improvements.
Third, the model's design to train a separate model for each The suggested ensemble works well in terms of both
time series is not scalable from a computational point of accuracy and speed. It gets a MAPE of 67.76%, which is
view, which makes it hard to use with large datasets, like 28.5% better than the best baseline. This means it can make
tens of thousands of time series. Fourth, the model better predictions. The ensemble always does better than the
presupposes that the training and testing datasets are competition, even when demand patterns change. It works
stationary. In real-world supply chain networks, though, best when the market is not very volatile (68.72%),
changes in the regime and shocks in demand can make moderately volatile (100.45%), or very volatile (49.52%).
stationarity not hold. Finally, demand time series with a lot The proposed framework is computationally efficient, taking
of zeros (more than 90%) make modeling harder, which only 23.8 minutes to train and 4.1 ms per series to infer.
shows how important it is to have models that can handle Because of this, it is good for big companies. Ablation
sparse and intermittent data. studies also show how important each part is. For instance,
. PatchTST is a good place to start for predictive performance,
residual learning helps you learn from your mistakes, and
There are many promising ways to improve and build on
adaptive weighting helps you get the best results for certain
the proposed framework. To begin with, we can apply the
series.
method to hierarchical forecasting, which needs to be
From a business perspective, the 28.5% MAPE reduction
consistent across different levels of product and geographic
enables substantial operational benefits: reduced inventory
locations. This is an important step in making the supply
carrying costs, improved service levels, and enhanced
chain work better in real life. Second, adding causal
planning efficiency. For typical mid-sized retailers, these
modeling and inference would let the framework show how
improvements translate to millions of dollars in annual cost
promotions, pricing plans, and other changes affect
savings.
outcomes, making the forecasts easier to understand and use.
Thirdly, creating probabilistic versions of PatchTST would Future research directions include extending to
give us full predictive distributions instead of point multivariate forecasting, developing probabilistic variants,
forecasts. This would give us better information about leveraging metalearning, and integrating causal inference.
uncertainty for tasks that come after. Fourth, meta-learning This work establishes error-adaptive ensembling as a
on several time series would help learn good ways to start, promising paradigm for next-generation demand forecasting
which would make it easier to quickly adapt to new or cold- systems, demonstrating that carefully designed hybrid
start products with little data. Fifth, using neural architecture architectures can overcome the limitations of both classical
search and other AutoML methods would make it possible to statistical methods and singlemodel deep learning
automatically find the best model architectures and approaches.
hyperparameters. Finally, adding ideas from robust REFERENCES
optimization would let the framework model forecast
[1] T. M. Choi, S. W. Wallace, and Y. Wang, “Big data analytics in
uncertainty directly, which would make supply chain operations management,” Production and Operations Management,
optimization more robust and aware of risk. vol. 27, no. 10, pp. 1868–1883, 2020.
[2] D. Ivanov and A. Dolgui, “Viability of intertwined supply networks:
D. Comparison with State-of-the-Art Extending the supply chain resilience angles towards survivability,”
International Journal of Production Research, vol. 58, no. 10, pp.
Our work contributes uniquely by: 2904– 2915, 2020.
• Focusing on highly volatile supply chain data [3] A. A. Syntetos, Z. Babai, J. E. Boylan, S. Kolassa, and K.
• Demonstrating ensemble benefits for intermittent Nikolopoulos, “Supply chain forecasting: Theory, practice, their gap
demand • Providing production-ready implementation and the future,” European Journal of Operational Research, vol. 252,
no. 1, pp. 1–26, 2016.
Our MAPE values (67-100%) appear higher than typical [4] T. Januschowski, J. Gasthaus, Y. Wang, D. Salinas, V. Flunkert, M.
benchmarks (∼10-20%), but this reflects the inherent Bohlke-Schneider, and L. Callot, “Criteria for classifying forecasting
methods,” International Journal of Forecasting, vol. 36, no. 1, pp. 167–
difficulty of supply chain forecasting with extreme volatility 177, 2020.
rather than model deficiency. Relative improvements over [5] G. E. Box and G. M. Jenkins, Time series analysis: Forecasting and
baselines (28.5%) align with gains reported in other control. Holden-Day, 1970.
ensemble forecasting studies [27]. [6] C. C. Holt, “Forecasting seasonals and trends by exponentially
weighted moving averages,” International Journal of Forecasting, vol.
VII. CONCLUSION 20, no. 1, pp. 5–10, 2004.
[7] S. Makridakis, E. Spiliotis, and V. Assimakopoulos, “The M4
This paper presents a novel error-adaptive ensemble competition: Results, findings, conclusion and way forward,”
framework for demand forecasting in supply chain International Journal of Forecasting, vol. 34, no. 4, pp. 802–808, 2018.
[8] S. J. Taylor and B. Letham, “Forecasting at scale,” The American [29] A. Timmermann, “Forecast combinations,” in Handbook of Economic
Statistician, vol. 72, no. 1, pp. 37–45, 2018. Forecasting, vol. 1. Elsevier, 2006, pp. 135–196.
[9] S. Hochreiter and J. Schmidhuber, “Long short-term memory,” Neural [30] J. M. Bates and C. W. J. Granger, “The combination of forecasts,”
Computation, vol. 9, no. 8, pp. 1735–1780, 1997. Journal of the Operational Research Society, vol. 20, no. 4, pp. 451–
[10] H. Hewamalage, C. Bergmeir, and K. Bandara, “Recurrent neural 468, 1969.
networks for time series forecasting: Current status and future [31] P. Newbold and C. W. J. Granger, “Experience with forecasting
directions,” International Journal of Forecasting, vol. 37, no. 1, pp. univariate time series and the combination of forecasts,” Journal of the
388–427, 2021. Royal Statistical Society: Series A (General), vol. 137, no. 2, pp. 131–
[11] A. Vaswani, N. Shazeer, N. Parmar, J. Uszkoreit, L. Jones, A. N. 146, 1974.
Gomez, Ł. Kaiser, and I. Polosukhin, “Attention is all you need,” in [32] D. H. Wolpert, “Stacked generalization,” Neural Networks, vol. 5, no.
Advances in Neural Information Processing Systems, 2017, pp. 5998– 2, pp. 241–259, 1992.
6008. [33] J. H. Stock and M. W. Watson, “Combination forecasts of output
[12] Y. Nie, N. H. Nguyen, P. Sinthong, and J. Kalagnanam, “A time series growth in a seven-country data set,” Journal of Forecasting, vol. 23,
is worth 64 words: Long-term forecasting with transformers,” in no. 6, pp. 405–430, 2004.
International Conference on Learning Representations, 2023. [34] J. H. Friedman, “Greedy function approximation: A gradient boosting
[13] F. Petropoulos et al., “Forecasting: Theory and practice,” International machine,” Annals of Statistics, vol. 29, no. 5, pp. 1189–1232, 2001.
Journal of Forecasting, vol. 38, no. 3, pp. 705–871, 2022. [35] T. Chen and C. Guestrin, “XGBoost: A scalable tree boosting system,”
[14] S. Smyl, “A hybrid method of exponential smoothing and recurrent in Proceedings of the 22nd ACM SIGKDD International Conference
neural networks for time series forecasting,” International Journal of on Knowledge Discovery and Data Mining, 2016, pp. 785–794.
Forecasting, vol. 36, no. 1, pp. 75–85, 2020. [36] G. Ke, Q. Meng, T. Finley, T. Wang, W. Chen, W. Ma, Q. Ye, and T. Y.
[15] S. Kolassa, “Evaluating predictive count data distributions in retail Liu, “LightGBM: A highly efficient gradient boosting decision tree,”
sales forecasting,” International Journal of Forecasting, vol. 32, no. 3, in Advances in Neural Information Processing Systems, 2017, pp.
pp. 788–803, 2016. 3146–3154.
[16] P. Montero-Manso and R. J. Hyndman, “Principles and algorithms for [37] V. Cerqueira, L. Torgo, and I. Mozetic, “Evaluating time series
forecasting groups of time series: Locality and globality,” International forecast-ˇ ing models: An empirical study on performance estimation
Journal of Forecasting, vol. 37, no. 4, pp. 1632–1653, 2021. methods,” Machine Learning, vol. 109, no. 11, pp. 1997–2028, 2020.
[17] R. J. Hyndman and G. Athanasopoulos, Forecasting: Principles and [38] D. P. Kingma and J. Ba, “Adam: A method for stochastic
practice, 2nd ed. OTexts, 2018. optimization,” in International Conference on Learning
[18] Z. Ma and T. Pan, "Adaptive Weight Tuning of EWMA Controller via Representations, 2015.
Model-Free Deep Reinforcement Learning," IEEE Transactions on
Semiconductor Manufacturing, vol. 36, no. 1, pp. 1–9, Feb. 2023, doi:
10.1109/TSM.2023.3236798.
[19] L. Kumar, S. Khedlekar, and U. K. Khedlekar, "A comparative
assessment of Holt-Winters exponential smoothing and autoregressive
integrated moving average for inventory optimization in supply
chains," Supply Chain Analytics, vol. 4, art. no. 100084, Sep. 2024,
doi: 10.1016/[Link].2024.100084.
[20] S. Makridakis, E. Spiliotis, and V. Assimakopoulos, “The M4
competition: 100,000 time series and 61 forecasting methods,”
International Journal of Forecasting, vol. 36, no. 1, pp. 54–74, 2020.
[21] K. Cho, B. Van Merrienboer, C. Gulcehre, D. Bahdanau, F. Bougares,¨
H. Schwenk, and Y. Bengio, “Learning phrase representations using
RNN encoder-decoder for statistical machine translation,” arXiv
preprint arXiv:1406.1078, 2014.
[22] D. Salinas, V. Flunkert, J. Gasthaus, and T. Januschowski, “DeepAR:
Probabilistic forecasting with autoregressive recurrent networks,”
International Journal of Forecasting, vol. 36, no. 3, pp. 1181–1191,
2020.
[23] S. Bai, J. Z. Kolter, and V. Koltun, “An empirical evaluation of generic
convolutional and recurrent networks for sequence modeling,” arXiv
preprint arXiv:1803.01271, 2018.
[24] A. van den Oord, S. Dieleman, H. Zen, K. Simonyan, O. Vinyals, A.
Graves, N. Kalchbrenner, A. Senior, and K. Kavukcuoglu, “WaveNet:
A generative model for raw audio,” arXiv preprint arXiv:1609.03499,
2016.
[25] B. Lim, S. O. Arık, N. Loeff, and T. Pfister, “Temporal fusion
transform-¨ ers for interpretable multi-horizon time series forecasting,”
International Journal of Forecasting, vol. 37, no. 4, pp. 1748–1764,
2021.
[26] H. Zhou, S. Zhang, J. Peng, S. Zhang, J. Li, H. Xiong, and [Link],
Informer: Beyond efficient transformer for long sequence timeseries
forecasting,” in Proceedings of the AAAI Conference on Artificial
Intelligence, vol. 35, no. 12, 2021, pp. 11106–11115.
[27] T. G. Dietterich, “Ensemble methods in machine learning,” in
International Workshop on Multiple Classifier Systems. Springer,
2000, pp. 1–15.
[28] R. T. Clemen, “Combining forecasts: A review and annotated
bibliography,” International Journal of Forecasting, vol. 5, no. 4, pp.
559–583, 1989.