0% found this document useful (0 votes)
13 views18 pages

Next Energy: Praveen Kumar Singh, Amit Saraswat, Yogesh Gupta

This research article explores deep learning and machine learning models for short-term forecasting of solar photovoltaic (SPV) power generation, addressing the challenges posed by variability and uncertainty in solar energy. The study evaluates various models, including Stacked LSTM, Bi-LSTM, and hybrid CNN-LSTM, using the DKASC Alice Springs dataset and finds that the Stacked LSTM model outperforms others in prediction accuracy. The paper highlights the importance of accurate forecasting for grid stability and energy management, providing a comprehensive analysis of model performance across different data sequences.

Uploaded by

Pranav Patil
Copyright
© All Rights Reserved
We take content rights seriously. If you suspect this is your content, claim it here.
Available Formats
Download as PDF, TXT or read online on Scribd
0% found this document useful (0 votes)
13 views18 pages

Next Energy: Praveen Kumar Singh, Amit Saraswat, Yogesh Gupta

This research article explores deep learning and machine learning models for short-term forecasting of solar photovoltaic (SPV) power generation, addressing the challenges posed by variability and uncertainty in solar energy. The study evaluates various models, including Stacked LSTM, Bi-LSTM, and hybrid CNN-LSTM, using the DKASC Alice Springs dataset and finds that the Stacked LSTM model outperforms others in prediction accuracy. The paper highlights the importance of accurate forecasting for grid stability and energy management, providing a comprehensive analysis of model performance across different data sequences.

Uploaded by

Pranav Patil
Copyright
© All Rights Reserved
We take content rights seriously. If you suspect this is your content, claim it here.
Available Formats
Download as PDF, TXT or read online on Scribd

Next Energy 11 (2026) 100531

Contents lists available at ScienceDirect

Next Energy
journal homepage: [Link]/journal/next-energy

Research article

Deep learning prediction models for short-term solar photovoltaic power


generation forecasting ]]
]]]]]]
]]

Praveen Kumar Singha, Amit Saraswatb, , Yogesh Guptac


a
Department of Computer Applications, Manipal University Jaipur, Rajasthan 303007, India
b
Department of Electrical Engineering, Manipal University Jaipur, Rajasthan 303007, India
c
School of Engineering and Technology, BML Munjal University, Gurugram, Haryana 122413, India

A R T I C L E I N F O A B S T R A C T

Keywords: The increasing concerns about the environmental impact of fossil fuels have emphasized the importance of clean
Solar photovoltaic power generation solar energy, which offers a pollution-free alternative for meeting growing energy needs. However, the accurate
Renewable energy sources prediction of solar photovoltaic (SPV) based power generation is a very challenging task because of its inherent
Short-term forecasting variability and uncertainty. To address this challenging problem, this paper applies several machine-learning,
Deep-learning methods
deep-learning, and their hybrid models such as: One-Dimensional Convolutional Neural Network (1D CNN), Bi-
Machine-learning methods
Directional Long Short-Term Memory (Bi-LSTM), Stacked LSTM, Artificial Neural Network (ANN), Linear
Support vector regression
Hybrid CNN-LSTM models Regression (LR), Support Vector Regression (SVR), XGBoost, and a hybrid CNN-LSTM model. These models are
examined and compared on four different data sequences of DKASC Alice Springs dataset. The prediction per­
formances of all these models are evaluated based on various error metrics: MAE (mean absolute error), ex­
plained variance, RMSE (root mean square error), R², and sMAPE (symmetric mean absolute percentage error).
The simulation results demonstrates that Stacked LSTM model outperforms all other benchmark forecasting
models and able to obtains average values of performance metrics i.e. MAE of 1.1157, RMSE of 2.3408, an
Explained Variance of 0.8998, R² of 0.9004, and sMAPE of 1.1795 as evaluated across all four different data
sequences. Moreover, a comprehensive statistical analysis, using Diebold Mariano Test and boxplots, confirms
the further superiority of Stacked-LSTM model to efficiently address inherent uncertainty of solar power gen­
eration.

1. Introduction when it comes to accurate estimation of photovoltaic power. The recent


evolution of powerful deep learning methods provides obvious solu­
The world has increasing concerns about the environmental impact tions for these accurate prediction problems of photovoltaic power
of ongoing fossil fuel consumption including oil, natural gas, and coal. generation which heavily depends on various climatic conditions.
By harnessing clean and environmentally friendly solar energy through In the modern RE integrated power grid, precise solar power gen­
photovoltaic power generation has been brought into the spotlight over eration forecasting is of utmost importance for reliable and economical
the last one decade. This renewable energy source offers an efficient grid operations [3]. It helps the grid operators and energy planners to
and sustainable alternative to meet ever increasing energy demands. effectively manage a balance of supply and demand of electricity, en­
Over the past few decades, urbanization has significantly increased suring grid stability and efficient utilization of RE resources [4]. An
energy and power demands, substantially impacting the environment. accurate estimating solar power generation, with its high penetration to
In response, many countries have promoted renewable energy (RE) the modern grids, is a tricky task because of the unpredictable nature of
policies, with solar energy emerging as a highly favored source due to factors such as environmental and weather conditions. However, the
its abundant supply and pollution-free generation [1]. The last decade maximum power output from a SPV, which inherently depends upon
has seen a remarkable rise in the adoption and deployment of solar various environment and climate conditions, may be obtained by
photovoltaic (SPV) systems for power generation [2]. On contrary, the adopting an appropriate power tracking scheme [5]. The output of an
variability and intermittency of solar energy pose significant challenges SPV system is influenced not only by the intensity of sunshine but also


Corresponding author.
E-mail address: [Link]@[Link] (A. Saraswat).

[Link]
Received 1 December 2025; Received in revised form 27 January 2026; Accepted 29 January 2026
2949-821X/© 2026 The Authors. Published by Elsevier Ltd. This is an open access article under the CC BY-NC-ND license ([Link]
nc-nd/4.0/).
P.K. Singh, A. Saraswat and Y. Gupta Next Energy 11 (2026) 100531

by various other atmospheric variables that are closely tied to time. The significant research work as reported in [24] applied SVM and ANN
solar power generations are inherently uncertain and have high de­ based machine learning methods for short-term predictions of solar
pendency upon these weather conditions, which can significantly affect radiation. Nam et al. [25] introduced a gated recurrent unit model for
daily grid operations management [6]. Accurate short-term forecasting solar power forecasting, which demonstrated superior performance
of SPV power generation is a critical concern for the renewable energy over traditional machine learning models.
integrated electrical grid operations. Deep learning models, which can Several research studies have been reported to test two main cate­
capture complicated patterns in vast datasets, have considerable ad­ gories of deep learning (DL) methods for precise solar power forecasting
vantages over standard forecasting methods as discussed above. i.e. Long Short-Term Memory (LSTM) based neural networks [26–28]
Therefore, the research work provides a comprehensive investigation and Convolutional Neural Network (CNN) [29,30]. Mishra et al. [28]
on algorithmic capabilities of deep learning-based prediction models introduced a hybrid model that combines wavelet transformation (WT)
for accurate short-term forecasting of SPV power generations, with the for feature extraction from time series solar power data with LSTM to
goal of enhancing grid dependability and efficiency through improved incorporate additional meteorological variables. A hybrid multichannel
forecast accuracy. CNN model [29] was developed which extracts both meteorological and
geographical features from SPV power plants, resulting in enhanced
1.1. Literature review forecasting precision and accuracy. Wang et al. [30] proposed predic­
tion models based on an end-to-end mapping between solar irradiance
A recent literature reports numerous studies on various forecasting data and concurrent sky images, and the results suggested that CNN-
methods aimed at achieving accurate solar power predictions. Some of based deep learning-driven mapping models effectively capture their
the very recent literature reviews [7,8] suggest that these methods are relationships, surpassing the performance of conventional ANN-based
primarily classified based on time horizons into: (a) Long-term (one models.
year and above), (b) Medium-term (a week to a few months), and (c) Moreover, several hybrid deep learning models have been also
Short-term (few hours or less up to few minutes). Medium-term fore­ proposed which combine various architectures consistently outperform
casts are mainly used for scheduling and planning, while short-term other regression models [31]. Among the various deep learning
forecasts are needed for real-time grid operations like power dis­ methods, the LSTM networks have shown the highest accuracy, parti­
patching and generation control. The short-term forecasts are essential cularly for short-term, one-step-ahead predictions [32]. Further, Hamad
for affordable energy, day-ahead markets, and storage management [9]. et al. [33] proposed CNN-LSTM based hybrid model for one hour ahead
The SPV generation forecasting models are fundamentally grouped time series solar power forecasting. Dhakad et al. [34] implemented
into four categories: (i) Physical models, (ii) Statistical models, (iii) LSTM model to predict solar power generation and validated their re­
Machine-learning models, and (iv) Deep-learning models. The physical sults in terms of accuracy across various seasons using different archi­
models rely on mathematical principles such as photovoltaic conversion tectural configurations. Zang et al. [35] proposed another hybrid CNN
efficiency and weather conditions but face limitations due to incon­ model for a day-ahead SPV generation forecasting model using different
sistent assumptions and inaccurate information [10]. On the other input time series metrological parameters. In that model, the varia­
hand, the statistical methods try to establish a correlation mapping or tional decomposition mode was applied to historical solar power data
relationships between output-input data using curve fitting approaches sequences to decompose it into various frequency bands to obtain better
on available processed historical data through correlation analysis [11]. prediction accuracy for solar power generations. Ray et al. [36] in­
A crucial limitation of statistical methods is large computation time to troduced a hybrid model that combines LSTM and CNN for forecasting
process the input data and therefore, not suitable for real time appli­ long-term solar power using 24 years of dataset. Wang et al. [37] pro­
cations [12]. In contrast, machine-learning and deep-learning techni­ posed a hybrid model that integrates wavelet transformation with a
ques leverage non-linearities and advanced simulations, extracting deep neural network for solar power forecasting.
complex features to provide precise forecasting [13]. Consequently, Another study on a deep recurrent neural network employing LSTM
these methods have gained popularity for accurate photovoltaic power architecture was proposed with PSO for forecasting both demand and
generation forecasting [14]. However, apart from the specific solar-PV supply-based load and solar power output [38]. Mansour et al. [39]
generation forecasting application as focused in the present paper, there conducted a comparative study of various deep learning methods such
are several other applications of machine learning methods that have as 1D-CNN, Bi-LSTM, and Gated Recurrent Unit (GRU) for two distinct
been also reported in recent literature [15–19] such as: thermal coal datasets where it was found that Bi-LSTM yielded superior results
futures trading volume predictions [15], house price forecasting [16], compared to the other competing models. Zheng et al. [40] adopted a
wholesale prices of yellow corn [17], and steel price predictions strategy to optimize the hyperparameter of LSTM using Particle Swarm
[18,19] which demonstrate their applicability for a wide range of ap­ Optimization (PSO) algorithm and conducted a sensitivity analysis with
plications. different architectures of LSTM. Wang et al. [41] proposed a hybrid
Many researchers recently focused on developing deep learning and model of LSTM and RNN with the principle of time correlation to en­
machine learning based models for the precise solar power generation hance the prediction accuracy. Luo et al. [42] introduced a hybrid LSTM
forecasting [20]. These methods are very complex procedures which model incorporating domain expertise in physical constraints to de­
require extensive data from various smart metering devices [21]. These monstrate that the deep learning models benefit from more than just
methods incorporate several procedural steps, including data prepara­ data-driven approaches; integrating domain knowledge can enhance
tion, processing, algorithmic training, and validation. Some very recent accuracy and resilience. Gao et. al. [43] also proposed an LSTM-based
and relevant research works are tabulated in terms of developed model for different weather datasets for short-term SPV power fore­
methods, validation metrics and results as presented in Table 1. casting. Other significant research on appropriate LSTM model devel­
Few researchers applied Support Vector Machine (SVM) approach opment for day-ahead solar irradiance forecasting was reported in [44].
[22,23] for regression analysis, classification, pattern recognition, and Li et al. [45] developed a hybrid deep learning model that merges
prediction in SPV power generation forecasting. A hybrid model com­ wavelet packet decomposition with LSTM networks by utilizing his­
bining fuzzy regression and SVM was reported for accurate long-term torical data such as SPV power generation and weather data. Moreover,
predictions of horizontal solar radiation [22]. Further, Pan et al. [23] a significant hybrid model was also proposed by integrating CNN and
developed an ant colony-optimization method for very short-term SPV LSTM to extract key features based on the input data sequences re­
power forecasting by considering various weather parameters such as corded at the same time on different dates concerning some critical
meteorological temperature, wind direction, humidity, global and dif­ weather changes impacting SPV generation [46,47]. Sharadga et al.
fuse horizontal radiation, and sampling time in their work. Another [48] proposed a Bi-LSTM for predicting large-scale SPV output,

2
P.K. Singh, A. Saraswat and Y. Gupta Next Energy 11 (2026) 100531

surpassing the performance of both traditional neural networks and models i.e. ANN, LR, SVR, XGBoost for a short-term photovoltaic
statistical models. power generation forecasting applications.
✓ Comprehensive testing of these prediction models on a benchmark
1.2. Research gap and key contributions dataset of 1B DKASC Alice Springs SPV system.
✓ Validation of all implemented prediction models using four distinct
While numerous studies have adopted deep-learning and machine- performance indices: RMSE, R, MAE, and Explained Variance,
learning methods to predict solar photovoltaic power output, most of sMAPE.
them have been focused on analyzing a single dataset. However, there ✓ Exhaustive statistical analysis using Diebold Mariano Test and
is limited examination of how these models perform across different boxplots.
datasets and sequences. These research works indicated, as listed in
Table 1, that there is still much scope to further enhance prediction 1.3. Paper organization
accuracy of deep-learning, machine-learning, and their hybrid models.
To address this important research gap, the present research work The organization of the subsequent paper is presented as follows:
compares the performance of various deep-learning and machine- Section 2 exhibits the proposed methodology incorporating various
learning models to obtain accurate forecast for SPV power output across procedural steps. A detailed explanation of all the applied prediction
four distinct data sequences. The presented case study and its findings models is given in Section 3. The experimental results and associated
may further aid for selecting appropriate forecasting models and en­ discussion, along with a comparative study, are presented in Section 4.
hance the understanding of how varying input data sequences impact Finally, a conclusion of the present research work is drawn in Section 5.
SPV power predictions. The major contributions of the present research
work are outlined as follows: 2. Research methodology

✓ Development of four deep-learning prediction models i.e. LSTM, Bi- As the timely and accurate 30-minute-ahead solar photovoltaic
LSTM, 1D-CNN, CNN-LSTM and four distinct machine learning (SPV) power forecasts are of great essence to the grid operators which

Table 1
Critical literature review of some recent and relevant research works.

Ref. Year Model Benchmarking Methods Horizon Data Error Metrics


(Proposed / Tested) Resolution

[22] 2017 FRF – SVM FRF-SVM-Lin 15 min 1 hr RMSE, MAE, MaxAE,


FRF-SVM-Pol and IQR-AE
FRF-SVM-Gauss
FRF-SVM-Sig
ANFIS
[23] 2020 I–ACO–SVM SVM, ACO – SVM 60 min 5 min MSE, MAE, RMSE, and
R2
[24] 2022 MLR SVM, ANN 30 min 15 min RMSE, R2 , MSE, and
MAE
[25] 2020 EMD – GRU EMD - LSTM, EMD - DNN, EMD-SARIMA, EMD- 7 days 1 day MASE and MAE
MLR
[27] 2019 CNN – LSTM CNN, LSTM 30 min 5 min RMSE, MAE, and MAPE
[28] 2020 WT – LSTM LR, RR, LASSO, ENR 1, 15, 30, and 60 days 5 min RMSE, MAE, MAPE,
and R2
[29] 2021 CNN ANN, MLR 1 month 1 month RMSE, MAPE, and R2
[30] 2019 CNN ANN, LSTM 15 min 15 min RMSE, MAE, and CORR
[31] 2020 5D CNN – LSTM 5D-LSTM, 2D-CNN - LSTM 10 min to 180 min 10 min MSE, RMSE, and MAE
[32] 2021 LSTM Bi-LSTM, GRU, Bi-GRU, 1D CNN, 1D-CNN-LSTM, 1, 5, 30, and 60 min 1 min RMSE, MAPE, R, MAE
1D-CNN-GRU, NAR, ENN
[33] 2023 DSCLANet CNN, LSTM, CNN-LSTM, CNN - GRU 1 hr 1 hr MSE, MAE, and RMSE
[34] 2023 LSTM BPNN 15 min 15 min RMSPE, MAE, MAPE,
and R2
[35] 2020 CNN (ResNet, SVR, MLP, CNN, RFR, ETS, THETA, Physical 1 day 5 min MSE, MAE, MSLE,
DenseNet) MASE, NIA, UI
[36] 2020 LSTM – CNN ANN, Random Forest, NME, 1 yr 1 hr RMSE, NRMSE, MAPE,
and R
[37] 2017 WT – DCNN -QR BPNN, SVM, WT-DCNN 15 min, 45 min, 1 hr, 15 min MAPE, RMSE, and MAE
and 2 hrs
[38] 2019 DRNN–LSTM- PSO MLP, SVM 24 hr 1 hr RMSE, MAE, MAPE,
and PCC
[39] 2019 BI-LSTM 1D CNN, GRU 5 min 5 min RMSE, MSE, MAE and
R2
[40] 2020 LSTM – PSO LSTM, ANN, XGBoost 1 hr 30 min MAE, and RMSE
[41] 2020 LSTM – RNN BPNN, SVM 1 day 15 min RMSE, MAE, and COR
[42] 2021 PC – LSTM ARMA, KNN, FCNN, LSTM 1 hr 1 hr MAE, MSE, and R2
[43] 2019 LSTM BP, LSSVM, WNN 1 day 15 min RMSE, and MAD
[44] 2018 LSTM LR, BPNN 1 day 1 hr RMSE
[45] 2019 WPD + LSTM LSTM, RNN, GRU, MLP 1 hr 5 min MBE, MAPE, and RMSE
[46] 2020 CNN – LSTM BPNN, RBFNN 15 min – 90 min 15 min MAE, and RMSE
[47] 2019 LSTM – CNN LSTM, CNN, CNN-LSTM 15 min 5 min MAE, MAPE, RMSE,
and SDE
[48] 2020 BI-LSTM LSTM, LRNN, MLP, SARIMA, ARIMA, ARMA 1, 2, 3 hrs 15 min R2 , and RMSE

3
P.K. Singh, A. Saraswat and Y. Gupta Next Energy 11 (2026) 100531

Fig. 1. Generalized research methodology for solar power forecasting.

can make real-time scheduling decisions to at least dispatch power in and CNN-LSTM. The basic descriptions of all these prediction models
the short run and to mitigate solar curtailment [49]. The efficient short- and their structures are further detailed as presented in Section 3.
term forecasts enable the grid operator to predict the quick changes in
the solar-generated power, arrange the resources under control, and
2.5. Output data validation and visualization
ensure the balance between supply and demand. In other words, it
manifests effective proactive dispatch-level decisions while minimizing
The algorithmic performance of every predictive model is generally
photovoltaic curtailments in real-time grid operations. A generalized
estimated and validated in terms of the following five distinct indices:
methodology is designed for SPV power short-term forecasting on 30-
(i) Root Mean Square Error (RMSE), (ii) R-squared (R2), (iii) Mean
minute ahead of real time basis as illustrated in Fig. 1. It incorporates
Absolute Error (MAE), (iv) Explained Variance, and Symmetric Mean
five distinct procedural steps as described in further subsections.
Absolute Percentage Error (sMAPE) [39].
2.1. Data collection N
[Z (i) Z (i ) ]2
RMSE = 1
To carry out the desired experiments for short-term SPV power a=1
N (1)
predictions, appropriate input historical data is always needed. For a
practical prediction model, this input data is acquired from the plant (Z (i ) Z (i ) )2
geographical information details, meteorological data, weather pre­ R2 = 1
(Z (i ) Z (i )avg )2 (2)
dictions, and SPV battery storage system, etc. However, a comprehen­
sive performance analysis requires a time series dataset to validate the N
prediction models. In the present research work, the input dataset is ( a=1
| Z (i ) Z (i) |)
MAE =
acquired from the 1B DKASC website located in Alice Springs [50], as N (3)
detailed in further subsection 4.1.
Var (Z (i ) Z (i ) )
Explained Variance = 1
2.2. Data Preparation Var (y ) (4)

Subsequently, in the second stage, the complete dataset is parti­ 1


N
|Zi Zi |
tioned into four sub-sequences to facilitate experiments for twelve sMAPE =
N (|Zi| + | Zi )/2 (5)
months, nine months, six months, and three months, respectively. i=1

where Z(i) denotes a true target value, Z (i) denotes its forecasted value.
2.3. Data pre-processing var represents variance. N represents the total number of samples.
These performance parameters, as mathematically defined by Eq.
In this procedural step, various data preprocessing methods such as (1) – Eq. (5), are generally adopted to confirm the proposed prediction
cleaning of input data through missing value handling, outlier detection model and to compare its algorithmic performance with other com­
and removal (as described in further Subsection 4.2.1), feature selec­ peting prediction models. The stated short-term photovoltaic power
tion, and normalization are carried out. Subsequently, the entire input forecasting approach's competency and prediction accuracy are eval­
dataset is split into training and testing subsets for model development uated using these metrics. A metric R2, also known as the coefficient of
and evaluation as suggested in [51], as described in further subsection determination, is used to show how closely the outcome to match the
4.2.4. regression line. The explained variance score falls within a limit of 0–1,
and a higher value signifies a more effective model. RMSE calculates
2.4. Deep learning based prediction model the typical size of the discrepancies of forecasted values to actual va­
lues. A smaller RMSE signifies a model that fits the data better. MAE
The fourth stage after an appropriate partitioning of the pre-pro­ calculates the typical size of the discrepancies between the actual va­
cessed dataset into distinct training and testing subsets, these data lues with their predicted values, without considering whether the errors
subsets are reshaped to align with the specifications of the chosen are in the positive or negative direction (overestimation or under­
predictive models. Further, these varying input data sequences are estimation). Finally, sMAPE is a metric used to evaluate the forecasts
applied to various prediction models: Stacked LSTM, 1D-CNN, Bi-LSTM, precision in terms of mean square errors percentages.

4
P.K. Singh, A. Saraswat and Y. Gupta Next Energy 11 (2026) 100531

2.6. Diebold Mariano test 3.2. Convolutional neural network (1D-CNN)

The Diebold-Mariano (DM) test is a statistical test to ascertain It is a specific class of deep learning based neural network archi­
whether two sets of predictions exhibit statistically equivalent forecast tecture as shown in Fig. 2. It is specifically designed for processing grid-
accuracy [52]. It is employed specifically to compare the forecast ac­ like data, such as image data. The CNN [46] is used in two forms such
curacy between two competing forecasting models. Let it be assumed as one-dimension (1D) CNN and two-dimension (2D) CNN. The 1D CNN
that the differences di between the first set and second set of predictions is used for time sequence data, whereas 2D CNN is generally used for
w.r.t. the actual values are e1 and e2 , respectively. As in the case of the image processing. The basic structure of 1D CNN includes various
absolute error loss, the loss function is defined by Eq. (6) as follows: layers: convolutional layer, input and output layers, excitation layer,
pooling layer, and fully connected layer as shown in Fig. 2.
L (e ) = |e| (6)
Convolutional Layer: – It is a primary building block of CNN ar­
So, the loss difference di between the first set and second set of chitecture, where most of the computation takes place. It can be for­
predictions is defined by Eq. (7). mulated as Eq. (10).
di = |e1| |e2 | (7) Nc Oc + 2P
Vc = +1
S (10)
According to the null hypothesis, the (DM) statistic follows a stan­
dard normal distribution. A null hypothesis H0 states that two distinct where, Vc is output of convolutional layer, Nc is the size of input, kernel
prediction models, let's say A and B, have equal forecast accuracy or size is shown by Oc , P denotes the amount of padding, while S indicates
predictive ability. the stride value for the convolution kernel.
H0 : F (diA ) = F (diB ) (8) Pooling Layer: – Down sampling and dimensionality reduction are
two main responsibilities of pooling layer in convolutional neural net­
The alternative hypothesis H1 asserts that one model exhibits su­ work. Average pooling and maximum pooling are frequently adopted as
perior forecast accuracy compared to the other, i.e., two different kinds of pooling operations. It is mathematically ex­
H1 : F (diA) F (diB ) (9) pressed by Eq. (11).
Nc X+1
Wo = + No
S *(Nw X + 1) (11)
3. The SPV power prediction models
where, Wo is output, Nc is feature map height, NW is feature map
The developed deep-learning models for the short-term predictions width, No is feature map channels, S if stride and X is filtering size.
of SPV generations are described as follows: Flatten Layer: - It plays a critical role by converting the multi-di­
mensional output generated by the earlier convolutional and pooling
3.1. Persistence clear sky index (PCSI) model layers into a single-dimensional vector. This transformation is vital
because fully connected layers, which come after the flatten layer,
It is most simple and effective baseline model for evaluating solar operate on one-dimensional data, and this step facilitates the connec­
forecasting as suggested in [53]. This PCSI model is also considered as a tion between the features derived from the convolutional layers and
benchmarking model for evaluating all deep-learning models, which these fully connected layers.
combines the concept of persistence forecasting with the clear sky Fully Connected Layer: – It is to amalgamate the acquired features
index. The clear sky index at time t is defined by Eq. (17) as follows: and derive advanced predictions from these features. These layers are
complemented by diverse activation functions and can be employed in
C (t )
PCSI (t ) = combination with additional techniques like dropout to enhance reg­
Cclear (t ) (17)
ularization.
where C (t ) is measured solar power at time instance t and Cclear (t ) is
theoretical solar power under the clear sky index. 3.3. LSTM based neural network
The PCSI model assumes that the future value is equal to most recent
value. In the clear index model, it is defined as It is another type of deep learning model commonly used for tack­
ling complex time series forecasting problems. In comparison to
PCSI (t + h) = CSI (t ) (18)
Recurrent neural network (RNN), it offers solutions for gradient ex­
where PCSI (t + h) is the predicted PCSI for the forecast horizon h. plosion and gradient disappearance problems. The limitation of RNN is

Fig. 2. 1D-Convolutional Neural Network Architecture.

5
P.K. Singh, A. Saraswat and Y. Gupta Next Energy 11 (2026) 100531

Fig. 3. Basic LSTM Architecture.

to learn long term dependencies in case the data undergoes changes Moreover, the cell state update and hidden state update are per­
over time. The memory cells, input gate, output gate, and forget gate formed by Eq. (15) and Eq. (16), respectively.
are the fundamental components of an LSTM as presented as shown in
Ct = Ft × Ct 1 + It × Ct (15)
Fig. 3. The sigmoid activation function for forget gate, output gate,
input gate and tanh activation functions for cell state. X is for ele­ Ht = Ot × tanh(Ct ) (16)
mentwise multiplication and + is for summation. c (t 1) is prior state
of cell, h (t 1) is prior hidden layer output. However, x (t ) , c (t ) and, where, Ct is new cell state, Ct 1is previous cell state, Ht is new hidden
c (t ) represent the input data series, current cell state output, and hidden state is computed using cell state with output gate.
layer output, respectively. The LSTMs excel in sequence modeling due to their capability to
A multi-layer perceptron (MLP) structure and LSTM structure both attain and preserve extended interdependence within the dataset, all
share a lot of similarities. The LSTM network comprises of three distinct the while addressing the vanishing gradient issue frequently found in
layers: output layer, hidden layer, and input layer. The hidden layer traditional Recurrent neural network [41].
memory unit of a LSTM network is a critical component [23]. In a LSTM
network, memory cells act as storage units for information, which is 3.4. CNN – LSTM hybrid model
then controlled by various gates within the memory unit as detailed
below: Both LSTM and 1D CNN have their own advantages in the time
Forget Gate: - It is employed to eliminate unnecessary details from series prediction task. Due to CNN's spatial abilities and LSTM's tem­
the cell state. This gate manages the decision of whether to remember poral abilities, a hybrid model is also developed in this work, which is
or discard information from the prior cell state. The forget gate output known as CNN-LSTM [47]. CNN extract important features from input
is mathematically computed by Eq. (12). data and reshape to convert it into a time sequence format. Then after,
these sequence of feature maps are passed to an LSTM network. LSTMs
Ft = (Wpf . [Ht 1, Xt ] + Bf ) (12) are well-suited for capturing temporal dependencies in data. Further,
were, α sigmoid activation function, Wpf weight matrix, [Ht 1, Xt ] as­ this hybrid model is trained using the training dataset and predictions
sociation of current input with the previous state and Bf is bias in forget are made.
gate.
Input Gate: - It is used to update and retain pertinent information. It 3.5. Bi-directional LSTM model
further controls the new information flow into cell state. A mathema­
tical expression for the input gate is given by Eq. (13). It is a further extension of the conventional LSTM based neural
network structure as shown in Fig. 4. It is also a class of RNN. The
It = (Wpi [Ht 1, Xt ] + Bt ) (13)
primary feature of a Bi-LSTM is its simultaneous input sequences pro­
where, Bt is bias in input gate, α sigmoid function and Wpi is weight cessing capability in both forward direction as well as backward di­
matrix for input gate. rection. This capability enables it to cater for both past as well as future
Output Gate: - It plays a role in deciding the information to be information contexts, which is especially valuable in tasks requiring a
output or utilized by the network. This gate regulates the selection of comprehensive understanding of context from both directions [54]. In
which information should be transmitted as the ultimate hidden state Bi-LSTM input takes from both directions, Xt at time t, Xt + 1 at the next
from the cell state. The output gate operation is mathematically defined step and it is managed information from forward as well as backward
by Eq. (14). direction.

Ot = (Wot[Ht 1, Xt ] + Bo ) (14)
3.6. The proposed Stacked LSTM model
where, Ht 1 takes the input from previous cell state, Xt is current input,
Wot is a matrix of weights for the output gate and Bo is bias for output The proposed stacked LSTM architecture consists of multiple layers
gate. of LSTM units, where each layer progressively learns more abstract and

6
P.K. Singh, A. Saraswat and Y. Gupta Next Energy 11 (2026) 100531

Fig. 4. Bi-Directional LSTM Structure.

intricate temporal patterns from the input data. This layered structure ht = ot (*)tanh(Ct ) (20)
enables the model to capture hierarchical temporal dependencies and
where Ct is the cell state at time step t, ft is the forget gate (output
refines the features learned by the previous one at each layer. The
between 0 and 1), it is the input gate (output between 0 and 1), Ct is the
proposed approach improves the model’s ability to exploit long-term
candidate cell state and Ct 1 is the previous cell state.
relationships and complex temporal dynamics in sequential data [55].
Once the final layer of stack LSTM has processed the sequence, the
The proposed stacked LSTM model for SPV power forecasting is pre­
output is passed to a dense output layer, which reshapes the informa­
sented, incorporating three distinct LSTM layers in sequence to handle
tion into the desired format for forecasting. During training, over 50
the time-series data as depicted in Fig. 5.
epochs with a batch size of 200, the model's predictions are evaluated
The developed Stacked LSTM based deep-learning model in­
against the true values. The resulting errors are then used to update the
corporates multiple LSTM units such as input layer, dropout regular­
model weights through backpropagation algorithm, allowing the model
ization layers, and a fully connected dense output layer. The input layer
to learn and improve its accuracy. The Adam optimizer is adopted to
is the first component of the model where data is fed into the network.
efficiently update the weights, for minimizing the loss function. It is a
In this work, the input consists of 48 past time steps, with each time
sophisticated gradient descent technique used in this model to dyna­
step containing 6 features. These features denote past values of various
mically adjust the learning rate by leveraging the first and second
parameters like weather conditions or solar power data. This data is
moments of the gradient estimates. The update rule for a parameter t
passed to the first LSTM layer, where it starts learning the temporal
at time t is given by Eq. (21) as follows:
relationships between the observations. The core of the model com­
prises stacked LSTM layers, which are essential for capturing the tem­ mt
t = t 1 *
poral dependencies within the sequential data. The first layer of stack vt + (21)
LSTM has 64 units, and it processes the input sequence and learns the
is the learning rate, mt bias connected to the first moment mean, vt
lower-level temporal features such as short-term trends and basic pat­
Second-moment variance estimates and is a small constant to prevent
terns and captures local temporal dependencies in the input data. The
division by zero. In this case, Mean Squared Error (MSE) is chosen as
argument `return sequence true` ensures that the hidden state at each
the loss function, as it is suitable for regression problems. The Adam
time step is passed on to the next layer of stack LSTM for further pro­
optimizer’s adaptive learning rate helps the model converge efficiently,
cessing.
making it especially effective for training complex architectures like
Further, the second and third layers of stack LSTM have 32 units
stacked LSTMs.
each and these are built upon the initial features learned by the first
stack layer. These layers capture more abstract and higher-level tem­
poral patterns by processing the output sequence from the preceding 4. Simulation results and discussion
layer. The second layer also describes intermediate temporal de­
pendencies and complex and global temporal dependencies across a A detailed case study is presented by conducting several experi­
longer time horizon are captured by third layer. Then after, the third ments for testing four machine-learning models (i.e. ANN, SVR, LR, and
stack layer passes the outputs to the final hidden state, which is then XGBoost) and four deep-learning models (1D-CNN, CNN-LSTM, Bi-
passed through a dropout layer before being fed into a fully connected LSTM, and Stacked LSTM), their simulation outcomes, statistical ana­
dense output layer for the final prediction. Dropout layers are placed lysis and associated discussions in the following subsections.
between the stacked LSTM to reduce the risk of overfitting. Dropout is a
regularization technique that randomly disables a portion of the neu­ 4.1. Description of input DKSC DATASET
rons during training, encouraging the model to generalize better by not
relying too heavily on any specific neurons. In this work, a dropout rate For the presented case study, the input dataset of a historical solar
of 0.01 is applied after each LSTM layer to ensure the model does not power data and other associated meteorological data is collected from
overfit and can perform well on unseen data. The cell state and the Desert Knowledge Australia Solar Center (DKASC) situated in Alice
hidden state for all three layers are updated by Eq. (19) and Eq. (20), Springs, Australia as downloaded from [50]. This dataset comprises of
respectively. the observations spanning over one year, specifically from April 2016 to
March 2017, with data recorded at 5-minute intervals. The dataset
Ct = ft (*) Ct 1 + it (*) Ct (19) primarily includes twelve distinct parameters such as temperature,
current phase average, active power, wind speed, relative humidity,

7
P.K. Singh, A. Saraswat and Y. Gupta Next Energy 11 (2026) 100531

Fig. 5. Architecture of the proposed Stacked LSTM model.

global horizontal radiation, diffuse horizontal radiation, wind direction, 0°–360° interval. Such anomalies (as indicated in Table 2) are attribu­
weather daily rainfall, radiation global tilted, radiation diffuse tilted, table to sensor faults, data acquisition errors, missing-value place­
and active energy delivered as listed in Table 2. All the twelve said holders, or transmission interruptions. Nevertheless, the measures of
parameters make significant information on the target variable. The central tendency and dispersion remain within plausible physical
statistical description of the raw input dataset in terms of Max (max­ bounds for most variables. For instance, wind speed records a median of
imum), Min (minimum), mean, median, Std (standard deviation), and 1.70 m/s with an interquartile range spanning 0.94–3.14 m/s, while
percentiles 25%, 75% is tabulated in Table 2. Moreover, it also com­ wind direction exhibits a median of 214.5°, indicating that the bulk of
prises high-resolution meteorological observations and photovoltaic the data lies within expected operational limits. Similarly, negative or
system measurements recorded at the Desert Knowledge Australia Solar near-zero values observed in active power output and solar irradiance
Centre. These detailed statistics also reveals few anomalies in the raw variables predominantly reflect nighttime periods, inverter-related
input dataset i.e. certain variables display extreme minimum or max­ noise, or minor sensor calibration offsets.
imum values that are not physically realistic, including negative wind In all these research simulations, input data of one-year duration is
speed measurements and wind direction values exceeding the standard further bifurcated into four distinct data sequences which are named as

8
P.K. Singh, A. Saraswat and Y. Gupta Next Energy 11 (2026) 100531

Table 2
Statistical description of input dataset.

Dataset Features Count Min Max Median Mean Std 25% 75% Missing value (%)

Active Power Output 105408 -0.054 23.41 0 5.87 7.64 0 14.00 -


Temperature 105408 -39.98 41.70 21.90 21.03 9.84 14.77 28.15 -
Wind Speed 58622 -459.67 54.38 1.70 2.08 3.10 0.94 3.14 44.38%
Daily Rainfall 105408 0 65.99 0 0.57 3.96 0 0 -
Global Radiation 105408 0 1507.06 4.42 242.26 351.82 2.10 447.03 -
Weather Humidity 105408 0 102.89 33.87 40.26 26.35 19.67 58.43 -
Wind Direction 105408 -9836.6 30152.33 214.50 210.44 280.00 122.89 281.24 -
Diffuse Radiation 105408 0 707.58 3.28 58.02 100.83 1.13 71.10 -
Radiation Diffuse Tilted 104634 0.03 703.41 6.77 62.51 102.84 1.55 77.26 0.73%
Radiation Global Tilted 104634 0.06 1433.83 8.44 269.39 372.79 2.87 513.82 0.73%
Energy Delivered 105408 375894 427524 399570 400766.75 14912.52 387381.5 413701 -
Phase Average 105408 0 97.07 1.39 25.25 30.96 1.39 58.12 -

Data Sequence-A: 3-Months (April 2016 – June 2016), Data Sequence-B: 4.2.1. Missing values and outliers treatment
6-Months (April 2016 – September 2016), Data Sequence-C: 9-months As evident from Table 2, the highest missing values are exhibited for
(April 2016 – December 2016), and Data Sequence-D: 12-Months (April wind speed of 44.38%, and other features i.e. Radiation Diffuse Tilted,
2016 – March 2017) for experiments. A graphical visualization for this Radiation Global Tilted, have less than 1% missing values. The outliers
historical time-series data for SPV power is presented in Fig. 6. Here, detection interquartile range (IQR) method is applied for identifying
each subplot represents four subsequences spanning a particular one these missing values as suggested in [56], with observations below Q1
week in each of the four data sequences. In all subplots, the solar power − 1.5 × IQR or above Q3 + 1.5 × IQR considered extreme values.
output is observed to peak around noontime, decrease at sunrise and in These identified outliers are temporarily marked as missing (NaN) to
the mid-afternoon, and approach zero at nightfall. enable consistent treatment. Moreover, the negative values of active
power output and wind speed, as well as wind direction measurements
4.2. Data preparation for input DKSC DATASET outside the valid 0–360° range, are also classified as nonphysical re­
cords and removed, whereas physically meaningful zeros, such as
Several procedural steps such as missing value handling, outlier nighttime power output, were retained. Subsequently, all numeric
detection and removal, dataset split for training and testing purposes, features values are standardized using Z-score normalization, and
feature selection, and normalization are applied as detailed in further missing values, along with marked outliers, imputed using a K-Nearest
subsections. Neighbors (KNN)–based approach using five nearest neighbors with

Fig. 6. A graphical presentation of actual SPV power data for a particular time frame: (a) Data Sequence-A of 3-Months, (b) Data Sequence-B of 6-Months, (c) Data
Sequence-C of 9-Months, and (d) Data Sequence-D of 12-Months.

9
P.K. Singh, A. Saraswat and Y. Gupta Next Energy 11 (2026) 100531

Fig. 7. Heatmap Correlation among the features.

distance-based weighting in the normalized feature space [56]. More­


over, a daylight-only subset spanning 06:00 AM to 07:00 PM is utilized
to assess effective model performance during these periods of active
power generation. Subsequently, feature selection and normalization
are performed, followed by splitting the dataset into training and
testing sets.

4.2.2. Feature selection


It is very important to observe the dependency of the target variable
on all various system parameters and their inter-dependence or corre­
Fig. 8. Splitting of input dataset for training and testing.
lations among all these parameters. Therefore, a heatmap is presented
to highlight their associations among all these features in the dataset as
depicted in Fig. 7. It clearly indicates that the first five climate para­ validation, as shown in Fig. 8. On the other hand, the testing data is
meters, such as wind speed, global horizontal radiation, diffuse hor­ distinct from the training data, has the same probability distribution as
izontal radiation, weather temperature, and relative humidity, have the training data. Here, it is worth noted that the input dataset is not
higher correlations and may have a strong influence on the SPV power randomly divided for training and testing for these experiments as it is
generation forecasting outcomes. Therefore, these five parameters are having a prominent time series dependency.
selected for further simulations of SPV power generation predictions.
However, the other parameters may also affect the SPV power fore­ 5. Experimental results
casting to some very slight extent and therefore, are not considered
here. 5.1. Comparative performance of error metrics

4.2.3. Data normalization In this research work, four deep-learning models, i.e., 1D-CNN, Bi-
The input dataset, which has huge variations in the ranges for many LSTM, CNN-LSTM, Stacked LSTM, and four contemporary machine-
features and therefore, is also normalized within 0–1 from the original learning models, i.e., SVR, LR, XGBoost [57], and ANN [58] have been
range using min-max approach, as mathematically formulated as Eq. (22). implemented to forecast the SPV power outputs. For a fair statistical
comparison, each of these competing forecasting models are executed
DAct Dmin
DNor = for total 10 independent runs on all four data sequences as normalized
Dmax Dmin (22)
with the same target scale as described in Subsection 4.2.3. Moreover, a
where DNor is a normalized value, DAct is an actual value, Dmin and Dmax walk-forward validation method with a fixed-length sliding window is
are the minimum value and maximum value in all feature sets, respec­ adopted for all these competing forecasting models. In every step, past
tively 72 observations are utilized to train these models for the predictions in
subsequent 6-time steps, or 30-minute horizon. It is then followed by a
4.2.4. Splitting DATASET for training and testing purposes methodical advancement of the training window by a single observa­
Finally, the complete input time sequence dataset is divided into tion. All the developed deep learning models are trained using 50
two portions of 80% and 20% for training and testing, respectively. The epochs. The purpose of using epochs in deep-learning models is to en­
first portion is treated as training dataset, and the second portion is hance their algorithmic convergence and computation accuracy during
considered as testing dataset. The training data consisting of discrete their training process. A comprehensive overview of the hyperpara­
instances that are considered to train any prediction model by learning meters used in deep learning models and machine learning models,
its desired patterns. The input training dataset is further divided into along with detailed descriptions for each parameter, are detailed in
two parts, with a division of 90% for training the models and 10% for Table 3. In this study, four different input sequences are used, namely

10
P.K. Singh, A. Saraswat and Y. Gupta Next Energy 11 (2026) 100531

Table 3
Hyperparameters for four deep learning models and parameters associated with machine learning models.

Model Name Model Parameters, Optimizer & Loss function

ANN Units = 64, 32; activation = relu; Loss = mse; optimizer = adam
LR fit_intercept = True; n_jobs = None; positive = False; Intercept = 0.01823687
XGBoost Objective = reg:linear; booste = gbtree; learning_rate =0.01; n_estimators = 1000;
base_score = 0.5; max_depth = 3
SVR kernel = rbf; gamma = 0.1; C = 1.0; epsilon = 0.1
1D-CNN kernel_size = 5; filters = 64; pool_size=2; Batch Size= 200; Activation=relu; Optimizer = adam; Loss = mse
CNN-LSTM Filters=64, 32; kernel_size=5,3; pool_size = 2; strides =1; LSTM units=128,64,32,16; Batch Size: 200; Dropout=. 01; Activation=relu; Optimizer
=adam; Loss = mse
Bi-LSTM Units = 50, 32; Batch Size = 200; Activation = relu; Optimizer = adam; Loss = mse
Stacked LSTM Units = 64, 32,32; Dropout =.01; Batch Size=200; Activation = relu;
Optimizer = adam; Loss = mse

Table 4
Further details of hyper parameters for the proposed prediction model.

Input Sequence Model Parameters Remarks

(a) Data Sequence-A of 3-Months, LSTM Units (64, 32, 32) A stacked structure (decreasing number of units per layer) can effectively capture both high-level
(b) Data Sequence-B of 6-Months, and finer temporal patterns.
(c) Data Sequence-C of 9-Months, Dropout (0.01) Prevent overfitting
(d) Data Sequence-D of 12-Months Batch Size (200) Stabilizes gradient updates and speeds up training on GPUs.
Activation (ReLU) Avoids the vanishing gradient problem and accelerates training by providing non-saturating
gradients
Optimizer (Adam) Adaptive learning makes it robust and efficient
Loss (MSE) To minimize the squared difference between predicted and actual values

(a) Data Sequence-A of 3-Months, (b) Data Sequence-B of 6-Months, (c) patterns within sequential time-series datasets. The architecture of the
Data Sequence-C of 9-Months, and (d) Data Sequence-D of 12-Months, 1D-CNN utilized in this case study comprises of multiple layers in­
as discussed earlier. In the developed stacked LSTM network, three cluding a maximum pooling layer, a flatten layer, an input convolu­
distinct LSTM layers are employed alongside three dropout layers. All tional layer, and a dense output layer. A Bi-LSTM neural network aims
hyperparameters are tuned using the Adam optimizer, known for its to effectively capture and utilize both past and future context within
adaptive optimization algorithm. The effectiveness of each epoch is sequential data. This bidirectional approach offers several advantages,
assessed using the MSE loss function, as it offers a quantitative gauge of including enhanced contextual understanding, improved memory and
model accuracy. This metric is crucial for gradient-based optimization information flow, better representation learning, and enhanced per­
algorithms employed in training deep learning models. formance on prediction tasks. The Bi-LSTM neural network utilized in
A comprehensive overview of the hyperparameters used in deep this study incorporates a bidirectional input layer, a bidirectional
learning models and machine learning models along with detailed de­ output layer, and three dense layers. Two convolutional 1D layers, one
scriptions for each parameter are detailed in Table 3. In the developed max pooling layer, two LSTMs, one flattens, two dense, and two
stacked LSTM network, three distinct LSTM layers are employed dropout layers are employed in the hybrid CNN-LSTM network. The
alongside three dropout layers. All hyperparameters are tuned using case study also includes some of the contemporary machine learning
Adam optimizer as in Table 4. The effectiveness of every epoch is as­ methods such as ANN, XGBoost, LR, and SVR as given in Table 3.
sessed based on MSE (mean squared error) loss function, as it offers a The numerical comparisons of the algorithmic performances in
quantitative gauge of model accuracy. This metric is crucial for gra­ terms of RMSE, Explained Variance, R2 , MAE, and sMAPE matrices are
dient-based optimization algorithms employed in training deep presented corresponding to all developed prediction models for various
learning models. A 1D-CNN is a specific class of neural network com­ input sequences in Table 5, Table 6, Table 7, Table 8, and Table 9,
monly adopted to process one-dimensional time series data. While respectively. It is observed that the model's accuracy is the lowest when
traditional CNNs are primarily used for image recognition tasks, 1D- the length of input time-series dataset is set to 3 months. The overall
CNNs are specifically designed to capture interdependencies and accuracy of a prediction model is typically evaluated using RMSE,

Table 5
Comparison of developed models in terms of RMSE.

Prediction Model Data Sequence-A Data Sequence-B Data Sequence-C Data Sequence-D AVG
(03-MONTHS) Mean ± Std. (06-MONTHS) Mean ± Std. (09-MONTHS) Mean ± Std. (12-MONTHS) Mean ± Std.

PCSI 5.6322 ± 0.00 3.8576 ± 0.00 3.5622 ± 0.00 3.4623 ± 0.00 4.1286
ANN 3.6691 ± 0.02 2.2883 ± 0.01 2.3683 ± 0.02 2.3636 ± 0.03 2.6723
LR 5.7502 ± 0.01 2.6533 ± 0.02 3.2069 ± 0.01 2.4966 ± 0.02 3.5267
XG Boost 5.3273 ± 0.02 1.8386 ± 0.01 2.0085 ± 0.02 1.5885 ± 0.03 2.6907
SVR 2.5612 ± 0.04 3.0112 ± 0.03 2.6431 ± 0.03 2.6044 ± 0.02 2.7975
1D CNN 5.6384 ± 0.03 2.3112 ± 0.02 2.5744 ± 0.02 2.5288 ± 0.01 3.2632
CNN - LSTM 4.7510 ± 0.02 2.4040 ± 0.01 2.6247 ± 0.02 2.7880 ± 0.02 3.1419
BI LSTM 4.3130 ± 0.04 2.0977 ± 0.03 2.6225 ± 0.02 2.6325 ± 0.01 2.9164
Stacked LSTM 2.5350 ± 0.01 2.1937 ± 0.02 2.3569 ± 0.01 2.2779 ± 0.02 2.3408

11
P.K. Singh, A. Saraswat and Y. Gupta Next Energy 11 (2026) 100531

Table 6
Comparison of developed models in terms of Explained Variance.

Prediction Model Data Sequence-A Data Sequence-B Data Sequence-C Data Sequence-D AVG
(03-MONTHS) Mean ± Std. (06-MONTHS) Mean ± Std. (09-MONTHS) Mean ± Std. (12-MONTHS) Mean ± Std.

PCSI 0.4293 ± 0.00 0.8691 ± 0.00 0.8273 ± 0.00 0.8352 ± 0.00 0.7402
ANN 0.7536 ± 0.02 0.913 ± 0.03 0.911 ± 0.05 0.9007 ± 0.03 0.8695
LR 0.4182 ± 0.02 0.8887 ± 0.02 0.8634 ± 0.03 0.8912 ± 0.04 0.7653
XG Boost 0.4125 ± 0.03 0.9444 ± 0.04 0.943 ± 0.02 0.9551 ± 0.04 0.8137
SVR 0.8695 ± 0.06 0.8806 ± 0.04 0.8753 ± 0.03 0.8769 ± 0.07 0.8755
1D-CNN 0.6464 ± 0.04 0.909 ± 0.03 0.8928 ± 0.02 0.8865 ± 0.05 0.8336
CNN - LSTM 0.5505 ± 0.03 0.9024 ± 0.04 0.8884 ± 0.03 0.8726 ± 0.04 0.8034
BI-LSTM 0.6182 ± 0.04 0.9274 ± 0.04 0.8914 ± 0.04 0.8779 ± 0.03 0.8287
Stacked LSTM 0.8664 ± 0.02 0.9183 ± 0.03 0.9056 ± 0.02 0.9091 ± 0.03 0.8998

Table 7
Comparison of developed models in terms of R2.

Prediction Model Data Sequence-A Data Sequence-B Data Sequence-C Data Sequence-D AVG
(03-MONTHS) (06-MONTHS) (09-MONTHS) (12-MONTHS)
Mean ± Std. Mean ± Std. Mean ± Std. Mean ± Std.

PCSI 0.4291 ± 0.00 0.8698 ± 0.00 0.8270 ± 0.00 0.8350 ± 0.00 0.7402
ANN 0.7197 ± 0.06 0.9141 ± 0.04 0.9097 ± 0.04 0.8913 ± 0.06 0.8587
LR 0.309 ± 0.04 0.8802 ± 0.03 0.8340 ± 0.05 0.8890 ± 0.06 0.7280
XG Boost 0.4069 ± 0.04 0.9424 ± 0.04 0.9349 ± 0.06 0.9551 ± 0.04 0.8098
SVR 0.8484 ± 0.05 0.8709 ± 0.04 0.8698 ± 0.04 0.8695 ± 0.03 0.8646
1D-CNN 0.6299 ± 0.04 0.9092 ± 0.03 0.8926 ± 0.03 0.8864 ± 0.04 0.8295
CNN - LSTM 0.5300 ± 0.03 0.9017 ± 0.04 0.8884 ± 0.04 0.8720 ± 0.04 0.7980
BI- LSTM 0.6126 ± 0.04 0.9264 ± 0.04 0.8914 ± 0.03 0.8769 ± 0.05 0.8268
Stacked LSTM 0.8662 ± 0.03 0.9181 ± 0.02 0.9100 ± 0.03 0.9076 ± 0.03 0.9004

Table 8
Comparison of developed models in terms of MAE.

Prediction Model Data Sequence-A Data Sequence-B Data Sequence-C Data Sequence-D AVG
(03-MONTHS) (06-MONTHS) (09-MONTHS) (12-MONTHS)
Mean ± Std. Mean ± Std. Mean ± Std. Mean ± Std.

PCSI 3.4230 ± 0.00 1.0871 ± 0.00 2.5780 ± 0.00 1.6912 ± 0.00 2.1948
ANN 3.8124 ± 0.04 1.2398 ± 0.03 1.232 ± 0.05 1.3144 ± 0.04 1.8996
LR 2.9091 ± 0.03 1.8719 ± 0.03 2.258 ± 0.06 1.6947 ± 0.06 2.1834
XG Boost 3.6347 ± 0.03 1.0114 ± 0.04 1.1114 ± 0.03 0.8999 ± 0.04 1.6643
SVR 0.0936 ± 0.04 0.0904 ± 0.05 0.0875 ± 0.07 0.0844 ± 0.05 0.0889
1D-CNN 2.7121 ± 0.03 1.0624 ± 0.06 1.5652 ± 0.04 1.2905 ± 0.04 1.6575
CNN - LSTM 2.5287 ± 0.06 1.0742 ± 0.05 1.2715 ± 0.05 1.3006 ± 0.03 1.5437
Bi-LSTM 4.1412 ± 0.05 1.0521 ± 0.06 1.2411 ± 0.05 1.2932 ± 0.04 1.9319
Stacked LSTM 1.2073 ± 0.03 0.9778 ± 0.04 1.1525 ± 0.03 1.1253 ± 0.02 1.1157

Table 9
Comparison of developed models in terms of sMAPE.

Prediction Model Data Sequence-A Data Sequence-B Data Sequence-C Data Sequence-D AVG
(03-MONTHS) (06-MONTHS) (09-MONTHS) (12-MONTHS)
Mean ± Std. Mean ± Std. Mean ± Std. Mean ± Std.

PCSI 1.5106 ± 0.00 1.2461 ± 0.00 1.4476 ± 0.00 1.4976 ± 0.00 1.4254
ANN 1.3867 ± 0.06 1.2234 ± 0.05 1.1445 ± 0.05 1.1369 ± 0.04 1.2228
LR 1.5676 ± 0.07 1.2241 ± 0.04 1.1381 ± 0.03 1.1255 ± 0.06 1.2638
XG Boost 1.3504 ± 0.05 1.2207 ± 0.05 1.1268 ± 0.07 1.1050 ± 0.06 1.2012
SVR 1.3763 ± 0.05 1.2426 ± 0.06 1.1937 ± 0.06 1.1236 ± 0.05 1.2215
1D-CNN 1.3721 ± 0.06 1.2123 ± 0.05 1.1321 ± 0.05 1.1455 ± 0.06 1.2155
CNN - LSTM 1.4459 ± 0.03 1.2273 ± 0.04 1.1344 ± 0.04 1.1377 ± 0.04 1.2363
Bi-LSTM 1.4877 ± 0.04 1.2422 ± 0.03 1.1189 ± 0.04 1.1355 ± 0.06 1.2460
Stacked LSTM 1.2227 ± 0.03 1.2098 ± 0.04 1.1180 ± 0.03 1.0964 ± 0.03 1.1795

where a lower value indicates better predictive performance, and zero lowest average value for all four sequences being 2.3408. It suggests
signifies a perfect matching between actual values and their predicted that the stacked LSTM can fit the dataset better than the other models.
values. According to Table 5, the stacked LSTM model shows the lowest The proportion of target variance described by a prediction model is
RMSE value across all four sequences among the models tested, with the generally measured by Explained Variance, which quantifies how well

12
P.K. Singh, A. Saraswat and Y. Gupta Next Energy 11 (2026) 100531

the model accounts for the variability in the data. According to Table 6, very clear that the stacked LSTM model consistently exhibits a better fit
the explained variance for all four input sequences in the Stacked LSTM to the actual data as compared to other competing models with en­
model surpasses that of other models in the study. Explained variance hanced prediction accuracy to capture the SPV power patterns.
values close to 1 indicate that the model's predictions align closely with Furthermore, the obtained training and validation errors by all the
actual values. The average explained variance across all data sequences competing models for various input sequences are also graphically
for the Stacked LSTM is 0.8998, indicating satisfactory performance compared in Fig. 10 and Fig. 11, respectively. These graphical com­
comparison to the rest seven competing prediction models as presented parisons somehow help us to judge the case of overfitting and under­
in Table 6. The determination coefficient, R2 quantifies the percentage fitting. These graphical representations depict the significant fluctua­
of variance in the target-dependent variable that can be predicted from tions and instability in the discrepancies exhibited by each of these
the important features in a dataset. It serves as an indicator of the ac­ models on both the validation set and training set when the input se­
curacy of a prediction model in relation to the variability in the input quence is set to 3 months. However, when the input sequence is ex­
data. According to Table 7, the Stacked LSTM model achieves an tended to 6 months, the errors become more stable. Furthermore, the
average R2 value of 0.9004, surpassing the performance of the other proposed stacked LSTM model has consistently performed superiorly
models across all sequences. The results indicate that the, R2 value among all five competing prediction models with all the input time
tends to increase with the number of input sequences, as illustrated in sequence lengths.
Table 7. The average absolute difference between predicted and actual
values, computed by MAE (Mean Absolute Error), indicates model
performance, with lower values suggesting higher accuracy and zero 5.1.1. Statistical measure using boxplots
indicating a perfect match between actual and predicted values. Ac­ As already stated, all competing prediction models are executed for
cording to Table 8, the stacked LSTM model demonstrates the lowest total 10 independent runs for a fair statistical comparison, and there­
average MAE value of 1.1157 among all the competing predictive fore the variability in all their respective considered output error me­
models. Moreover, the sMAPE is a metric that evaluates the accuracy of trics are depicted using boxplots in Fig. 12. A box plot is a graphical tool
forecasting models by comparing actual and predicted values, where that effectively illustrates the centering, spread, and distribution of a
smaller values signify higher accuracy. From Table 9, it is further evi­ continuous dataset. It provides visual insights into how samples are
dent that the stacked LSTM model consistently achieves the lowest distributed and separated. These boxplots present five key statistics: the
average sMAPE value of 1.1795 across all four sequences, out­ minimum, lower quartile (Q1), median (Q2), upper quartile (Q3), and
performing the other competitive models. From the presented com­ maximum. In Fig. 12, the red horizontal lines represent the medians of
prehensive analysis of all obtained numerical results, it clearly signifies the RMSE, Explained Variance, R2 , and MAE metrics. The dashed lines
that the stacked LSTM model consistently outperforms other competing extending from the boxes (whiskers) depict the spread and shape of the
models across various error indices. From Fig. 9, we can visually distribution of these metrics. Any individual data points beyond the
compare the actual output SPV power with the predicted SPV power whiskers (bubbles) are considered outliers, providing additional in­
generated by all models utilized across all four input sequences. It is sights into the variability of the independent runs for each algorithm.

Fig. 9. Actual and predicted value plots (a) Data Sequence-A: 3-Months, (b) Data Sequence-B: 6-Months, (c) Data Sequence-C: 9-months, and (d) Data Sequence-D:
12-months.

13
P.K. Singh, A. Saraswat and Y. Gupta Next Energy 11 (2026) 100531

Fig. 10. A comparative illustration of training errors obtained for (a) Data Sequence-A: 3-Months, (b) Data Sequence-B: 6-Months, (c) Data Sequence-C: 9-months,
and (d) Data Sequence-D: 12-months.

Fig. 11. A comparative illustration of validation errors obtained for (a) Data Sequence-A: 3-Months, (b) Data Sequence-B: 6-Months, (c) Data Sequence-C: 9-months,
and (d) Data Sequence-D: 12-months.

5.1.2. Diebold Mariano test results The DM test results on absolute error loss obtained by ANN, LR,
This section provides a further statistical comparison of forecast XGBoost, SVR, 1D-CNN, CNN-LSTM with respect to Stacked-LSTM are
performance for all the competing prediction models using the Diebold- presented in Table 10. The higher absolute value of the DM statistics
Mariano (DM) test. This statistical test evaluates two hypotheses as indicates a significant difference in predictive accuracy between the
described by Eqs. (6)-(9). The null hypothesis assumes that there is no two models. The consistently negative DM statistics for all the seven
difference in forecast accuracy between two competing models on the benchmark models across all four input data sequences indicate that the
same input dataset. stacked LSTM archives superior predictive performance w.r.t. all other

14
P.K. Singh, A. Saraswat and Y. Gupta Next Energy 11 (2026) 100531

Fig. 12. Boxplots for performance metrics for a data sequence of 3 months.

Table 10
Diebold Mariano Test Results for all benchmark forecasting models w.r.t. Stacked LSTM.

Prediction Model Absolute-error loss function, |dt| with stacked LSTM

3 MONTHS 6 MONTHS 9 MONTHS 12 MONTHS

DM – Test p – value DM - Test p – value DM - Test p - value DM - Test p - value

ANN -21.0953 0.0000 -31.2889 0.0000 -12.5628 0.0000 -18.1171 0.0000


LR -2.4005 0.0164 -42.7302 0.0000 -41.1679 0.0000 -29.8710 0.0000
XGBoost -44.7440 0.0000 -0.7869 0.0431 -3.4948 0.0005 -35.7860 0.0000
SVR -25.4429 0.0000 -91.1154 0.0000 -63.9570 0.0000 -89.8933 0.0000
1D-CNN -26.4332 0.0000 -8.7852 0.0000 -43.1779 0.0000 -6.3162 0.0000
CNN-LSTM -26.4332 0.0000 -10.9237 0.0000 -19.1543 0.0000 -4.9442 0.0000
BI-LSTM -32.4639 0.0000 -19.3708 0.0000 -17.9218 0.0000 -14.8844 0.0000

competing models while evaluated using the absolute error loss func­ compared to the rest competing prediction models i.e. ANN, LR,
tion. These DM test results confirm that all the p-values of DM statistics XGBoost, SVR, 1D-CNN, CNN-LSTM.
are less than 0.05, which suggests that the null hypothesis is rejected,
and the alternative hypothesis (i.e. stacked LSTM model) is accepted.
More precisely, the observed differences are statistically significant, 5.2. Evaluating uncertainty quantification / prediction intervals
indicating that the forecasting accuracy of the stacked LSTM model is
superior to that of the other models. These DM test results further Moreover, the output prediction interval performances of all the
support superior predictive accuracy of Stacked LSTM model as competing forecasting models are also validated using Prediction
Interval Coverage Probability (PICP) and Mean Prediction Interval
Table 11
Comparison of Prediction Interval Coverage Probability (PICP) and Mean Prediction Interval Width (MPIW) for all forecasting models at the 90% confidence level.

Prediction Model Data Sequence A Data Sequence B Data Sequence C Data Sequence D

PICP MPIW PICP MPIW PICP MPIW PICP MPIW

ANN 0.906 19.070 0.900 6.429 0.900 7.629 0.900 9.123


LR 0.900 17.924 0.901 7.979 0.902 8.929 0.902 8.291
XG Boost 0.900 18.066 0.900 6.351 0.900 5.782 0.900 8.581
SVR 0.900 14.066 0.900 8.979 0.904 9.629 0.900 8.991
1D CNN 0.900 13.861 0.900 6.979 0.900 13.702 0.906 8.641
CNN - LSTM 0.900 14.811 0.900 6.779 0.901 9.129 0.900 9.491
BI LSTM 0.900 15.708 0.900 6.911 0.900 8.929 0.903 8.856
Stacked LSTM 0.900 13.811 0.900 6.466 0.900 8.826 0.900 8.291

15
P.K. Singh, A. Saraswat and Y. Gupta Next Energy 11 (2026) 100531

Fig. 13. Revised data pre-processing pipeline avoiding any potential data leakage.

Table 12
Numerical comparison of prediction models with revised data pre-processing pipeline avoiding any potential data leakage for 12-month data sequence.

Prediction Model RMSE Explained Variance R2 MAE sMAPE


Mean ± Std. Mean ± Std. Mean ± Std. Mean ± Std. Mean ± Std.

ANN 2.8156 ± 0.06 0.8407 ± 0.07 0.8397 ± 0.05 1.4254 ± 0.04 1.2819 ± 0.06
LR 2.8756 ± 0.05 0.8312 ± 0.05 0.8310 ± 0.06 1.7447 ± 0.05 1.2695 ± 0.06
XG Boost 2.9185 ± 0.04 0.8551 ± 0.04 0.8511 ± 0.05 1.2988 ± 0.04 1.2910 ± 0.04
SVR 2.9245 ± 0.06 0.8519 ± 0.07 0.8501 ± 0.05 1.2918 ± 0.06 1.2936 ± 0.07
1D-CNN 2.9388 ± 0.05 0.8515 ± 0.05 0.8498 ± 0.04 1.3105 ± 0.04 1.2875 ± 0.05
CNN - LSTM 3.0810 ± 0.05 0.7926 ± 0.04 0.7909 ± 0.06 1.3876 ± 0.05 1.2847 ± 0.04
BI-LSTM 2.9125 ± 0.06 0.8319 ± 0.05 0.8301 ± 0.06 1.3032 ± 0.06 1.2855 ± 0.04
Stacked LSTM 2.7681 ± 0.03 0.8791 ± 0.05 0.8711 ± 0.04 1.2453 ± 0.05 1.2514 ± 0.04

Width (MPIW), as suggested in [59], for all four data sequences as ✓ These deep-learning models require a very large amount of data to
presented in Table 11. The output PICP values have consistently re­ work effectively. In the context of SPV forecasting applications, it
mained close to the nominal 0.90 for all competing models and data necessitates considerable size of historical weather data, solar ra­
sequences, indicating that these prediction intervals are well calibrated diation data, and SPV system output records.
and reliably capture the true values at the intended 90% confidence ✓ These deep-learning models frequently require lengthy training cy­
level. Besides, the output MPIW values vary across models and data cles, particularly with huge datasets, making them unfeasible for
sequences, reflecting differences in the precision of uncertainty quan­ rapid deployment or adaptation to changing conditions which is an
tification. Notably, Stacked LSTM model demonstrates comparatively inherent characteristic of SPV generation.
narrower prediction intervals width as compared to other forecasting ✓ These methods are prone to overfitting, which means they perform
models with higher confidence in their forecasts. well on training data but badly on unseen data. It is especially
problematic in solar forecasting, as the algorithm may memorize
5.3. Evaluating impact of potential data leakage patterns in past data rather than learning generalizable properties.
Many factors influence SPV production, including seasonal fluc­
A generalized research methodology adopted for solar power fore­ tuations, weather variability, and regional disparities. These deep-
casting as shown in Fig. 1, may suffer with potential data leakage. It learning models may struggle to generalize under varied settings,
may be due to feature selection and normalization applied to complete resulting in overfitting to specific instances.
input dataset rather than only training data which reveals the test data ✓ A model based on data from one specific area may underperform
to the prediction model during its training stage. For further in­ when applied to another place with different weather patterns,
vestigations, an alternative preprocessing pipeline has been tried by terrain, or SPV system configurations. This reduces the scalability
avoiding any potential data leakages where the data splitting step is and generalizability of deep learning models.
carried out on input dataset prior to feature selection and data nor­
malization as shown in Fig. 13. Here, firstly the feature selection has 6. Conclusion
been performed only on the training data using Pearson correlation
analysis, and subsequently min-max scale has been applied to nor­ This paper presents a comprehensive performance analysis of
malize the training data by keeping test data exclusively unrevealed to four deep learning-based prediction models i.e. 1D-CNN, Stacked
the forecasting models during their training stage. LSTM, Bi-LSTM, and a hybrid CNN-LSTM for short-term SPV power
All the competing prediction models are again executed for ten in­ generation forecasting application. To further aid this comprehen­
dependent runs with the revised data pre-processing pipeline avoiding any sive case study, four distinct machine learning models, namely ANN,
potential data leakage, and their respective output performance metrics are LR, SVR, and XGBoost, are also implemented and compared. All
compared as presented in Table 12. The output performance metric values these eight prediction models are tested and validated on meteor­
are inferior as compared to the same as obtained with the previous full ological data collected from the DKASC Alice Springs site, covering a
dataset normalization and feature selection as presented in Tables 5–9. complete year (April 2016 to March 2017). This one-year dataset is
However, this revised pipeline ensures that no information from the test categorized into four time-series sequences: 3-month duration (Data
data is used during feature selection or normalization, thereby preventing Sequence-A), 6-month duration (Data Sequence-B), 9-month dura­
data leakage and providing a more rigorous assessment of performance tion (Data Sequence-C), and 12-month duration (Data Sequence-D),
generalization. Moreover, it is further evident that the stacked LSTM model to facilitate extensive numerical analysis and performance evalua­
still outperformed all other competing prediction models in terms of all the tions. These data sequences are preprocessed to remove noise, and
five metrics of performance evaluation as depicted in Table 12. the developed models were then applied to predict SPV power
generation for 30 minutes ahead of real time. The performance of the
5.4. Limitations of deep-learning models for solar-PV generation predictions models was evaluated using five error indices: RMSE, MAE, R2,
Explained Variance, and SMAPE. Among these eight models, the
Besides having several advantages of deep-learning methods for SPV Stacked LSTM demonstrated the highest prediction accuracy, parti­
power generation forecasting, they also come with a few algorithmic cularly when trained on the 6-month data sequence. More specifi­
limitations as detailed below: cally, the Stacked LSTM model is capable to achieve superior

16
P.K. Singh, A. Saraswat and Y. Gupta Next Energy 11 (2026) 100531

average values for RMSE, MAE, R2, Explained Variance, and sMAPE 102326, [Link]
as 2.3408, 1.1157, 0.9004, 0.8998, and 1.1795, respectively, which [10] M.G. De Giorgi, P.M. Congedo, M. Malvoni, D. Laforgia, Error analysis of hybrid
photovoltaic power forecasting models: A case study of Mediterranean climate,
indicates that it outperforms the other competing prediction models Energy Convers. Manag. 100 (2015) 117–130, [Link]
across all the data sequences. enconman.2015.04.078.
Here it is worth mentioning that the present research work has a few [11] J. Alonso-Montesinos, F.J. Batlles, Solar radiation forecasting in the short-and
medium-term under all sky conditions, Energy 83 (2015) 387–393, [Link]
limitations such as a single location data is considered, and only short- 10.1016/[Link].2015.02.036.
term forecasting is performed. Moreover, this work is also not focused [12] X. Ruhang, The restriction research for urban area building integrated grid-con­
on external influences such as shading effects, dust accumulation, or nected PV power generation potential, Energy 113 (2016) 124–143, [Link]
org/10.1016/[Link].2016.07.035.
grid constraints, which may impact the solar power generation. For [13] M.A.F. Lima, P.C. Carvalho, L.M. Fernández-Ramírez, A.P. Braga, Improving solar
future research directions, the use of larger and more diverse datasets is forecasting using deep learning and portfolio theory integration, Energy 195 (2020)
recommended to enhance the robustness of the prediction models. In 117016, [Link]
[14] G.M. Yagli, D. Yang, D. Srinivasan, Automatic hourly solar forecasting using ma­
this study, the data was collected from a single location, which may
chine learning models, Renew. Sustain. Energy Rev. 105 (2019) 487–498, https://
limit the generalizability of the models when applied to other regions [Link]/10.1016/[Link].2019.02.006.
with different weather patterns, terrain, or SPV system configurations. [15] B. Jin, X. Xu, Y. Zhang, Thermal coal futures trading volume predictions through
Therefore, training these models on heterogeneous datasets from mul­ the neural network, J. Model. Manag. (2024), [Link]
2023-0207.
tiple sources could improve their performance and adaptability. [16] X. Xu, Y. Zhang, House price forecasting with neural networks, Intell. Syst. Appl. 12
(2021) 200052, [Link]
Ethical Statement [17] B. Jin, X. Xu, Forecasting wholesale prices of yellow corn through the Gaussian
process regression, Neural Comput. Applic 36 (2024) 8693–8710, [Link]
10.1007/s00521-024-09531-2.
This study does not contain any studies with human or animal [18] B. Jin, X. Xu, Predictions of steel price indices through machine learning for the
subjects performed by any of the authors. regional northeast Chinese market, Neural Comput. Applic 36 (2024)
20863–20882, [Link]
[19] B. Jin, X. Xu, Machine learning predictions of regional steel price indices for east
Data availability China, Ironmak. Steelmak. (2024), [Link]
[20] F. Harrou, F. Kadri, Y. Sun, Forecasting of photovoltaic solar power production
using LSTM approach, Adv. Stat. Model. Forecast. Fault Detect. Renew. Energy Syst.
The data that support the findings of this study are openly available (2020), [Link]
in DKASC. (n.d.). Alice Springs at [Link] [21] M. Seyedmahmoudian, E. Jamei, G. Thirunavukkarasu, T. Soon, M. Mortimer,
locations/alice-springs?source=1B B. Horan, A. Stojcevski, S. Mekhilef, Short-term forecasting of the output power of a
building-integrated photovoltaic system using a metaheuristic approach, Energies
11 (5) (2018) 1260, [Link]
Declaration of Competing Interest [22] F. Baser, H. Demirhan, A fuzzy regression with support vector machine approach to
the estimation of horizontal global solar radiation, Energy 123 (2017) 229–240,
The authors declare that they have no known competing financial [Link]
[23] M. Pan, C. Li, R. Gao, Y. Huang, H. You, T. Gu, F. Qin, Photovoltaic power fore­
interests or personal relationships that could have appeared to influ­
casting based on a support vector machine with improved ant colony optimization,
ence the work reported in this paper. J. Clean. Prod. 277 (2020) 123948, [Link]
123948.
[24] P.K. Singh, A. Saraswat, Y. Gupta, S.K. Goyal, Prediction of short-term solar ra­
Acknowledgement
diation using machine learning methods, Lect. Notes Electr. Eng. 863 (2022)
181–192, [Link]
The authors are grateful to recognize the software simulation and [25] K. Nam, S. Hwangbo, C. Yoo, A deep learning-based forecasting model for renew­
data analysis facility available at Power System Laboratory of Manipal able energy scenarios to guide sustainable energy policy: A case study of Korea,
Renew. Sustain. Energy Rev. 122 (2020) 109725, [Link]
University Jaipur, Rajasthan, India to carry out the present research 2020.109725.
work. [26] H. Wang, Z. Lei, X. Zhang, B. Zhou, J. Peng, A review of deep learning for renewable
energy forecasting, Energy Convers. Manag. 198 (2019) 111799, [Link]
10.1016/[Link].2019.111799.
References [27] K. Wang, X. Qi, H. Liu, A comparison of day-ahead photovoltaic power forecasting
models based on deep learning neural network, Appl. Energy 251 (2019) 113315,
[1] W. VanDeventer, E. Jamei, G.S. Thirunavukkarasu, M. Seyedmahmoudian, [Link]
T.K. Soon, B. Horan, S. Mekhilef, S. Stojcevski, Short-term PV power forecasting [28] M. Mishra, P.B. Dash, J. Nayak, B. Naik, S.K. Swain, Deep learning and wavelet
using hybrid GASVM technique, Renew. Energy 140 (2019) 367–379, [Link] transform integrated approach for short-term solar PV power prediction,
org/10.1016/[Link].2019.02.087. Measurement 166 (2020) 108250, [Link]
[2] S.R. Ola, A. Saraswat, S.K. Goyal, S.K. Jhajharia, B. Khan, O.P. Mahela, H. Haes 108250.
Alhelou, P. Siano, A protection scheme for a power system with solar energy pe­ [29] J. Heo, K. Song, S. Han, D.E. Lee, Multi-channel convolutional neural network for
netration, Appl. Sci. 10 (4) (2020) 1516, [Link] integration of meteorological and geographical features in solar power forecasting,
[3] A. Dairi, F. Harrou, Y. Sun, S. Khadraoui, Short-term forecasting of photovoltaic Appl. Energy 295 (2021) 117083, [Link]
solar power production using variational auto-encoder driven deep learning ap­ 117083.
proach, Appl. Sci. 10 (23) (2020) 8400, [Link] [30] F. Wang, Z. Lei, X. Zhang, B. Zhou, J. Peng, J. Li, Deep learning-based irradiance
[4] M. Elsaraiti, A. Merabet, Solar power forecasting using deep learning techniques, mapping model for solar PV power forecasting using sky image, IEEE Ind. Appl. Soc.
IEEE Access 10 (2022) 31692–31698, [Link] Annu. Meet. (2019), [Link]
3160484. [31] M. Tovar, M. Robles, F. Rashid, PV power prediction using CNN-LSTM hybrid
[5] B.P. Singh, S.K. Goyal, S.A. Siddiqui, A. Saraswat, R. Ucheniya, Intersection point neural network model: Case study of Temixco-Morelos, México, Energies 13 (24)
determination method: A novel MPPT approach for sudden and fast changing en­ (2020) 6512, [Link]
vironmental conditions, Renew. Energy 200 (2022) 614–632, [Link] [32] A. Mellit, A.M. Pavan, V. Lughi, Deep learning neural networks for short-term
1016/[Link].2022.09.056. photovoltaic power forecasting, Renew. Energy 172 (2021) 276–288, [Link]
[6] P.K. Singh, A. Saraswat, Y. Gupta, S.K. Goyal, Y. Gupta, A comparative study of org/10.1016/[Link].2021.02.166.
deep learning methods for short-term solar radiation forecasting, Lect. Notes Electr. [33] A. Hamad, H. Shabana, I. Muhammad, Solar power prediction using dual stream
Eng. 1065 (2023), [Link] CNN-LSTM architecture, Sensors 23 (2) (2023) 945, [Link]
[7] U.K. Das, K.S. Tey, M. Seyedmahmoudian, S. Mekhilef, M.Y.I. Idris, s23020945.
W. VanDeventer, B. Horan, A. Stojcevski, Forecasting of photovoltaic power gen­ [34] D.K. Dhaked, S. Dadhich, D. Birla, Power output forecasting of solar photovoltaic
eration and model optimization: A review, Renew. Sustain. Energy Rev. 81 (2018) plant using LSTM, Green. Energy Intell. Transp. 2 (5) (2023) 100113, [Link]
912–928, [Link] org/10.1016/[Link].2023.100113.
[8] M.N. Akhter, S. Mekhilef, H. Mokhlis, N.M. Shah, Review on forecasting of pho­ [35] H. Zang, L. Cheng, T. Ding, K.W. Cheung, Z. Wei, G. Sun, Day-ahead photovoltaic
tovoltaic power generation based on machine learning and metaheuristic techni­ power forecasting approach based on deep convolutional neural networks and
ques, IET Renew. Power Gener. 13 (7) (2019) 1009–1023, [Link] meta-learning, Int. J. Electr. Power Energy Syst. 118 (2020) 105790, [Link]
1049/iet-rpg.2018.5649. org/10.1016/[Link].2019.105790.
[9] X. Luo, D. Zhang, An adaptive deep learning framework for day-ahead forecasting [36] B. Ray, R. Shah, M.R. Islam, S. Islam, A new data-driven long-term solar yield
of photovoltaic power generation, Sustain. Energy Technol. Assess. 52 (2022) analysis model of photovoltaic power plants, IEEE Access 8 (2020)

17
P.K. Singh, A. Saraswat and Y. Gupta Next Energy 11 (2026) 100531

136223–136233, [Link] Energy 185 (2019) 387–405, [Link]


[37] H. Wang, H. Yi, J. Peng, G. Wang, Y. Liu, H. Jiang, W. Liu, Deterministic and [58] C. Chatfield, Prediction intervals for time-series forecasting, A handbook for re­
probabilistic forecasting of photovoltaic power based on deep convolutional neural searchers and practitioners, Principles of forecasting, Springer US, Boston, MA,
network, Energy Convers. Manag. 153 (2017) 409–422, [Link] 2001, pp. 475–494, [Link]
enconman.2017.10.008. [59] K.D. Poti, R.M. Naidoo, N.T. Mbungu, R.C. Bansal, Optimal hybrid power dispatch
[38] L. Wen, K. Zhou, S. Yang, X. Lu, Optimal load dispatch of community microgrid through smart solar power forecasting and battery storage integration, J. Energy
with deep learning-based solar power and load forecasting, Energy 171 (2019) Storage 86 (2024) 111246, [Link]
1053–1065, [Link]
[39] A.A. Mansour, A. Tilioua, M. Touzani, Bi-LSTM, GRU, and 1D-CNN models for
short-term photovoltaic panel efficiency forecasting: Case amorphous silicon grid-
connected PV system, Results Eng. 21 (2024) 101886, [Link] PRAVEEN KUMAR SINGH holds an MCA, MBA, and [Link].
rineng.2024.101886. in Physics, and is currently pursuing a Ph.D. in Computer
[40] J. Zheng, H. Zhang, Y. Dai, B. Wang, T. Zheng, Q. Liao, Y. Liang, F. Zhang, X. Song, Applications from Manipal University Jaipur. He is pre­
Time series prediction for output of multi-region solar power plants, Appl. Energy sently serving as an Assistant Professor at Maharaja Agrasen
257 (2020) 114001, [Link] Institute of Management Studies, Delhi. His research areas
[41] F. Wang, Z. Xuan, Z. Zhen, K. Li, T. Wang, M. Shi, A day-ahead PV power fore­ of interest include Artificial Intelligence, Machine Learning,
casting method based on LSTM-RNN model and time correlation modification under Deep Learning, and Renewable Energy Technologies.
partial daily pattern prediction framework, Energy Convers. Manag. 212 (2020)
112766, [Link]
[42] X. Luo, D. Zhang, X. Zhu, Deep learning-based forecasting of photovoltaic power
generation by incorporating domain knowledge, Energy 225 (2021) 120240,
[Link]
[43] M. Gao, J. Li, F. Hong, D. Long, Day-ahead power forecasting in a large-scale
photovoltaic plant based on weather classification using LSTM, Energy 187 (2019)
115838, [Link]
[44] X. Qing, Y. Niu, Hourly day-ahead solar irradiance prediction using weather fore­
AMIT SARASWAT received his Ph.D. (Electrical Power
casts by LSTM, Energy 148 (2018) 461–468, [Link]
Systems), [Link]. (Engineering Systems) and [Link].
2018.01.177.
(Electrical Engineering) from Faculty of Engineering,
[45] P. Li, K. Zhou, X. Lu, S. Yang, A hybrid deep learning model for short-term PV power
Dayalbagh Educational Institute, (Deemed University),
forecasting, Appl. Energy 259 (2020) 114216, [Link]
Agra, India in 2013, 2006 and 2003 respectively. Presently,
2019.114216.
he is currently working as Professor in Department of
[46] G. Li, S. Xie, B. Wang, J. Xin, Y. Li, S. Du, Photovoltaic power forecasting with a
Electrical Engineering, Manipal University Jaipur,
hybrid deep learning approach, IEEE Access 8 (2020) 175871–175880, [Link]
Rajasthan, India since July 2014. His research area includes
org/10.1109/ACCESS.2020.3025860.
Grid Integration Studies for Renewable Energy Systems,
[47] K. Wang, X. Qi, H. Liu, Photovoltaic power forecasting-based LSTM-Convolutional
Reactive Power & Congestion Management, Competitive
Network, Energy 189 (2019) 116225, [Link]
Electricity Markets, Competitive Bidding, Multi-objective
[48] H. Sharadga, S. Hajimirza, R.S. Balog, Time series forecasting of solar power gen­
Evolutionary Methods and Applications of Soft-Computing
eration for large-scale photovoltaic plants, Renew. Energy 150 (2020) 797–807,
Techniques in Power System Optimization etc. He has au­
[Link]
thored many research articles in high-impact peer-reviewed International/National
[49] DKASC. (n.d.). Alice Springs. Retrieved from 〈[Link]
Journals and conference proceedings, and many chapters in book series conference. He is
locations/alice-springs?source=1B〉.
a life member of the Indian Society for Technical Education (India), and Senior Member
[50] Z.A. Khan, T. Hussain, I.U. Haq, F.U.M. Ullah, S.W. Baik, Towards efficient and
of IEEE (USA) and IEEE Power & Energy Society. He was a recipient of “ITSR Foundation
effective renewable energy prediction via deep learning, Energy Rep. 8 (2022)
Award-2023” for Excellence in Leadership Awarded facilitated by Institute of Technical
10230–10243, [Link]
and Scientific Research, and The Institution of Engineers (India) – Rajasthan State Centre,
[51] F.X. Diebold, R.S. Mariano, Comparing predictive accuracy, J. Bus. Econ. Stat. 20
Jaipur, India.
(1) (2002) 134–144 〈[Link]
[52] P. Lauret, R. Alonso-Suárez, J. Le Gal La Salle, M. David, Solar forecasts based on
the clear sky index or the clearness index: Which is better? (MDPI), Solar 2 (4)
(2022, October) 432–444, [Link] YOGESH GUPTA is a distinguished academician and re­
[53] T. Peng, C. Zhang, J. Zhou, M.S. Nazir, An integrated framework of Bi-directional searcher, currently serving as a Professor in the Department
long-short term memory (BiLSTM) based on sine cosine algorithm for hourly solar of Computer Science and Engineering at BML Munjal
radiation forecasting, Energy 221 (2021) 119887, [Link] University, Gurugram, India. He holds a Bachelor’s degree
energy.2021.119887. in Information Technology and a Ph.D. from Dayalbagh
[54] A.Y. Barrera-Animas, L.O. Oyedele, M. Bilal, T.D. Akinosho, J.M.D. Delgado, Educational Institute, Agra. With expertise in Information
L.A. Akanbi, Rainfall prediction: A comparative analysis of modern machine Retrieval, Machine Learning, Big Data, and Soft Computing
learning algorithms for time-series forecasting, Mach. Learn. Appl. 7 (2022) Techniques, Prof. Gupta has made significant contributions
100204, [Link] to the field through numerous publications in reputed
[55] A.S. AlSalehy, M. Bailey, Improving Time Series Data Quality: Identifying Outliers journals and international conferences. His research focuses
and Handling Missing Values in a Multilocation Gas and Weather Dataset, Smart on leveraging advanced computational techniques to ad­
Cities 8 (3) (2025) 82, [Link] dress complex real-world problems. In addition to his aca­
[56] D.J. Bae, B.S. Kwon, K.B. Song, XGBoost-based day-ahead load forecasting algo­ demic and research endeavors, he is a lifetime member of
rithm considering behind-the-meter solar PV generation, Energies 15 (1) (2021) the Computer Society of India, reflecting his commitment to the advancement of com­
128, [Link] puter science and engineering. Prof. Gupta continues to inspire and mentor students,
[57] S. Pereira, P. Canhoto, R. Salgado, M.J. Costa, Development of an ANN-based correc­ fostering innovation and excellence in the field of computing.
tive algorithm of the operational ECMWF global horizontal irradiation forecasts, Sol.

18

You might also like