0% found this document useful (0 votes)
21 views25 pages

Solar Radiation Prediction Using ML & DL

This research article presents a multi-model approach for solar radiation prediction using machine learning (ML) and deep learning (DL) techniques, focusing on meteorological data from 1988 to 2022. The study identifies gradient boosting regressor (GBR) as the best model for long-term forecasts and recurrent neural network (RNN) for short-term predictions, while also introducing a hybrid GBR-RNN model that outperforms individual models. The findings enhance solar irradiance forecasting accuracy and address input uncertainty, contributing to improved energy management in renewable energy systems.

Uploaded by

22jr1a1293
Copyright
© All Rights Reserved
We take content rights seriously. If you suspect this is your content, claim it here.
Available Formats
Download as PDF, TXT or read online on Scribd
0% found this document useful (0 votes)
21 views25 pages

Solar Radiation Prediction Using ML & DL

This research article presents a multi-model approach for solar radiation prediction using machine learning (ML) and deep learning (DL) techniques, focusing on meteorological data from 1988 to 2022. The study identifies gradient boosting regressor (GBR) as the best model for long-term forecasts and recurrent neural network (RNN) for short-term predictions, while also introducing a hybrid GBR-RNN model that outperforms individual models. The findings enhance solar irradiance forecasting accuracy and address input uncertainty, contributing to improved energy management in renewable energy systems.

Uploaded by

22jr1a1293
Copyright
© All Rights Reserved
We take content rights seriously. If you suspect this is your content, claim it here.
Available Formats
Download as PDF, TXT or read online on Scribd

RESEARCH ARTICLE | MAY 01 2025

Solar radiation prediction: A multi-model machine learning


and deep learning approach
C Vanlalchhuanawmi; Subhasish Deb  ; Md. Minarul Islam  ; Taha Selim Ustun

AIP Advances 15, 055201 (2025)


[Link]


View Export
Online Citation

Articles You May Be Interested In

Development of a solar radiation measuring instrument for building energy management system
Rev. Sci. Instrum. (May 2025)

Optimized solar power forecasting: A multi-decomposition framework using VMD and swarm techniques
AIP Advances (September 2025)

Game-based resource allocation in heterogeneous downlink CR-NOMA network without subchannel


sharing

26 December 2025 13:17:38


AIP Advances (January 2025)
AIP Advances ARTICLE [Link]/aip/adv

Solar radiation prediction: A multi-model


machine learning and deep learning approach
Cite as: AIP Advances 15, 055201 (2025); doi: 10.1063/5.0237246
Submitted: 4 September 2024 • Accepted: 12 April 2025 •
Published Online: 1 May 2025

C Vanlalchhuanawmi,1 Subhasish Deb,1,a) Md. Minarul Islam,2,a) and Taha Selim Ustun3

AFFILIATIONS
1
Department of Electrical Engineering, Mizoram University, Aizawl, Mizoram-796004, India.
2
Department of Electrical and Electronic Engineering, University of Dhaka, Dhaka 1000, Bangladesh
3
Fukushima Renewable Energy Institute, AIST (FREA), Koriyama 9630298, Japan

a)
Authors to whom correspondence should be addressed: subhasishdeb30@[Link] and mimislam-eee@[Link]

ABSTRACT
The increasing integration of renewable energies into electrical grids necessitates accurate forecasting of meteorological variables, particu-
larly solar irradiance. This study presents a novel long-term solar irradiance forecasting approach, utilizing meteorological data from the
National Renewable Energy Laboratory spanning 1988–2022. Focusing on five input variables—solar irradiance, dew point, temperature,

26 December 2025 13:17:38


relative humidity, and wind speed—this study evaluates the predictive performance of 13 data-driven models, comprising ten machine learn-
ing (ML) and three deep learning (DL) algorithms. Among them, gradient boosting regressor (GBR) and recurrent neural network (RNN)
emerged as top performers in ML and deep learning, respectively. In order to choose the most suitable model for the long and short term,
four forecast time-horizons (1, 8, 16, and 24 h) were also taken into consideration for the accurate models. A feature selection process using
Pearson’s coefficient identified the most relevant inputs, while quantile regression was employed for uncertainty assessment, mean predic-
tion interval, and prediction interval coverage probability models. This study demonstrates that RNN excels in short-term predictions, while
GBR is more effective for long-term forecasts. A new hybrid approach GBR-RNN model was developed, achieving superior performance in
terms of RMSE, MAE, and R2 metrics. This multi-model approach, integrating both ML and DL techniques, enhances solar irradiance fore-
casting by addressing input uncertainty and considering various forecast horizons. The findings contribute to the ongoing advancement of
renewable energy forecasting by providing robust, accurate, and uncertainty-aware predictive models. Moreover, this approach helps identify
the best-performing model, enabling more reliable and precise solar irradiance forecasts for energy management. This highlights both the
improvement in forecasting methods and the importance of selecting the best model for accuracy.
© 2025 Author(s). All article content, except where otherwise noted, is licensed under a Creative Commons Attribution (CC BY) license
([Link] [Link]

I. INTRODUCTION impact on leaf size, growth rate, flower and fruit development, soil
temperature, and moisture content.1 Although working on numeri-
A vital renewable energy source for many uses, including agri- cal estimating models is a necessary substitute because it is difficult
culture, climate modeling, and the design of solar energy systems, to obtain measurements of solar energy directly, due to both tech-
is solar radiation (SR). Engineers may determine the best kind, nological and economic constraints, such as the high expense of
size, and solar panel orientation through the use of SR data in the installing and calibrating recording equipment and the requirement
development and optimization of solar energy systems for specific for continuous maintenance, direct measurements of SR are uncom-
environments. The understanding of how Earth’s temperature and mon in most places. As a result, modeling strategies for SR are
weather patterns are affected is crucial for climate modeling, which regarded as important.
likewise largely depends on SR. Since photosynthesis, the process With an emphasis on sensor networks for forecasting, some
by which plants turn sunlight into energy depends on sunshine, review research investigates solar irradiance resources and solar
SR is essential to agricultural plant development and crop produc- forecasting. An overview of forecasting techniques, forecast error
tion. Plant growth and agricultural operations can be impacted by measures, radiometers, sensor network datasets, and sun irradiance
the kind and amount of sunshine received, which can also have an resources is provided. Three primary categories may be used to

AIP Advances 15, 055201 (2025); doi: 10.1063/5.0237246 15, 055201-1


© Author(s) 2025
AIP Advances ARTICLE [Link]/aip/adv

group solar forecasting techniques: data-driven approaches, image- for solar forecasting and, giving considerable gains in accuracy and
based methods, and numerical weather predictions (NWPs). While efficiency over older techniques.11 Projecting the amount of solar
image-based approaches are favored for intra-hour and 0.5–6-h energy that photovoltaic (PV) systems will produce over a range of
forecasting, NWPs are appropriate for predicting 6 to 48 h in periods, from minutes to days, is known as solar forecasting. For
advance.2 Based on input data, data-driven approaches may be used solar power to be successfully integrated into the electrical grid,
for a broad variety of forecast horizons. Self-powered sensors are grid stability to be increased, and the operation and maintenance
now a part of wireless sensor technology, allowing for real-time of solar power plants to be optimized, accurate solar projections
parameter adjustments. A vastly dispersed limitless supply is offered are essential. DL substantially boosts solar forecasting by apply-
by energy harvesting technologies, including piezoelectricity, solar ing sophisticated techniques including recurrent neural networks
light, electromagnetic fields, radio frequency, and physical move- (RNNs), long short-term memory networks (LSTMs), convolu-
ments.1 The best ambient source of solar energy is thought to be tional neural networks (CNNs), artificial neural networks (ANN),
photovoltaic cell modules because of their high-power density, effec- and hybrid models. These models employ multiple data sources,
tiveness of conversion, and interoperability with integrated circuit including historical solar output, weather data, satellite photos, sky
technologies. Since the energy sector in the EU contributes to more cameras, and NWP models. Applications of DL in solar forecasting
than 75% of greenhouse gas emissions, improving the contribution include grid management, energy trading, operational optimization,
of renewable energy in all sectors is essential to attaining a climate- and microgrid integration, all of which benefit from better accuracy
neutral continent by 2050 and a net reduction of emissions by at least and efficiency. However, issues such as data quality, model inter-
55% by 2030.3 pretability, and computational needs persist. Future developments
Complex non-linear correlations may be found by using in the field focus on transfer learning, explainable artificial intel-
machine learning (ML) algorithms to examine vast datasets and ligence, and real-time forecasting to further boost the usefulness
find patterns and interactions between SR and the input parameters. of DL in solar energy applications. Researchers categorize the esti-
Several ML models, such as support vector regression, random for- mates of solar power and irradiance according to several elements,
est (RF), and k-nearest neighbor algorithm (KNN), have been used including satellite imagery, regional and meteorological features,
for SR estimation.4–6 The electromagnetic waves that the sun emits and cloud imaging.12 Although for solar forecasting, there are not
are known as solar radiation, which include visible light, ultraviolet any recognized categorization standards, forecast scales, historical
(UV) rays, and infrared (IR) radiation. Solar irradiance measures the data, and meteorological data models form the basis of the majority

26 December 2025 13:17:38


power of solar energy received per unit area at a specific moment, of forecasts. Time horizon is the primary categorization criterion,
while SR typically refers to the total energy received over a period and projections for different time horizons have a major impact on
of time and includes the broader spectrum of solar emissions. A lit- different grid operating components.13,14
tle portion of this radiation makes it to Earth’s atmosphere, where The field of machine learning includes the general applica-
gases, particles, and clouds can absorb, scatter, or reflect it. Some tion of computers to learn from data and the implicit programming
of this radiation makes it to the surface of the Earth, where the of computer algorithms to carry out specific tasks. Nevertheless,
Sun powers all life. Sea level variations are influenced by pro- deep learning is a development of machine learning that uses
cesses including evaporation, condensation, and precipitation, all of ANN-mimicked models for feature mapping and data learning. In
which depend on as a primary energy source in the earth’s water Ref. 15, the author proposes a robust approach to power usage pre-
cycle.7 In addition, it fuels atmospheric instability, which produces diction considering different trends in ML and DL models. The
extreme weather occurrences and catastrophic weather phenom- energy balance of Earth is dependent on solar radiation, which also
ena such as storms and hurricanes. The principal energy source affects heat fluxes, hydrological cycles, ecosystems, and tempera-
for ecosystems and natural processes on earth is the most plentiful ture. Compared to fossil fuels, it has less of an adverse effect on
and easily accessible energy source. In addition to being safe, eas- the environment when used to generate electricity.16,17 Furthermore,
ily available, and non-polluting, it also has the ability to lessen the detrimental effects on the environment such as global warming,
intensity of the greenhouse effect. There are three primary compo- and the production of renewable energy are becoming more and
nents of solar radiation: extra radiation, diffuse radiation, and direct more dependent on renewable resources. Predicting solar energy
radiation. Gaining an understanding of these elements is essential can be done with ML, which enables more precise estimates, better
for researching microclimates, constructing energy efficient build- planning, and grid integration. This contributes to the rising usage
ing lighting, and improving solar energy systems. Serious problems of solar energy by increasing efficiency, lowering costs, generat-
such as industrial air pollution, global warming, and environmen- ing employment opportunities, and reducing carbon emissions.18,19
tal damage are being faced by the world. Because non-renewable Based on the findings of the articles evaluated, different algorithms
fuels such as coal, oil, and natural gas have negative effects on the were employed to estimate solar radiation in different areas; yet,
environment, finding other options is crucial. A more sustainable a comparison between machine learning and deep learning mod-
future may be achieved by addressing these problems head-on using els was not taken into consideration. Further research is necessary
clean energy, including solar power. Making the switch to clean for the forecasting of diffuse and beam solar radiation on inclined
energy may reduce global warming, clean up the air, increase energy surfaces.
security, protect ecosystems, reduce energy poverty, and encourage Research on DL-based forecasting models has been extensive,
long-term economic growth.8 but few studies have addressed how meteorological data creation
Another tool is deep learning (DL)9 which is used in differ- affects the precision of PV power forecasts. Due to short datasets
ent engineering contexts.10 Recently, it has become a potent tool that are insufficient for training deep neural networks, the majority

AIP Advances 15, 055201 (2025); doi: 10.1063/5.0237246 15, 055201-2


© Author(s) 2025
AIP Advances ARTICLE [Link]/aip/adv

of research produces forecasting models that suffer from inaccuracy, PV performance.24 To provide more accurate predictions of solar
especially during harsh weather. The generated model was tested radiation, researchers are still examining various machine learn-
using data from 34 meteorological stations, and the coefficient of ing and deep learning techniques, despite the enormous body of
determination was determined to be 0.98.20 Some studies lay the literature already available on the subject. Various approaches,
groundwork for future research in several areas, including energy including machine learning and deep learning, have been pro-
consumption forecasting across medium and long timeframes, solar posed as the best for forecasting GSR in a variety of situations
power, and wind power.21 It emphasizes the production and bal- with short-term data that contains the majority of just around ten
ancing of meteorological data to enhance prediction performance. years.25
However longer-term models currently lack training data and the This research investigates many patterns in three DL and ten
majority of research focuses on short-term estimates.22 Therefore, ML models. Alongside more standard ML methods such as lin-
a lot of research has been done on neural networks and DL mod- ear regression, Lasso, ridge, elastic net, random forest, Extra Trees,
els for forecasting solar irradiance. Nevertheless, there is a dearth XGBoost, LightGBM, K-Neighbors, and GBR, it examines the effi-
of thorough analyses about the importance of DL in this context. cacy of numerous DL approaches such as LSTM, RNN, and Gated
The research employs an ensemble feature selection method based Recurrent Unit (GRU). This work is the first to analyze the perfor-
on Pearson’s correlation coefficient and assesses prediction accu- mance of machine learning and deep learning models, even though
racy with multiple metrics. It highlights the limitations of existing earlier studies in the literature have incorporated all the models
methods, identifies key factors affecting forecasting accuracy, and taken into consideration for solar radiance prediction. To the best
suggests directions for future research, emphasizing the importance of the authors’ knowledge, no prior research has taken into account
of advanced ML and DL techniques in enhancing solar irradiance their uncertainties on ML and DL models, or the forecast of the
predictions.23 DL techniques were shown to be appropriate for the optimal prediction model for both short- and long-term model
different prediction tasks connected to solar energy. In addition, a predictions. The majority of earlier work has only addressed SR pre-
quick and effective DL method for solar irradiance prediction was diction in machine learning. Considering that the data use is spread,
proposed, including multi-reservoir echo state computation. When the resulting models are tested and trained on hourly data collected
compared to Elman neural networks, the model was shown to be over 20 years, and their performance is verified using 4-year data,
better appropriate for solar prediction tasks. The various systems’ outperforming that of previous research datasets. It identifies GBR as
methods of measuring solar energy in solar energy-based techno- the best ML model for long-term forecasting and RNN as the top DL

26 December 2025 13:17:38


logical applications vary. While the performance of certain other model for short-term predictions, using a Pearson coefficient-based
solar energy-based technologies is based on diffuse solar radiation feature selection and quantile regression for uncertainty assessment.
(DSR), global solar radiation (GSR) is employed for the system The models are evaluated using the coefficient of determination
computation in solar PV modeling. PV panels are vital for renew- (R2), MAE, RMSE, and MAPE. This section provides a thorough
able energy but suffer efficiency losses due to dust buildup and analysis of solar irradiance prediction while accounting for previ-
environmental factors. This article reviews dust causes, impacts, ous research findings Fig. 1 shows the flow chart of the suggested
models, cleaning methods, and sustainable solutions to maintain techniques.

FIG. 1. Flow chart for the suggested methodology.

AIP Advances 15, 055201 (2025); doi: 10.1063/5.0237246 15, 055201-3


© Author(s) 2025
AIP Advances ARTICLE [Link]/aip/adv

A. Focus and potential of the work included the years 2011–2021. A few figures show solar irradiance
The manuscript highlights the development of a hybrid fore- information.27,28 The Jupyter and Scikit–learn libraries are used to
casting model (GBR-RNN) aimed at enhancing the accuracy and generate the forecasting model for each period. This study’s dataset
reliability of long-term solar irradiance and energy demand pre- is divided into training and testing groups. The data are used to
dictions. To ensure clarity from the outset, the core focus of the calculate solar irradiance, dew point, temperature, relative humid-
study—namely, the integration of ML and DL models for improved ity and wind speed. Some of the calculations establish the optimal
forecasting—and to emphasize the potential of this approach in input variables to employ in model training.
advancing energy management strategies with the advance ML-DL
model. B. Characteristics of solar irradiation forecasting
From a practical standpoint, different prediction horizons serve
B. Motivation specific needs in decision-making within smart grids and micro-
grids. For activities such as PV storage control and power selling,
This research is motivated by the need to address key gaps in very short-term forecasting—from seconds to minutes—is essential.
solar irradiance forecasting, including limited model generalization, This timeframe has gained heightened importance in today’s smart
lack of uncertainty quantification, and narrow evaluation across grid and microgrid environments. To make decisions on energy
forecasting horizons. While previous studies have often focused markets and power system operations, such as unit commitment
on either ML or DL models in isolation and used small or local- and economic load dispatch, short-term forecasting, which lasts up
ized datasets, this work provides a comprehensive comparison of to two or three days, is crucial. Predicting for the medium term,
13 ML and DL models using more than 20 years of hourly data, up to seven days ahead, aids in scheduling maintenance for various
with performance validated on a separate 4-year dataset. By iden- power generation and transmission assets. Long-term forecasting,
tifying GBR as the best model for long-term prediction and RNN spanning from months to years, supports planning for solar energy
for short-term prediction, and incorporating quantile regression to projects and PV plant development.29,30 Both the prediction hori-
assess uncertainty, this study delivers a robust, accurate, and practi- zon and the choice of input variables have an impact on a prediction
cal forecasting framework. The goal is to offer a deployable solution model’s accuracy. Typically, important variables such as historical
suitable for consumer-level systems, advancing both the reliability data on PV generation and meteorological factors such as solar irra-
and applicability of solar forecasting models. diance, dew point, temperature, relative humidity, and wind speed

26 December 2025 13:17:38


are utilized in forecasting models. To develop accurate predictions, a
C. Contribution structured approach is employed, involving steps such as acquiring
A major contribution of this study is the development of a novel a comprehensive dataset; pre-processing to eliminate outliers and
hybrid model, GBR-RNN, which integrates the strengths of GBR normalize data; selecting relevant variables and their lag values using
and RNN to improve forecasting accuracy, as detailed in Table I. ensemble feature selection; partitioning data into training, valida-
The research highlights key gaps and limitations in existing stud- tion, and test sets; employing diverse ML algorithms; and evaluating
ies and underscores the importance of aligning model selection with forecasting algorithms using statistical metrics. This approach min-
forecasting horizons and accounting for input uncertainty. The pro- imizes bias toward extreme values and ensures precise forecasts. In
posed hybrid approach is designed to deliver accurate long-term this study, an ML-DL technique was evaluated independently using
predictions of solar irradiance and energy demand in hybrid energy measurement records spanning from 1998 to 2022, the majority of
systems, ultimately aiming to enhance the efficiency and reliability which included sun irradiance data.
of energy management.
The creation of these models, the research area containing the C. Solar irradiance components
SR components, and pre-processing are all thoroughly explained The essential variables, tools, and jargon that are often
in Sec. II. The models that have been carefully studied provide employed in solar irradiance forecasting techniques are explained.
Sec. III with support for the assessment measures used in this study. Through the process of absorption, reflections, and re-emissions, as
Section IV digs into the discussion of each model’s performance and it lowers, the sun’s total extraterrestrial beam irradiance (EBI) on
comparisons were made, while Sec. V gives the major conclusions Earth’s atmosphere decreases.
from the entire research. Direct normal irradiance (DNI) and direct horizontal irradi-
ance (DHI) make up the EBI incident on Earth’s surface. Global
horizontal irradiance (GHI), or solar irradiance incident on a hor-
II. MATERIALS AND METHODS
izontal plane on Earth’s surface, is the geometric sum of these
A. Area of study and description components. In fixed photovoltaic systems, GHI is employed for
A historical dataset of solar activity from January 1, 1998, to measure–correlate–predict assessments through comparisons with
December 31, 2022, has been collected from the NREL in order to the solar database.31 Usually, silicon reference cells and horizontal
evaluate the forecasting models’ effectiveness. Operating in GMT 0 pyranometers are used for the measurement as in Eq. (1),
time zone, NREL is the principal national laboratory for research GHI = DHI + DNI cos (θ). (1)
on renewable energy in the United States. It uses a baseline mea-
suring system at latitudes of −64.76 west and −5.03 north, at an Here, solar zenith angle is represented by θ.
elevation of 80 m, as depicted in Fig. 2. The testing dataset included A pyrheliometer and a rotating shadow-band irradiometer are
the last month record for the year 2022, while the training dataset used to measure the DNI. DHI is the shadow-free solar energy that

AIP Advances 15, 055201 (2025); doi: 10.1063/5.0237246 15, 055201-4


© Author(s) 2025
TABLE I. Categorization of the affined articles using different models including limitations.

Paper no. 1 12 17 18 25 26

© Author(s) 2025
Year 2023 2020 2023 2022 2023 2022
Field Focuses on improving Long-term or Energy load forecasting Forecasting direct solar The field of this study Time-horizons are
AIP Advances

solar forecasting large-scale error: and optimization using radiation and global falls under ML applied considered: 1, 2, and
through advanced MAE,MBE, RMSE, machine learning and solar radiation (GSR) to solar energy systems 3 h ahead n Pearson
methodologies, data Kolmogrov, forecast deep learning using machine learning and solar radiation coefficient, random
processing techniques, skill-Smimov test techniques for smart and deep learning forecasting. Error: forest, mutual
and emerging integral buildings and smart models across multiple RMSE and R2 information, and relief;
technologies such as grids geographic locations errors: RMSE, R2 and
5 G and AI. MAE, MAPE
Approach 128 algorithms used. Solar irradiance such Random forest (RF), Machine learning: PR, Models: multivariate Models: support vector

AIP Advances 15, 055201 (2025); doi: 10.1063/5.0237246


Categorized into four as numerical weather support vector SVR, RF deep learning time series model using regression, extreme
groups, namely, prediction methods, regression (SVR), models: ANN, CNN, recursive feature gradient boosting,
statistical, physical, satellite-based extreme gradient and RNN elimination with a categorical boosting,
hybrid, and others with approaches, boosting (XGBoost), decision tree, Pearson and voting-average
relevant application cloud-image based multilayer perceptron correlation, logistic
conditions and features methodologies, (MLP), long regression, gradient
data-driven methods, short-term memory boosting models, and a
and (LSTM), and temporal random forest
sensor-network-based convolutional network
approaches (Conv-1D)
Limitations This paper highlights This paper lacks This study’s findings This study focuses only This study introduces a This study is limited by
trends in solar novelty, misses are limited to a specific on hourly forecasts and feature selection empirically selected
forecasting, such as AI-driven advances, context and may not lacks evaluation of framework for solar features, reliance on
hybrid methods, and overlooks practical generalize across daily or monthly forecasting but relies Brazilian data, and
spatial averaging, and implementation. It different power averages. It also does on high-quality data, high dataset
probabilistic highlights biases in profiles. The models not consider hybrid or lacks generalizability, dependency. Future
forecasting. It calls for remote-sensed data used a narrow set of optimized AI models. and struggles with work should refine
more economic without solutions and hyperparameters, Future research should variability. It faces feature selection, test
analysis, better data lacks detailed method potentially limiting include these models trade-offs in accuracy with varied datasets,
ARTICLE

pre-processing, and the comparisons, broader effectiveness. and broader time scales vs interpretability, and adapt the
use of emerging uncertainty analysis, Future work should for improved solar requires high approach for broader
technologies such as and real-world validate the methods in radiation forecasting computational applications such as
5 G and AI. However, integration insights real-time systems and resources, and needs wind speed and load
it lacks concrete investigate advanced further validation and forecasting
solutions or architectures for better benchmarking
methodologies to forecasting and
address these integration
challenges
[Link]/aip/adv

15, 055201-5
26 December 2025 13:17:38
AIP Advances ARTICLE [Link]/aip/adv

26 December 2025 13:17:38


FIG. 2. Location of the proposed model.

is gathered by a surface per unit area and travels omnidirection- ical variables. It highlights the challenges of traditional regression
ally to the Earth’s surface through atmospheric particle scattering. models, such as overfitting and optimism bias, which occur when
It is utilized in GHI redundancy estimates and permanent PV sys- models are overly complex and over-estimate their explanatory
tems.32 DHI is the shadow-free solar energy that is gathered by a power. To address these issues, regularized regression techniques are
surface per unit area and travels omnidirectionally to the Earth’s employed. This study evaluates ten different ML models and three
surface through atmospheric particle scattering. It is utilized in DL models, selected based on their effectiveness for the task at hand.
GHI redundancy estimates and permanent PV systems. A revolv- The aim of these ML algorithms is to develop predictive models
ing shadow-band irradiometer and a pyranometer mounted in a sun that accurately estimate specific types of data. This process requires
tracker are used to measure it.33 a large dataset to help the algorithms understand system behav-
ior.34 The ML workflow includes several phases: data acquisition,
data cleaning, and data segregation. The collected data are divided
D. Pre-processing into training, testing, and blind sets. The models are trained on the
Imputation is a technique used by the forecasting model to fill training set, evaluated and optimized on the testing set, and finally
in missing values in time series and adjust model hyperparameters. validated on the blind set. The research employs a technique called
80% of the data from 1998 to 2018 is covered by the training set and endogenous forecasting, which utilizes ML-based time series mod-
10% is covered by the test set from 2019 to 2022. Due to data par- els and DL. This approach uses previously recorded solar irradiance
titioning based on the training and testing sets for the model, the data as input parameters for forecasting future values. The ten fore-
validation and training sets are not continuous. This method over- casting models examined in this study, along with the methodologies
estimates the performance of the model while reducing prediction employed, are detailed in the sub sections.
error.

A. Machine learning (ML)


III. COLLECTIVE APPROACHES TO FORECASTING Several ML algorithms’ performance is assessed using the sug-
SOLAR IRRADIANCE gested approach to forecast solar radiation. Briefly stated here are the
This section discusses the use of various ensemble ML and LR, lasso, ridge, elastic net, random forest, extra tress, XGB, LGBM,
DL models for forecasting solar irradiance and other meteorolog- gradient boosting, and KNN regressor algorithms utilized.

AIP Advances 15, 055201 (2025); doi: 10.1063/5.0237246 15, 055201-6


© Author(s) 2025
AIP Advances ARTICLE [Link]/aip/adv

1. Linear regressor to shrink some components of the solution in ̂ β(lasso) to zero for
LR is the most basic and often used type of regression algo- appropriately chosen values of λ, while simultaneously regulariz-
rithm. Through the use of linear predictor functions, linear regres- ing the least squares fit. In comparison with the well-known LARS
sion is used to display the correlation between the input and output method,37 the cyclical coordinate descent approach computes the
variables. The unknown model parameters are estimated using the complete lasso solution routes for λ for the lasso estimator more
least squares method using the given data. Either a set of linear equa- quickly and efficiently. The lasso is a visually beautiful and widely
tions may be solved or iterative techniques such as gradient descent used variable selection technique because of these qualities.
can be used to estimate the values of the parameters. With a probability that tends to 1, the subset of real parameters
LR, a supervised ML method, generates continuous output pre- with zero coefficients may be estimated using an oracle approach
dictions with a consistent slope. Instead of categorizing data, it as precisely zero, just as if the subset model itself were known
focuses on predicting values within a continuous range, such as in advance. In addition, the nonzero coefficients are estimated by
prices or sales. Early studies on time series predominantly operated using an Oracle estimator in a normally distributed and asymptot-
under the assumption of a deterministic world, particularly in the ically unbiased manner and selects variables in an asymptotically
19th century. Yule proposed in 1927 that any time series may be consistent and efficient manner. The super-efficiency phenomena
seen as a manifestation of a stochastic process, thereby introduc- are strongly associated with the oracle characteristic.17 In addi-
ing stochasticity to time series analysis. This basic idea has since tion to possessing the oracle quality, optimal estimators also need
served as the foundation for the development of other time series to meet several crucial additional regularity requirements, namely,
approaches. The World’s decomposition theorem made it easier to continuous shrinkage. The lasso lacks the oracle characteristic, even
formulate and solve linear forecasting issues.35 yet it estimates the larger nonzero coefficients with asymptotically
Time series analysis has since produced a large body of lit- non-ignorable bias and can only pick variables correctly provided
erature covering a wide range of subjects, including identification, the predictor matrix (or the design matrix) fits a fairly rigorous
forecasting, model validation, and parameter estimation. As shown condition.
in Eq. (2), in this study, the equation for LR demonstrates how one or
more independent variables are correlated linearly and the outcome 3. Ridge regressor
of the dependent variable that is numerical, A kind of linear regression called ridge regression penalizes big
coefficients in order to keep the loss function from overfitting. Ridge
y = α + βx,

26 December 2025 13:17:38


(2) regression is the best choice when many predictors are present, all
where β represents the slope of the line and α denotes the y-intercept of which are drawn from a normal distribution and have nonzero
in the linear relationship between y and x. coefficients. It performs best when there are several predictors, each
with a little influence. In addition, it prevents poorly specified and
2. Lasso regressor highly unpredictable coefficients in linear regression models with a
A penalty term is added to the ordinary least squares objective large number of linked variables. The coefficients of connected pre-
function in the Lasso regressor, also known as the Least Absolute dictors are uniformly shrunk toward zero via Ridge regression. The
Shrinkage and Selection Operator. It is a type of linear regres- reason Ridge regression does not cause coefficients to disappear, it is
sion model. The alpha parameter governs this penalty term, which unable to choose a model that includes only the most relevant and
promotes sparsity in the model by reducing or eliminating some accurate subset of predictors as in Eq. (4),
coefficients completely. Consequently, Lasso regression not only
̂
β(ridge) = arg min∥y − Xβ∥22 + λ∥β∥22. (4)
predicts target variables but also performs feature selection, making
it particularly useful when dealing with high-dimensional datasets
n
where identifying relevant features is crucial. Through its regu- The i-th row of X is represented by yit ; ∥β22 ∥ = ∑ β2j is the
larization mechanism, Lasso regression balances between model j=1
complexity and generalization performance, making it a versatile ℓ2-norm penalty on b, represented by β2j ; and the tuning (penalty,
tool widely used in various fields such as economics, genetics, and regularization, or complexity) parameter is λ ≥ 0. This parameter
signal processing.36 controls the intensity of the punishment (linear shrinkage) by weigh-
Lasso regression techniques are commonly used in genomics ing the respective contributions of the penalty term and the data-
due to large datasets and quick algorithms. Though, it can break dependent empirical error. The quantity of shrinking increases with
down when predictors are identical and are not resistant to large cor- a bigger value of λ. Since the value of λ depends on the data, data-
relations. In addition, randomly select one predictor and disregard driven techniques such as cross-validation may be used to find
others. The lasso penalty only expects a small portion of coefficients it.36 Although the phenotypes are centered around their means, the
to be higher, with the majority near zero. For sparse optimization, intercept in (4) is taken to be zero.
the lasso estimator employs a penalized least squares criterion as
shown in Eq. (3), 4. Elastic net regressor
̂
β(lasso) = arg min1 ∥y − Xβ∥22 + λ∥β∥. (3) The elastic net (ENET) is an extension of the lasso that is resis-
tant to very high correlations between the predictors.18 The ENET
When the tuning parameter is λ ≥ 0, sparsity in the solution is was introduced for analyzing large dimensional data in order to pre-
p
induced by the ℓ1 penalty in ∥β∥ = ∑ ∣βi ∣ which allows the Lasso vent the lasso solution paths from becoming unstable in the presence
i=1 of strongly correlated predictors. The ENET may be expressed as in

AIP Advances 15, 055201 (2025); doi: 10.1063/5.0237246 15, 055201-7


© Author(s) 2025
AIP Advances ARTICLE [Link]/aip/adv

Eq. (5) and employs a combination of the ℓ1 (lasso) and ℓ2 (ridge collection of weak prediction models—usually decision trees—to
regression) penalties: produce a more reliable model when generating predictions. If a
GBR has Q trees, where γT represents the scaling factor and hT
̂
β(enet) = arg min1 ∥y − Xβ∥22 , denotes the weak learner, then the following prediction equation is
(5)
subject to Pα (β) = (1 − α)∣β∣ + λ∥β∥22 ≤ s, provided in Eq. (7):7
Q
where Pα (β) is the ENET penalty; α = 1, the ENET simulates basic f Q (P j ) = ∑ hT γT (x j ). (7)
ridge regression; and when α = 0, it simulates the lasso. Automatic T
variable selection is carried out by the ENET’s ℓ1 portion, whereas
grouped selection is encouraged and solution routes are stabilized in 7. XGBoost
relation to random sampling by the ℓ2 portion, which enhances pre- XGBoost, also known as Extreme Gradient Boosting, is a scal-
diction. The ENET may choose groups of correlated features when able ML system specifically designed for tree boosting. It efficiently
the groups are unknown in advance by inducing a grouping effect handles distributed gradient boosting, quickly determining the sig-
during variable selection, which increases the likelihood that a set of nificance of all input characteristics. This approach has been refined
highly connected variables would have comparable magnitude coef- and enhanced by subsequent re searchers. By consolidating mul-
ficients. In contrast to the lasso, the elastic net chooses more than tiple low-accuracy prediction models into a single high-accuracy
n variables when p >> n. The elastic net does not, yet, possess the model, boosting improves ML models. The XGBoost model excels
oracle quality.37 at achieving acceptable prediction accuracy even with extensive
5. Random forest (RF) regressor or complex datasets, potentially requiring fewer iterations or rep-
etitions to achieve the desired accuracy level compared to other
RF regression mixes numerous decision trees to categorize or
approaches. Overall, XGBoost is recognized as an efficient and reli-
predict variable values using bagging to create varied subsets of
able gradient-boosting machine approach and is expressed as in
training data. This boosts tree variety and captures correlations,
Eq. (8),39
making RF successful for both regression and classification prob-
lems, especially with many variables relative to data. RF incorporates N
(T−1)
randomized decision trees and averages their predictions, making it obj (XGb)(T) = ∑ l(y j , ŷ j + f T (x j ) + Ω( f T ) + constant, (8)
j=1
ideal for large-scale issues. It is adaptive and gives valuable feature

26 December 2025 13:17:38


significance metrics. RF is resilient against overfitting, handles miss- where xj is the feature vector, N is the sample size, yj is the observed
ing variables effectively, and resists noise and outliers. Thus, it can (T−1)
have interpretability concerns and significant processing costs with values, ŷ j is the predicted value of the previous iteration, fT is a
huge datasets. Despite these disadvantages, RF remains a popular new function that the model learns, and Ω( f T ) is the regularization
ML method due to its balance of use and performance.38 term that protects the model against complexity. The loss function,
In addition, RF’s adaptability makes it possible for it to pro- represented by the letter l, determines the output of the new tree by
cess both continuous and categorical data, expanding the range of computing the discrepancy between the label and the phase-prior
scenarios in which it can be used. In this scenario, the training set prediction.
is denoted by p, and RF generates S regression trees. Thus, Q(p) is 8. Light gradient boosting machine
created using an S number tree. The forecasting formula is given by
Eq. (6) The Decision Tree Algorithm is used by the Light Gradient
Boosting Machine (LGBM), a gradient-boosting construction, for
1 S
f′
S ranking and classification tasks. It uses ensemble learning to create a
(p) = ∑ Q(p). (6)
rf S S=1 strong single learner by combining several weak learners. LGBM’s
fast processing speed makes it perfect for managing big datasets.
6. Gradient boosting regressor In terms of accuracy and time efficiency, it performs better than
XGBoost. LGBM uses Decision Trees to split nodes and identify
A supervised regression problem is solar irradiance forecasting, the nodes with the best information gain by applying leaf-splitting
and a ML method for numerical value prediction is called gradient- techniques. Features are recorded in a histogram and binned to save
boosting regression (GBR). GBR is an ensemble technique that computation and storage expenses. LightGBM algorithm optimiza-
builds a more robust and complicated model by combining many tion entails modifying important parameters, including learning
basic regression models. Iteratively adding new models while fixing rate, minimum data in a leaf, feature fractions, and number of leaves
mistakes made by previous models is how it operates. The discrep- per trees. Equation (9) shows the weight function measuring the
ancy between predicted and actual values is determined by the loss quality of tree structure q(x).40 The objective function is eventually
function, which must be minimized. The parameters of the regres- obtained by integrating, where I l and I r are samples of the left and
sion models are updated using gradient descent in each iteration. right branches, respectively,
In addition, GBR can perform both linear and non-linear regres-
⎛ ′ ′ ′
2 2 2
sion tasks and work with a range of regression techniques, such as ⎞
neural networks, decision trees, and linear regression, in addition to ⎜ ( ∑ gi) ( ∑ gi) (∑ g i ) ⎟
1 ⎜ iεI l iεI r iεI ⎟
managing missing data and outliers in training data. On the other G= ⎜ ′ + ′ + ′ ⎟. (9)
2⎜⎜ ∑ hi + λ ∑ hi + λ ∑ hi + λ ⎟

hand, the approach may require substantial computing resources
due to its computational demands. This ML model “boosts” a ⎝ iεI l iεI r iε ⎠

AIP Advances 15, 055201 (2025); doi: 10.1063/5.0237246 15, 055201-8


© Author(s) 2025
AIP Advances ARTICLE [Link]/aip/adv

9. K-Nearest neighbor algorithm (KNN) where xT represents the input at time (T), hT is the hidden neu-
A well-liked supervised learning technique for problems ron at time t, U is the weight matrix for the hidden layer at time t,
involving regression and classification is the KNN algorithm. By pre- and W represents the transition weights between hidden layers. The
dicting the relationship between previously unknown data points RNN merges the current input and the previous hidden state, using
and the dataset, it places newly discovered data in the most related the tanh function to process this information. The result is a brand-
category. The carefully chosen “k” value, which is established by new hidden state that stores information from previous inputs and
repeatedly running the algorithm with various “k” values and choos- functions as a memory.
ing the configuration that performs best on the training dataset, Gradient disappearance and explosion are common challenges
has a direct impact on how successful the method is. A distance for RNNs during training. Gradient explosion may be solved by
function determines the majority decision of nearest neighbors in halting backpropagation at a certain moment, yet, since not all
the classification process, which is based on the Pearson correla- weights are updated, this yields subpar results. The vanishing gra-
tion analysis between meteorological variables. Various variations dient problem can be lessened with the help of proper weight
can employ distinct distance functions, including the Minkowsky, initialization.38
Euclidean, and Mahalanobis distances. The Euclidean distance func- 2. Long-short term memory (LSTM)
tion was selected for the estimate procedure in this investigation.
The vanishing gradient and gradient explosion problems that
The KNN method, shown in Eq. (10), is renowned for being straight-
cause long-term data dependencies and time-series forecasting are
forward, simple to use, and flexible enough to adjust to changes in
addressed by the LSTM network, an enhanced version of the RNN.
the data. Its drawbacks include noisy and high-dimensional data
LSTMs are widely utilized in power systems for forecasting renew-
affecting performance, susceptibility to outliers, and the compu-
able power generation, load, and demand response. In order to
tational expense of forecasting new instances, particularly for big
overcome the drawbacks of a diminishing gradient, the LTSM net-
datasets where calculating distances to all data points is necessary,38
work concept was presented. Information flow is controlled by the
√ inputs, outputs, forget gates, and memory cells that make up the
y
aij = (axij )2 + (aij )2 , (10)
network. While the input gate updates cells and decides their next
where the axial distances in the x- and y-axis directions between the concealed state, the forget gate divides data between deleted and
y
centroids of departments i and j are denoted by (axij )2 and (aij )2 , preserved data. The activation function of the gates is the sigmoid
function, which produces values between 0 and 1 to allow infor-

26 December 2025 13:17:38


respectively.
mation to travel through them only in certain combinations. When
10. Extra trees regressor the structure is open, all information can pass through it; when it
While there are some noteworthy distinctions, the extra trees is closed, it acts as a gate. Equation (13) describes the information
regressor is an ensemble method that is comparable with random distribution and mathematical representation.26 The output gates
forests (RFs). In contrast to RFs, extra trees regressor builds decision control a cell’s output and combine it with a cell state that is trig-
trees by employing the entire dataset rather than just bootstrapping gered by the tanh function to produce the final output, ht, which is
samples and choosing splits completely at random. By increasing represented in Eqs. (14) and (15),

CT = IT ∗ CT′ + f T ∗ CT−1 ,
the variety among the trees, this method may improve the model’s
(13)
generalization and robustness when compared to conventional RFs.
The algorithm, shown in Eq. (11), builds B decision trees using
dataset D. It selects split points randomly for each tree split and then OT = σ(w0 ⋅ hT−1 , w0 ⋅ xT + b0 ), (14)
chooses the best split point from these random selections, instead of
searching for the optimal split point,
hT = tan h(CT ).OT . (15)
Eb (y) = ag ⋅ [{t(x; θy,i )∣i = 1, 2, . . . , n}]. (11)
Here, t(x; θy,i ) indicates the ith tree in the yth ensemble, which is 3. Gated recurrent unit (GRU)
specified by θ. The aggregation method used is usually the average Compared to LSTM, GRU is a less complex RNN that offers
for regression tasks.12 greater computation and learning efficiency and simplicity. Long-
term dependencies can be recalled and captured by using both
B. Deep learning models, but because GRU contains fewer features, it can compute
1. Recurrent neural network more quickly and has a lower complexity. It can solve vanish-
An artificial neural network type called a recurrent neural ing gradients by efficiently learning long-term dependency data.
network (RNN) is particularly good at training on sequential or Because of the functional mechanism and design similarities, GRU
time-series data, which contains temporal information that ordi- is regarded as an LSTM variation. Although the gate mechanism is
nary neural networks cannot detect. Sequence data are divided into used by both GRU and LSTM to govern information flow, GRU only
components using RNNs, which also preserve a state that allows the has two gates: the reset gate (rT ) and the update gate (ZT ). These
data to be represented at various times. As shown in Eq. (12), RNN gates establish what data should be removed and kept for later use,
include inputs, hidden neurons, and an activation function, accordingly.26 GRU models require a lot of training and underfitting
because of their poor learning efficiency and delayed convergence.
hT = tan h (U ⋅ xT + W ⋅ hT − 1), (12) Even with their advantages, GRUs might have low learning efficiency

AIP Advances 15, 055201 (2025); doi: 10.1063/5.0237246 15, 055201-9


© Author(s) 2025
AIP Advances ARTICLE [Link]/aip/adv

and delayed convergence, which can result in long training times 2. Coefficient of determination (R2)
and possible underfitting. As shown in Eqs. (16)–(19), the reset gate As expressed in Eq. (21), high R2 value indicates good model
is represented by RT , the update gate by zt, the memory component performance,41
by AT , the activation function by tanh, and the final memory by hT
at the current time step (T),41 N
∑ (xi − yi )
2

RT = σ(Ur hT−1 + br + WR XT ), (16) R2 = 1 − N


i=1
. (21)
∑ (xi − yi,avg )
2
i=1
ZT = σ(bZ WZ ), (17)
3. Mean absolute error (MAE)
Expressed in Eq. (22), Forecast performance evaluation in the
AT = tan h(Wh ⋅ XT + Uh bh + (rT ∗ hT−1 )), (18)
renewable energy business and regression problems frequently use
the MAE, a global error metric used to assess uniform forecast
mistakes. Higher projections are shown by smaller values, which
hT = (1 − ZT ) ∗ hT−1 + ZT ∗ AT . (19)
indicate the total forecast accuracy,41
N
IV. MODEL EVALUATION ∑ ∣xi − yi ∣
MAE = i=1
. (22)
A. Statistical metrics and analysis N
To accurately assess the expected accuracy of the models that
were previously described, three statistical quality criteria were used. 4. Mean absolute percentage error (MAPE)
These formulas serve as the basis for calculating these measure- It is the segmentation of MAE according to capacity or demand.
ments. Implementing these measures was facilitated by Scikit-learn, In addition, an average of absolute deviations of percentage errors is
a renowned Python library celebrated for its capacity to run multi- shown as in Eq. (23),42
ple ML models. Scikit-learn, an open-source toolkit, is specifically

26 December 2025 13:17:38


N
designed to furnish advanced capabilities for tasks involving pre-
∑ (xi − yi )
2
dictive data analysis. The functionalities inherent within Scikit-learn
MAPE = i=1
× 100%. (23)
enable accurate and efficient data processing.42 In addition, the para- N
meters of the algorithms can be fine-tuned to optimize efficiency
further. Solar irradiance forecasts are assessed using error measures 5. Uncertainity
such as R2, MAE, MAPE and RMSE; the error analysis shown in
To evaluate the models’ prediction uncertainty, this study used
Table III and Fig. 2 respectively. A number of data points is denoted
quantile regression (QR). This approach is particularly helpful since
by N; the mean value of the computed variable is denoted by yi,avg ;
it offers information about the whole spectrum of prediction uncer-
the actual value at time step i is denoted by yi ; and the value that the
tainty and makes distributional errors visible. The quantiles of
forecasting model approximates after that is indicated by xi . These
conditional functions are computed using a linear statistical method
are common error metrics.
known as QR analysis. The sequence of the ML models is shown
1. Pearson’s correlation coefficient (R2) in Tables IV and V The chance that an input pattern’s target falls
inside the prediction bounds is known as the PICP, and it is cal-
Selecting features using filters is a crucial stage in creating a ML
culated using the matching frequency in the manner described in
algorithm model. It was applied to determine which input dataset
the following. It is also possible to add one more performance met-
columns had the highest predictive potential. The filter selection
ric for the forecast limit. This is the test dataset’s MPI, computed
metric in this study was the Pearson correlation.40 The linear cor-
across all points. It may be calculated by estimating the capacity to
relation coefficient, or Pearson correlation coefficient, between two
encapsulate target values inside the forecast boundaries. It offers a
variables, X and Y, is measured using Pearson correlation. It is calcu-
substantial comprehension of plausible causal relationships over the
lated as in Eq. (20), by dividing the two variables’ covariance by the
whole dataset. Utilizing the uncertainty measures for MPI and PICP,
sum of their standard deviations. The Pearson correlation coefficient
initially introduced by Ref. 28, the efficacy of the QR method was
is unaffected by modifications to the values of the two variables,
evaluated. The PICP and MPI formulas are shown in Eqs. (24) and
n (25): in this case, yi,avg represents the upper and lower prediction
′ ′
∑ (Xi − X )(Yi − Y ) upper
limits, which are denoted, respectively, by PLt and PLlower
t ,
rXY = ¿ i=1
√ , (20)
Án n ⎧
Á ′
À ∑ (Xi − X )2 ′
∑ (Yi − Y )
2
1 N ⎪ 1, PLt PLlower
⎪ t
upper
< yi,avg < PLt ,
PICP = ∑ U, U ⎨ (24)
i=1 i=1
N t=1 ⎪

⎩ 0, Otherwise,
where the correlation function is represented by n, while the covari-
ance is denoted by rXY . The sample mean is Y ′ , and the individual 1 N upper
MPI = ∑ {PLt − PLlower
t . (25)
sample points are X i and Y i , which are indexed by i. N N=1

AIP Advances 15, 055201 (2025); doi: 10.1063/5.0237246 15, 055201-10


© Author(s) 2025
AIP Advances ARTICLE [Link]/aip/adv

V. PREDICTED RESULTS AND DISCUSSION assumed by Pearson correlation, the association may not be effec-
tively represented in directions. Subsequently, the following sections
A. Sensitivity test
will elucidate the ongoing research efforts. The Pearson correla-
One of the most important steps in using ML models is select- tion matrix shown in Fig. 3 indicates that temperature has a strong
ing features using filters. It assists in determining which input positive correlation with solar irradiance, suggesting that higher
dataset columns have the highest predictive potential. The filter temperatures are associated 10 with higher levels of solar radiation.
selection parameter in this investigation was the Pearson correlation. Conversely, relative humidity shows a strong negative correlation
The accuracy of a model or technique usually hinges on how care- with solar irradiance, meaning that higher humidity is linked to
fully suitable input parameters are chosen. To do this, a Pearson cor- lower solar radiation. Other variables, such as dew point and wind
relation analysis was carried out between the output variable—solar speed, have weaker correlations with solar irradiance, with dew
radiance—and the meteorological input variables—dew point, tem- point showing a slight positive relationship and wind speed show-
perature, relative humidity, and wind speed. Pearson correlation, ing a moderate negative relationship. These insights are useful for
often known as Pearson’s or the Pearson correlation coefficient, is predicting solar irradiance based on meteorological conditions.
a measurement of the linear relationship between two continuous Based on the Fig. 3 matrix, temperature is strongly positively
variables.43 correlated with solar irradiance (0.74), making it a strong predic-
The strength and direction of the linear relationship between tor. Relative humidity has a strong negative correlation (−0.82),
the variables are both measured. A Pearson correlation value that indicating that it might also be a significant predictor but with an
ranges from −1 to 1 shows a weak or non-existent linear link; a inverse relationship. These variables can be prioritized when build-
positive correlation indicates a perfect positive linear relationship; ing predictive models due to their strong associations with the target
a negative correlation indicates a perfect negative linear relation- variable. Wind speed shows a moderate negative correlation with
ship. The degree of the linear relationship between the variables is solar irradiance (0.51), which might be useful but not as strong as
indicated by the correlation coefficient’s proximity to +1 or −1; a temperature or relative humidity. Dew point and wind speed have
close correlation of 0 indicates a weak or non-existent linear rela- low correlations, suggesting both might be less useful predictors on
tionship. If the connection between the variables is not linear, as their own.

26 December 2025 13:17:38

FIG. 3. Pearson correlation plot for different variables.

AIP Advances 15, 055201 (2025); doi: 10.1063/5.0237246 15, 055201-11


© Author(s) 2025
AIP Advances ARTICLE [Link]/aip/adv

B. Performance evaluation models Similarly, the figures in the following represent the final month’s
in machine learning record for the year 2022.
Every experiment has been conducted on a single platform. The baseline models serve as a benchmark to assess how
Several experimental results were described in the current section. machine learning techniques can enhance prediction accuracy.
Training has made use of historical data from NREL over the last Figures 4–8 illustrate the comparison between the baseline and
six years. With considerable caution, the tests were conducted using actual values for dew point, temperature, relative humidity, wind
Python 3.0 and the Scikit learn (sklearn) module to use machine speed, and solar irradiance, respectively. This baseline approach
learning methods. The Lenovo workstation, which included an Intel helps identify the most and least accurate models among the ten
Core i5 1235U CPU with 10 cores and 16 GB RAM, was used for the machine learning models. Several criteria influence which model
experiment. 80% of the data from 1998 to 2018 are covered by the is best: overfitting, complexity, interpretability, and performance
training set and 10% are covered by the test set from 2019 to 2022. measures such as RMSE, MAE, MAPE, and R-squared on the test

26 December 2025 13:17:38


FIG. 4. Comparison of baseline prediction and real value for dew point.

FIG. 5. Comparison of baseline prediction and real value for temperature.

AIP Advances 15, 055201 (2025); doi: 10.1063/5.0237246 15, 055201-12


© Author(s) 2025
AIP Advances ARTICLE [Link]/aip/adv

FIG. 6. Comparison of baseline prediction and real value for relative humidity.

26 December 2025 13:17:38


FIG. 7. Comparison of baseline prediction and real value for wind speed.

FIG. 8. Comparison of baseline prediction and real value for solar irradiance.

AIP Advances 15, 055201 (2025); doi: 10.1063/5.0237246 15, 055201-13


© Author(s) 2025
AIP Advances ARTICLE [Link]/aip/adv

set. Higher R-squared values imply better fit, whereas performance as random forest, extra trees, XGB regressor, and LGBM regressor
measurements such as RMSE, MAE, and MAPE suggest greater also have low RMSE and MAE. The choice depends on prioritizing
performance.44 When a model outperforms the test set on the train- simplicity or predictive performance. Validating the chosen model’s
ing set, it is said to be overfitting. Another thing to think about performance is crucial before making a final decision. Table II sum-
is complexity; simpler models, such lasso or linear regression, are marizes a comparison of the current machine-learning methods with
favored over more intricate models, such as random forest or gra- other approaches in the solar irradiance forecast. Based on the low-
dient boosting. Finally, interpretability is a criterion that should be est errors (RMSE, MAE, and MAPE) and the greatest R2 values,
taken into account. Elastic net and gradient boosting regressor are the best-performing models are identified for each target variable
potential models for predicting data and has a low RMSE and MAE illustrated in Figs. 5–7 Figs. 4–8. The best model for solar irradi-
on the test set, indicating a good predictive performance. GBR has ance is gradient boosting, which has the highest R2 (0.607) and the
the lowest RMSE but higher MAE and MAPE. Other models such lowest RMSE (0.685), MAE (0.530), and MAPE (0.0217). Gradient

TABLE II. Performance metrics for different models.

Model Target variable RMSE test MAE test MAPE test R2 test

Solar irradiance (W/m2 ) 16.4929 13.9615 2.98 × 1016 0


Dew point (○ C) 1.0940 0.9072 0.036 99 −2.74 × 10−5
Baseline Temperature (○ C) 1.9678 1.6987 0.066 64 −0.01
Relative humidity (%) 7.8740 6.7837 0.075 22 −0.02
Wind speed (degrees) 0.1691 0.1450 3 × 1014 −0

Solar irradiance (W/m2 ) 0.9351 0.7559 0.030 88 0.2692


Dew point (○ C) 1.5664 1.3194 0.052 03 0.3592
Linear regression Temperature (○ C) 7.3180 5.9459 0.065 03 0.1207
2.81 × 1014

26 December 2025 13:17:38


Relative humidity (%) 0.1564 0.1336 0.1428
Wind speed (degrees) 14.9731 11.3857 2.06 × 1016 0.1728

Solar irradiance (W/m2 ) 0.9562 0.7789 0.031 85 0.2359


Dew point (○ C) 1.5365 1.2873 0.050 50 0.3834
Lasso Temperature (○ C) 7.2727 5.9220 0.064 86 0.1315
Relative humidity (%) 0.1691 0.1450 3.01 × 1014 −0.0011
Wind speed (degrees) 14.9522 11.3536 2.04 × 1016 0.1751

Solar irradiance (W/m2 ) 0.9351 0.7559 0.030 88 0.2692


Dew point (○ C) 1.5664 1.3194 0.052 03 0.3592
Ridge Temperature (○ C) 7.3180 5.9459 0.065 03 0.1207
Relative humidity (%) 0.1564 0.1336 2.81 × 1014 0.1428
Wind speed (degrees) 14.9731 11.3857 2.06 × 1016 0.1728

Solar irradiance (W/m2 ) 0.9438 0.7646 0.031 29 0.2556


Dew point (○ C) 1.5432 1.2962 0.050 93 0.3781
ElasticNet Temperature (○ C) 7.3063 5.9677 0.065 27 0.1235
Relative humidity (%) 0.1691 0.1450 3.01 × 1014 −0.0011
Wind speed (degrees) 14.9707 11.4442 2.09 × 1016 0.1731

Solar irradiance (W/m2 ) 0.7950 0.5751 0.023 67 0.4719


Dew point (○ C) 0.9826 0.7650 0.029 58 0.7478
Random forest Temperature (○ C) 5.0070 2.8866 0.033 42 0.5884
Relative humidity (%) 0.1478 0.0916 1.50 × 1014 0.2348
Wind speed (degrees) 8.4070 4.3857 0.4816 0.7392

Solar irradiance (W/m2 ) 0.6855 0.5300 0.021 71 0.6074


Dew point (○ C) 0.8249 0.6683 0.025 83 0.8223
Gradient boosting Temperature (○ C) 4.1217 2.5615 0.029 26 0.7211
Relative humidity (%) 0.1268 0.0893 1.86 × 1014 0.4372
Wind speed (degrees) 7.9225 4.2300 4.74 × 1014 0.7684

AIP Advances 15, 055201 (2025); doi: 10.1063/5.0237246 15, 055201-14


© Author(s) 2025
AIP Advances ARTICLE [Link]/aip/adv

TABLE II. (Continued.)

Model Target variable RMSE test MAE test MAPE test R2 test

Solar irradiance (W/m2 ) 0.8373 0.6054 0.024 96 0.4142


Dew point (○ C) 0.9625 0.7487 0.029 05 0.7581
Extra trees Temperature (○ C) 5.4635 3.1183 0.036 03 0.5099
Relative humidity (%) 0.1450 0.0893 1.33 × 1014 0.2635
Wind speed (degrees) 9.4835 4.8260 0.5010 0.6682

Solar irradiance (W/m2 ) 0.8502 0.6184 0.025 47 0.3959


Dew point (○ C) 1.0434 0.8430 0.032 75 0.7157
XGBoost Temperature (○ C) 4.8657 3.2052 0.036 62 0.6113
Relative humidity (%) 0.1350 0.0897 1.61 × 1014 0.3619
Wind speed (degrees) 8.2628 4.7973 1.61 × 1015 0.7481

Solar irradiance (W/m2 ) 0.7137 0.5264 0.021 76 0.5744


Dew point (○ C) 0.8587 0.6655 0.025 88 0.8074
LightGBM Temperature (○ C) 4.2436 2.6402 0.030 33 0.7043
Relative humidity (%) 0.1254 0.0831 1.53 × 1014 0.4491
Wind speed (degrees) 7.8507 4.2117 2.20 × 1014 0.7726

Solar irradiance (W/m2 ) 0.7933 0.5804 0.023 98 0.4740


Dew point (○ C) 0.9127 0.7100 0.027 71 0.7824
KNeighbors Temperature (○ C) 4.4015 2.7496 0.031 53 0.6819
Relative humidity (%) 0.1405 0.0907 1.35 × 1014 0.3089
Wind speed (degrees) 8.4759 4.4312 1.42 × 1014 0.7349

26 December 2025 13:17:38


boosting also performs the best for dew point, having the highest working performance of GBR, which contains short and long peri-
R2 (0.822) and the lowest RMSE (0.825), MAE (0.668), and MAPE ods such as 1, 8, 16 and 24 h. The evaluation of the GBR model
(0.0258). Gradient boosting leads once more in temperature pre- in Table III for short-term (1-h) predictions highlights its balanced
diction, with the lowest RMSE (4.122), MAE (2.561), and MAPE performance. With an RMSE of 0.72, the model’s predictions show
(0.0293), along with the greatest R2 (0.721). LightGBM is the best a moderate level of precision, averaging 0.72 units away from the
model for relative humidity, with the lowest MAE (0.0831), RMSE actual values. The MAE of 0.53 further indicates fair consistency in
(0.125), and robust R2 (0.449). LightGBM excels in wind speed, hav- these predictions, as the average absolute error between predicted
ing the lowest RMSE (7.851), MAE (4.212), and greatest R2 (0.772). and actual values remains relatively low. The MAPE of 2.21% under-
Based on the analysis of test metrics, gradient boosting performs the scores the model’s reliable predictive accuracy, as the percentage
best for the majority of the target variables, particularly for solar irra- deviation from the actual values is minimal. In addition, the R2
diance, dew point, and temperature. This is likely due to its ability to value of 0.57 signifies that the model explains 57% of the variance
handle complex, non-linear relationships and capture interactions in the data, demonstrating a reasonable fit. Overall, these metrics
between features. For relative humidity and wind speed, LightGBM suggest that the GBR model is a reliable option for short-term fore-
shows superior performance, possibly because of its efficiency in casting and long term for providing a good balance of precision and
dealing with large datasets and its capability to model intricate accuracy.
patterns.
Both gradient boosting and LightGBM are ensemble methods C. Performance metrics of DL
that combine multiple weak learners to create a strong learner, which Strong performance is demonstrated by DL algorithms, partic-
can lead to better generalization and robustness in various scenar- ularly when working with temporal and spatial data. Small datasets,
ios. Then, Figs. 10 and 11 provide a figure-based understanding of however, could not provide sufficient or appropriate representative
Table II. GBR can be chosen as the best performer here. In addition, training data, which could lead to overfitting of the model. DL lacks
Fig. 13 depicted the comparison of the overall result actual value and adequate data for trainings, which are insufficient when dealing with
the predicted values, that is, by all the ten different ML methods of extreme weather types. One tactic to improve the performance of
solar irradiance alone. deep learning, and particularly deep networks, is data augmentation.
Following a rigorous learning process, GBR surpasses the other By providing fresh data composed of either old data or new copies
ten models and is currently drifting toward thoroughly learning of existing data that have been modified, this approach increases the
the model—GBR—for solar irradiance prediction, which now takes amount of training data that are accessible. This paper focused on
into account four distinct time horizons. Figure 14 demonstrates the the application of many well-known DL models, including RNN,

AIP Advances 15, 055201 (2025); doi: 10.1063/5.0237246 15, 055201-15


© Author(s) 2025
AIP Advances ARTICLE [Link]/aip/adv

FIG. 9. RMSE values of optimum models for various input combinations in a train and test dataset.

26 December 2025 13:17:38

FIG. 10. MAE values of optimum models for various input combinations in a train and test dataset.

AIP Advances 15, 055201 (2025); doi: 10.1063/5.0237246 15, 055201-16


© Author(s) 2025
AIP Advances ARTICLE [Link]/aip/adv

FIG. 11. MAPE values of optimum models for various input combinations in a train and test dataset.

26 December 2025 13:17:38

FIG. 12. R2 values of optimum models for various input combinations in a train and test dataset.

AIP Advances 15, 055201 (2025); doi: 10.1063/5.0237246 15, 055201-17


© Author(s) 2025
AIP Advances ARTICLE [Link]/aip/adv

26 December 2025 13:17:38


FIG. 13. Solar irradiance performance by various ML models.

FIG. 14. Solar irradiance prediction of gradient boosting.

AIP Advances 15, 055201 (2025); doi: 10.1063/5.0237246 15, 055201-18


© Author(s) 2025
AIP Advances ARTICLE [Link]/aip/adv

TABLE III. Performance of different DL prediction models. Minimal overfitting or underfitting for all models is suggested by the
2 tight alignment of training and validation losses.
MODEL RMSE MAE R Overall, the RNN, LSTM, and GRU models demonstrate sim-
RNN 0.300 610 965 0.266 959 38 0.008 343 347 ilar performance in terms of loss reduction and final loss values,
LSTM 0.302 093 561 0.268 533 459 −0.001 462 363 indicating that any of these models could be suitable for the fore-
GRU 0.301 453 221 0.268 708 013 0.002 778 684 casting task with minor differences in learning stability. Based on
the provided loss curves, the LSTM model appears to be the best
performer. It shows the most stable and consistent validation loss,
indicating better generalization to unseen data compared to the
LSTM, and GRU, to solar irradiance forecasting. In addition, this RNN and GRU models, which exhibit more fluctuations.
study addressed the problem of overfitting and highlighted improve- Thus, the LSTM model is likely the most reliable for the fore-
ments in model training by augmenting the training data. The table casting task. Based on RMSE and MAE, RNN performs the best
further provides details about the models’ architecture, advantages, overall because its values are the lowest for both measures. In addi-
and disadvantages. tion, compared to the other models, RNN has the highest (but still
The RNN has superior prediction accuracy, as evidenced by extremely low) R2 value, meaning it explains a little bit more vari-
its lowest RMSE and MAE. Among the three models, i.e., RNN, ance in the data. As a result, of the three models, RNN performs the
LSTM, and GRU models, the RNN performs best with the low- best in terms of accuracy. Therefore, RNN shows the different time
est RMSE (0.3006) and MAE (0.2670), indicating the most accurate horizons predictions (1, 8, 16, and 24 h), as is depicted in Fig. 16.
predictions. However, its R2 value (0.0083) is very low, showing Although the RNN model achieves low RMSE and MAE values, its
limited ability to explain data variability. The LSTM model has high MAPE and negative or near-zero R2 values suggest poor over-
slightly higher RMSE (0.3021) and MAE (0.2685) and a negative R2 all predictive performance, particularly over longer time horizons
(−0.0015), suggesting poorer performance. The GRU model’s per- Fig. 17.
formance is similar to LSTM but with a slightly positive R2 (0.0028), The model has difficulty providing accurate and reliable predic-
indicating marginally better variance explanation shown in Table III. tions, as indicated by the high percentage errors and low explained
However, all models show limited effectiveness in capturing data variance. The RNN model performs consistently when evaluated
patterns. over a range of time horizons in terms of RMSE and MAE, but its

26 December 2025 13:17:38


Figure 15 shows the training and validation loss curves for three extraordinarily high MAPE and negative or almost zero R2 values
types of neural network models: RNN, LSTM, and GRU, across 50 raise red flags, as observed in Table IV. Over several projection peri-
epochs. In the first few epochs, all three models show a sharp decline ods, the RMSE and MAE values stay consistent at 0.28 and around
in loss, which is then followed by stabilization at a loss value of 0.1. 0.24, respectively, suggesting a rather constant level of inaccuracy.

FIG. 15. Loss curves of DL models.

AIP Advances 15, 055201 (2025); doi: 10.1063/5.0237246 15, 055201-19


© Author(s) 2025
AIP Advances ARTICLE [Link]/aip/adv

FIG. 16. RNN forecast for different time horizons.

26 December 2025 13:17:38


FIG. 17. Hybrid model prediction comparison on solar irradiance.

However, the MAPE values show a large percentage divergence it comes to longer-term projections when it is unable to produce
from the actual values, particularly for the 8 h (933.09%) and 24 h accurate and dependable predictions.
(867.32%) projections, which are noticeably high. This implies that When comparing the best-performing machine learning and
the RNN model has trouble forecasting real values accurately, even deep learning models—GBR and RNN, respectively—Tables II and
if it can maintain a constant error margin. Furthermore, a weak fit IV reveal that each model is suited to different forecasting horizons.
is shown by the negative or almost zero R2 values, which show that The model optimized for short-term forecasts excels in predicting
the model explains little to none of the variation in the data. In sum- the immediate future (1 h ahead) but does not perform as well for
mary, the RNN model performs less well than ideal, especially when longer-term predictions (24 h ahead). For short-term prediction, the

AIP Advances 15, 055201 (2025); doi: 10.1063/5.0237246 15, 055201-20


© Author(s) 2025
AIP Advances ARTICLE [Link]/aip/adv

TABLE IV. Various time prediction performances of RNN error metrics. predictions, while the GBR model is preferable for accurate and
2 stable long-term forecasting.
Time horizon RMSE MAE MAPE (%) R
1h 0.28 0.24 197.06 −0.03
8h 0.28 0.24 933.09 −0.01 D. Performance of the hybrid model
16 h 0.28 0.24 574.25 −0.01 The proposed hybrid GBR-RNN model effectively combines
24 h 0.28 0.22 867.32 0.01 the strengths of GBR and RNN to enhance forecasting accuracy.
GBR excels at capturing complex patterns in structured, non-
linear data and is robust to outliers with fast convergence, while
RNN specializes in modeling sequential data, learning temporal
RNN model has a significantly lower RMSE and MAE for short- dependencies, and capturing short-term patterns.
term predictions compared to the GBR, which suggests that it may By integrating these two approaches, the hybrid model com-
be more effective for immediate (1-h) forecasts. However, the nega- pensates for the limitations of each when used alone. This syn-
tive R2 value and extremely high MAPE indicate that its predictive ergy makes the GBR-RNN model particularly well-suited for time-
accuracy and consistency are poor. For longer-term forecasts (8, 16, series forecasting tasks such as solar irradiance prediction, offer-
and 24 h), the GBR model is superior. It exhibits higher R2 values, ing improved reliability and performance. Figure 18 illustrates the
indicating a better fit to the data, and its MAPE is significantly lower, pseudo-code of the Hybrid GBR-RNN model, detailing the step-
demonstrating greater accuracy in predictions relative to actual val- by-step flow from GBR training and prediction to RNN sequence
ues. Overall, the RNN model may offer more precise short-term modeling and final output generation.

26 December 2025 13:17:38

FIG. 18. Pseudo-code of hybrid GBR-RNN model.

AIP Advances 15, 055201 (2025); doi: 10.1063/5.0237246 15, 055201-21


© Author(s) 2025
AIP Advances ARTICLE [Link]/aip/adv

The hybrid GBR-RNN model is a two-stage forecasting the other two models. The R2 values indicate that both the GBR
approach designed to improve the accuracy of solar irradiance pre- and hybrid models explain 77% of the variance in the data, while
diction. First, GBR is used to capture the complex non-linear trends the RNN explains slightly less at 75%. This shows that GBR and the
in historical data by combining multiple weak learners (usually deci- hybrid model are equally effective in capturing the variability in the
sion trees). The GBR model is trained on input features such as data, while the RNN is slightly less effective.
weather parameters and previous irradiance values. It builds an Therefore the hybrid model performs similarly to the GBR
ensemble of decision trees where each new tree corrects the errors model, with slight advantages in MAE. Both models outperform
of the previous ones. The final prediction from GBR is given by the RNN in terms of RMSE, MAE, and R2 , indicating that these
are more reliable for this particular forecasting task. The hybrid
Q
model’s performance being nearly identical to GBR suggests that
f Q (x j ) = ∑ γT ⋅ hT (x j ), (26)
T=1
while the hybridization does not drastically improve the metrics, it
maintains strong performance across all evaluated aspects. There-
where fQ (xj ) is the predicted value at input xj , Q is the number of fore, the hybrid model offers a balanced approach that combines the
trees, hT is the weak learner at iteration T, and γT is the learning rate strengths of both GBR and RNN, without sacrificing accuracy. They
or scaling factor. improve accuracy and reliability of solar irradiance forecasting.
In the second stage, the RNN takes the sequence of GBR
outputs as its input to learn temporal patterns over time. RNNs E. Uncertainity
maintain a hidden state that is updated at each time step using The performance of SR, which may be a statistical or ML model,
the current input and the previous hidden state. This allows RNNs is evaluated using the Prediction Interval Coverage Probability
to effectively model dependencies across time, which is crucial for (PICP) and the Mean Prediction Interval (MPI).
time-series forecasting such as solar irradiance. The hidden state
update in the RNN is calculated as (a) PICP: measures the likelihood that the actual outcome of an
input falls within the predicted boundaries. It is calculated
hT = tanh (U ⋅ xT + W ⋅ hT−1 ), (27) based on a frequency formula.
(b) MPI: quantifies the ability to capture target values within pre-
where xT is the input at time T, hT−1 is the previous hidden state, diction bounds and is determined for all points in the test
U and W are weight matrices, and tanh is the activation function. data.

26 December 2025 13:17:38


By combining GBR’s trend learning with RNN’s temporal modeling,
the hybrid model improves prediction accuracy and handles both Residuals of the model are used to estimate uncertainty. Both
non-linearity and time-dependence in solar data. training and testing datasets provide PICP and MPI values, which
Extensive testing has revealed that the hybrid model con- are obtained using the QR technique and presented in Table VI,
sistently outperforms other traditional and advanced forecasting showing the solar irradiance performance. The table provides an
methods as Eqs. (26) and (27) show the model. Figure 16 pro- analysis of various weather-related variables—the MPI values indi-
vides a detailed comparison of the performance metrics for the cate the average range within which the predicted values fall, with
hybrid model, GBR, and RNN. The results indicate that the hybrid lower values signifying tighter and more precise predictions. The
approach not only enhances predictive accuracy but also demon- PICP values reflect the probability that the true values lie within the
strates robustness across different scenarios. The improved perfor- prediction intervals, with higher values suggesting greater reliability.
mance metrics underscore the effectiveness of integrating ML and
DL methodologies, validating the potential of hybrid models in
achieving superior forecasting outcomes. TABLE VI. Evaluation of model uncertainty.
Table V shows the RMSE values, which indicate that the GBR
model has the lowest error, closely followed by the hybrid model, Target variable Model MPI PICP
while the RNN model has the highest RMSE. This suggests that
Linear 58.430 607 0.950 268 8
GBR is slightly better at minimizing the overall error in predic-
regression
tions, although the hybrid model is very close in performance. The
Lasso 58.412 471 0.948 924 7
MAE values show that the hybrid model has the lowest mean abso-
Ridge 58.430 607 0.950 268 8
lute error, slightly better than the GBR, while the RNN again has
ElasticNet 58.411 848 0.950 268 8
the highest MAE. The hybrid model’s lower MAE suggests that it
Random forest 32.879 851 0.909 946 2
is more accurate in terms of average prediction error compared to
Solar irradiance (w/m2 ) Gradient 31.000 028 0.913 978 5
boosting
Extra trees 37.389 42 0.909 946 2
TABLE V. Performance of different DL prediction models.
XGBoost 32.390 316 0.907 258 1
MODEL RMSE MAE R2 LightGBM 30.760 028 0.904 569 9
KNN 33.179 569 0.913 978 5
GBR 7.87 4.30 0.77 LSTM 32.935 577 0.897 849 5
RNN 8.16 4.58 0.75 RNN 31.549 742 0.915 322 6
Hybrid 7.89 4.28 0.77 GRU 31.996 078 0.900 537 6

AIP Advances 15, 055201 (2025); doi: 10.1063/5.0237246 15, 055201-22


© Author(s) 2025
AIP Advances ARTICLE [Link]/aip/adv

The linear models, including linear regression, lasso, ridge, the NREL (1988–2022). Focusing on the five input variables—solar
and elastic net, exhibit similar MPI values of ∼58.4. This consis- irradiance, dew point, temperature, relative humidity, and wind
tency indicates that these models generate relatively wide prediction speed—with a primary emphasis on solar irradiance, this study sys-
intervals, suggesting a conservative approach that accounts for a tematically evaluates the predictive capabilities of three DL and ten
significant level of uncertainty in their predictions. ML algorithms.
In addition, the PICP for these models is around 0.95, mean- Through an extensive comparative analysis, this study not only
ing that 95% of the actual observations fall within the prediction identifies the top-performing individual models—GBR, XGBoost,
intervals. This balance demonstrates that the prediction intervals are and elastic net for ML, and RNN for DL—but also demonstrates
neither excessively wide nor too narrow, effectively capturing the how ensemble methods can significantly enhance forecasting accu-
variability present in the data. racy. The incorporation of Pearson correlation for feature selection
In contrast, tree-based models such as random forest, gradi- and the use of quantile regression to account for uncertainty further
ent boosting, extra trees, XGBoost, and LightGBM generally display underscore the robustness and depth of this work.
lower MPI values, ranging from about 30 to 37, with gradient boost- A standout contribution of this research is the development
ing and LightGBM showing particularly low MPI values. These of a hybrid GBR-RNN model, combining the strengths of the best-
lower MPI values suggest that tree-based models produce nar- performing ML and DL models. This hybrid approach consistently
rower prediction intervals, reflecting greater confidence in their outperforms individual models across key performance metrics (R2 ,
predictions. However, the PICP for these models is slightly reduced RMSE, and MAE), particularly for long-term solar irradiance fore-
compared to the linear models, falling between 0.90 and 0.91, with casting, affirming its superior predictive capabilities. The hybrid
LightGBM having the lowest PICP at 0.90. This indicates that while model not only improves accuracy but also enhances the reliability
tree-based models are more confident, they may occasionally under- and efficiency of energy management in hybrid renewable systems.
estimate the uncertainty. This results in fewer actual values being Moreover, this study sets a new benchmark in solar irradiance fore-
captured within the prediction interval. casting, showcasing clear superiority over conventional techniques.
The KNN model has an MPI of 33.18, which signifies a moder- By comparing the proposed hybrid method with four other exist-
ate level of confidence in its predictions. The PICP for KNN is 0.91, ing approaches, it is evident that the developed model significantly
suggesting that the model does a reasonable job of capturing the enhances prediction quality and operational decision-making in
actual observations within its prediction intervals, although there is energy systems.

26 December 2025 13:17:38


still some room for improvement in fully accounting for uncertainty. While the seasonal forecasting capability has not been explored
Neural network models, including LSTM, RNN, and GRU, in this study, it represents a key limitation that future work should
demonstrate relatively low MPI values, around 31–33. This indicates address to better understand model performance across different
that these models are confident in their predictions, producing nar- climatic conditions. In addition, the current analysis is based on a
rower prediction intervals. The PICP for these models ranges from single location’s dataset, which may limit the generalizability of the
0.89 to 0.92. In particular, GRU and RNN exhibit slightly higher results. To build on this work, future research should expand the
PICP values compared to LSTM, while the LSTM model has a PICP scope by evaluating model performance across multiple geographi-
of 0.89. This lower PICP for LSTM suggests that its prediction inter- cal locations and climate zones. Exploring additional meteorological
vals might be too narrow, leading to a higher likelihood of missing inputs and incorporating more advanced ensemble techniques could
actual values. further improve forecasting accuracy. The findings of this study,
Nonetheless, linear models offer a conservative approach with particularly the success of the GBR-RNN hybrid model, provide a
wide prediction intervals and high coverage probability, effectively strong foundation for such advancements. This research is expected
balancing confidence and uncertainty. Tree-based models provide to influence future developments by encouraging the design of
more confident predictions with narrower intervals but may slightly more adaptive, uncertainty-aware, and location-specific forecasting
underestimate uncertainty. KNN strikes a moderate balance, while frameworks that support more intelligent and responsive energy
neural network models are highly confident but risk underestimat- management in hybrid renewable energy systems.
ing uncertainty, particularly the LSTM model. Selecting the appro-
priate model depends on the specific requirements for confidence
and the acceptable level of uncertainty in the predictions. AUTHOR DECLARATIONS
If the goal is to minimize uncertainty while maintaining high Conflict of Interest
coverage, a balance between MPI and PICP should be sought. Mod-
els with high PICP and moderate MPI (such as linear models) offer The authors have no conflicts to disclose.
a conservative approach, while tree-based models and neural net-
works may provide more confident predictions but with a slight risk Author Contributions
of underestimating uncertainty.
C Vanlalchhuanawmi: Conceptualization (equal); Investigation
(equal); Writing – original draft (equal); Writing – review & editing
VI. CONCLUSIONS (equal). Subhasish Deb: Conceptualization (equal); Investigation
The increasing integration of renewable energies into the elec- (equal); Writing – original draft (equal); Writing – review & editing
trical system demands improved forecasting of critical meteorolog- (equal). Md. Minarul Islam: Conceptualization (equal); Investiga-
ical variables. This research presents a novel and superior approach tion (equal); Writing – original draft (equal); Writing – review
to long-term forecasting using meteorological observation data from & editing (equal). Taha Selim Ustun: Conceptualization (equal);

AIP Advances 15, 055201 (2025); doi: 10.1063/5.0237246 15, 055201-23


© Author(s) 2025
AIP Advances ARTICLE [Link]/aip/adv

21
Investigation (equal); Writing – original draft (equal); Writing – K. Sudharshan et al., “Systematic review on impact of different irradi-
review & editing (equal). ance forecasting techniques for solar energy prediction,” Energies 15(17), 6267
(2022).
22
M. Alizamir et al., “A new insight for daily solar radiation prediction by meteo-
DATA AVAILABILITY rological data using an advanced artificial intelligence algorithm: Deep extreme
The data that support the findings of this study are available learning machine integrated with variational mode decomposition technique,”
Sustainability 15(14), 11275 (2023).
within the article. 23
M. K. Behera, I. Majumder, and N. Nayak, “Solar photovoltaic power fore-
casting using optimized modified extreme learning machine technique,” Eng. Sci.
Technol. Int. J. 21(3), 428–438 (2018).
24
REFERENCES H. M. Khalid et al., “Dust accumulation and aggregation on PV panels: An
integrated survey on impacts, mathematical models, cleaning mechanisms, and
1
Bo. Yang et al., “Classification and summarization of solar irradiance and power possible sustainable solution,” Sol. Energy 251, 261–285 (2023).
25
forecasting methods: A thorough review,” CSEE J. Power Energy Syst. 9, 978 H. Hissou et al., “A novel machine learning approach for solar radiation
(2021). estimation,” Sustainability 15(13), 10609 (2023).
2 26
C. Voyant et al., “Machine learning methods for solar radiation forecasting: A E. S. Solano, P. Dehghanian, and C. M. Affonso, “Solar radiation forecasting
review,” Renewable Energy 105, 569–582 (2017). using machine learning and ensemble feature selection,” Energies 15(19), 7049
3
S. Vadivel, S. Ramasamy, S. Mikkili, M. Ahsan, J. Haider et al., “Hybrid social (2022).
27
grouping algorithm-perturb and observe power tracking scheme for partially M. S. Alam et al., “Ensemble machine-learning models for accurate prediction
shaded photovoltaic array,” Int. J. Energy Res. 2023, 1–8. of solar irradiation in Bangladesh,” Processes 11(3), 908 (2023).
4 28
T. Zahid, K. Xu, and W. Li, “Machine learning an alternate technique to estimate J. Avanija et al., “Prediction of house price using XGBoost regression
the state of charge of energy storage devices,” Electron. Lett. 53(25), 1665–1666 algorithm,” Turk. J. Comput. Math. Educ. 12(2), 2151–2155 (2021).
(2017). 29
R. Chang, L. Bai, and C.-H. Hsu, “Solar power generation prediction
5
T. S. Ustun, S. M. S. Hussain, A. Ulutas, A. Onen, M. M. Roomi, and D. Mashima, based on deep learning,” Sustainable Energy Technol. Assess. 47, 101354
“Machine learning-based intrusion detection for achieving cybersecurity in smart (2021).
grids using IEC 61850 GOOSE messages,” Symmetry 13, 826 (2021). 30
H. Zheng and Y. Wu, “A XGBoost model with weather similarity analysis and
6
T. S. Ustun, S. M. S. Hussain, L. Yavuz, and A. Onen, “Artificial intelligence based feature engineering for short-term wind power forecasting,” Appl. Sci. 9(15), 3019
intrusion detection system for IEC 61850 sampled values under symmetric and (2019).
asymmetric faults,” IEEE Access 9, 56486–56495 (2021). 31
M. Rana, “Overview of data warehouse architecture, big data and green
7
Y. Zhou et al., “A review on global solar radiation prediction with machine computing,” J. Comput. Sci. Technol. Stud. 5, 213–217 (2024).

26 December 2025 13:17:38


learning models in a comprehensive perspective,” Energy Convers. Manage. 235, 32
Ü. Ağbulut, A. E. Gürel, and Y. Biçen, “Prediction of daily global solar radi-
113960 (2021). ation using different machine learning algorithms: Evaluation and comparison,”
8
A. Sharma and A. Kakkar, “Forecasting daily global solar irradiance genera- Renewable and Sustainable Energy Rev. 135, 110114 (2021).
tion using machine learning,” Renewable Sustainable Energy Rev. 82, 2254–2269 33
X.-H. Le et al., “Quantifying predictive uncertainty and feature selection in
(2018). river bed load estimation: A multi-model machine learning approach with particle
9
B. Biswal et al., “Review on smart grid load forecasting for smart energy man- swarm optimization,” Water 16(14), 1945 (2024).
agement using machine learning and deep learning techniques,” Energy Rep. 12, 34
A. M. Assaf et al., “A review on neural network based models for short term
3654–3670 (2024). solar irradiance forecasting,” Appl. Sci. 13(14), 8332 (2023).
10 35
B. Roy et al., “Deep learning based relay for online fault detection, classification, M. M. Alam et al., “Deep learning based optimal energy management for photo-
and fault location in a grid-connected microgrid,” IEEE Access 11, 62674–62696 voltaic and battery energy storage integrated home micro-grid system,” Sci. Rep.
(2023). 12(1), 15133 (2022).
11 36
R. Girimurugan et al., “Application of deep learning to the prediction of solar K. Albeladi, B. Zafar, and A. Mueen, “Time series forecasting using LSTM and
irradiance through missing data,” Int. J. Photoenergy 2023, 1–17. ARIMA,” Int. J. Adv. Comput. Sci. Appl. 14(1), 313–320 (2023).
12 37
D. S. Kumar et al., “Solar irradiance resource and forecasting: A comprehensive F. Sheng and Li. Jia, “Short-term load forecasting based on SARIMAX-LSTM,”
review,” IET Renewable Power Gener. 14(10), 1641–1656 (2020). in 2020 5th International Conference on Power and Renewable Energy (ICPRE)
13
P. Chaudhary et al., “Forecasting solar radiation: Using machine learning (IEEE, 2020).
algorithms,” J. Cases Inf. Technol. 23(4), 1–21 (2021). 38
A. E. Gürel et al., “A state of art review on estimation of solar radiation with
14
H. Musbah, H. H. Aly, and T. A. Little, “Energy management of hybrid energy various models,” Heliyon 9(2), e13167 (2023).
system sources based on machine learning classification algorithms,” Electr. 39
Z. Wang, T. Hong, and M. A. Piette, “Building thermal load prediction
Power Syst. Res. 199, 107436 (2021). through shallow machine learning and deep learning,” Appl. Energy 263, 114683
15
P. Amitasree, G. R. Vamshi, and V. K. Devi, “Electricity consumption fore- (2020).
casting using machine learning,” in 2021 2nd International Conference on Smart 40
S. K. Nayak, “Exploring and forecasting of solar radiation with machine learning
Electronics and Communication (ICOSEC) (IEEE, 2021). methods,” World J. Adv. Res. Rev. 20(3), 824–828 (2023).
16 41
M. G. M. Abdolrasol et al., “Optimal PI controller based PSO optimization for E. Küçüktopçu, B. Cemek, and H. Simsek, “Comparative analysis of single
PV inverter using SPWM techniques,” Energy Rep. 8, 1003–1011 (2022). and hybrid machine learning models for daily solar radiation,” Energy Rep. 11,
17
M. Cordeiro-Costas et al., “Load forecasting with machine learning and deep 3256–3266 (2024).
learning methods,” Appl. Sci. 13(13), 7933 (2023). 42
V. Chandran et al., “State of charge estimation of lithium-ion battery for elec-
18
O. Bamisile et al., “Comparison of machine learning and deep learning algo- tric vehicles using machine learning algorithms,” World Electr. Veh. J. 12(1), 38
rithms for hourly global/diffuse solar radiation predictions,” Int. J. Energy Res. (2021).
46(8), 10052–10073 (2022). 43
D. L. Shrestha and D. P. Solomatine, “Machine learning approaches for estima-
19
Z. Farooq, A. Rahman et al., “Power generation control of renewable energy tion of prediction interval for the model output,” Neural Networks 19(2), 225–235
based hybrid deregulated power system,” Energies 15, 517 (2022). (2006).
20 44
E. Jumin et al., “Machine learning versus linear regression modelling approach J. O. Ogutu, T. Schulz-Streeck, and H.-P. Piepho, “Genomic selection using reg-
for accurate ozone concentrations prediction,” Eng. Appl. Comput. Fluid Mech. ularized linear regression models: Ridge regression, lasso, elastic net and their
14(1), 713–725 (2020). extensions,” BMC Proc. 6, S10 (2012).

AIP Advances 15, 055201 (2025); doi: 10.1063/5.0237246 15, 055201-24


© Author(s) 2025

You might also like