Solar Radiation Prediction via ML
Solar Radiation Prediction via ML
Intitulé
Le jury composé de :
this thesis.
Dedication
I dedicate this modest work to my dear father for his
Bouguerra Oussama
Dedication
I dedicate this modest work to my dear father for his
Benslimane Oussama
Table of Contents
LISTE OF FIGUERS ...............................................................................................................V
LIST OF TABLES...................................................................................................................VI
NOTATION AND ABBREVIATED TERMS …………………...….…………………....VII
INTRODUCTION..................................................................................................................... 1
[Link].................................................................................................................20
[Link] LEARNING .....................................................................................................20
[Link] Principle ...............................................................................................................21
[Link] of Learning in Machine Learning..............................................................21
II.1.2. Regression And Classification..............................................................................22
III. ALGORITHMS OF MACHINE LEARNING...............................................................23
III .1. Random Forest ..........................................................................................................23
III .1.1. Definition Of The Model ....................................................................................23
III .1.2. How It’s Work.....................................................................................................24
III .[Link] The Random Forest .............................................................................26
III .2. Gradient Boosting Machine......................................................................................28
III .2.1. Definition of The Model .....................................................................................28
III .2.2. The Cart Decision Tree ......................................................................................29
III .2.3. Gradient Boosting Machine Algorithm ............................................................30
III .2.4. Hyperparameter Tuning Gradient Boosting Machine....................................31
III .[Link] Learning versus Deep Learning ................................................................33
[Link] LEARNING.............................................................................................................34
IV.2. History of Deep Learning...........................................................................................35
IV.3. Why Deep Learning?..................................................................................................36
IV.4. Deep Learning Application Areas.............................................................................37
V- ALGORITHMS OF DEEP LEARNING .........................................................................37
[Link] Neural Networks ..................................................................................................37
[Link] ...............................................................................................................39
[Link]-layer Perceptron ..........................................................................................39
[Link] Functions ..............................................................................................40
[Link] functions...........................................................................................42
[Link] Regularization .........................................................................................44
[Link] Neural Networks .........................................................................................44
[Link] Short-Term Memory (LSTM)............................................................................47
[Link] ............................................................................................................47
[Link] Principle ......................................................................................................48
[Link] Algorithm ....................................................................................................51
[Link] of LSTM .............................................................................................52
V.4. Bidirectional Long Short-Term Memory (BLSTM) ................................................52
[Link]...................................................................................................................54
[Link].................................................................................................................55
II. Proposed System .................................................................................................................55
III. DESCRIPTION OF THE DATASET .............................................................................56
II .[Link]..............................................................................................................57
[Link] Matrix ......................................................................................................57
[Link] Charts....................................................................................................................58
IV. Preprocessing Dataset .......................................................................................................58
IV.1. Steps involved in data preprocessing ........................................................................58
IV.2. Data Standardization.....................................................................................................59
[Link] EVALUTION...............................................................................................59
[Link] ENVIRONMENT..........................................................................59
[Link] Colab ...........................................................................................................59
[Link] ......................................................................................................................60
[Link] Notebook ...................................................................................................60
V.1.4. Visual Studio Code ................................................................................................61
[Link] Presentation ..................................................................................................61
[Link] Software .............................................................................................................61
V.2. EVALUATION CRITERIA........................................................................................63
[Link] Of Pre section Error.............................................................................63
V.3. SIMULATION RESULTS ..........................................................................................65
[Link] Forest........................................................................................................65
[Link] Boosting Machine (GBM)......................................................................67
[Link] LSTM (BI-LSTM) ...........................................................................69
[Link] Neural Network (DNN) ................................................................................71
[Link] Short Term Memory (LSTM) .....................................................................73
[Link] OF RESULTS.............................................................................75
[Link]...................................................................................................................76
LISTE OF FIGURES
𝑮𝒃 : beam irradiance
𝜽𝒛 : zenith angle
𝜹 : declination
r: distance
𝝎 : Time Angle
TSV : dedicated solar time
𝑨𝒛 : Azimuth
DC : direct current
AC : alternating current
ML : Machine Learning
DL : Deep learning
AI : artificial intelligence
CART : Classification And Regression Trees
RF : Random forest
GBM : Gradient Boosting Machine
According to the International Energy Agency (IEA) [1], global renewable electricity
capacity is projected to rise by over 1 TW, a 46 percent increase over the period 2018 to
2023.
Solar photovoltaic (PV) represents more than half of this expansion and dominates the
growth of renewable ability.
However, because the energy output of PV panels depends on weather conditions such
as cloud cover and solar irradiance, the PV panels' energy output is unstable. To
understand and manage the output variability is of interest for several actors in the energy
market.
In the case of Renewable Energy Source (RES), these machine learning models are
used to forecast the energy produced in a power station or even to forecast the actions of
weather conditions. In the energy industry, such predictions are of great significance. The
problem of Solar Irradiance variability and unpredictability which reaches the surface of
the Earth is well known. A precise forecast of this variable will therefore facilitate better
planning and operation of power delivery at the economic level or at the level of energy
1
output, either by making alternative arrangements for traditional power and overall
timetables or by investing the correct amount of energy resources and reserves to minimize
the operating costs of the power system.
The purpose of this memory is to research the viability of machine learning algorithms
to forecast day-ahead solar irradiance of the next hour and hourly. The studied machine
learning algorithms are the networks Random Forest, Gradient Boosting Machine, Deep
Neural Network ,LSTM, and Bidirectional LSTM. The research uses historical weather
data from the HI-SEAS weather station (Dataset).
2
Chapter I introduces a few of the features and behaviors of this RES. In this chapter,
we will see the value of this form of technology increasing and understand why we should
invest in it.
Chapter II offers an overview of how deep neural networked machine learning works.
It begins with a general machine learning summary and then continues with a more
thorough take on the core components of machine learning. The structure of a neural
network and the algorithm for gradient descent optimization are discussed in some detail.
Chapter III focuses on the empirical results, analysis, and discussion of the results are
presented.
Finally, the last section describes the key conclusions taken from the analysis along
with implications for further field studies.
3
CHAPTER I
I. INTRODUCTION
Renewable energies are inexhaustibly provided by the Sun, the wind, the heat of the
Earth, waterfalls, tides, or the growth of plants. These are the energies of the future. Today,
they are underexploited about their potential. For example, renewable energy accounts for
only 20% of global electricity consumption. Control of the random nature of renewable energy
sources such as solar radiation on the ground could allow the proper sizing of solar systems of
all kinds and allow power grid operators to integrate them better. Unfortunately, radiation data
are not widely available. With this in mind, we have endeavored during this study to
contribute to the research of modeling methodologies for the prediction of solar radiation. This
type of prediction is essential because it could, in the long term, make it possible to exploit
better the solar Renewable Energy (RE), whose intermittence heavily penalizes its use. To
make this prediction, it is necessary to have statistical and mathematical tools dedicated to this
type of analysis[17].
[Link] RADIATION
II.1. Definition
Solar radiation is the energy per unit vicinity obtained from the Sun in the shape of
electromagnetic radiation. The SI unit of photovoltaic irradiance is watt per rectangular meter
W/m2. The find out about and size of photovoltaic irradiance is fascinating for the prediction
of the power era of photovoltaic energy plants[18].
Is the radiation measured on a horizontal floor on Earth, coming from mild scattered
by way of the atmosphere. It measures radiation from all points in the sky except for radiation
4
Chapter I State of the Art
from the solar disk. In the absence of atmosphere, there has to be nearly no diffuse sky
radiation[18].
[Link] Radiation
Is the total irradiance from the solar on a horizontal floor on Earth. It is the sum of
DHI, DNI (after accounting for the photovoltaic zenith perspective of the solar 𝜃𝑧 ) and
mirrored radiation. However, due to the fact reflected radiation is commonly insignificant in
contrast to direct and diffuse radiation for all realistic functions, international horizontal
radiation is stated to be the sum of direct and diffuse radiation solely :
𝑮 = 𝑮𝒅 + 𝑮𝒃 𝒄𝒐𝒔(𝜃𝑧 ) (I.1)
the location G denotes the Global Horizontal Irradiance,𝑮𝒅 the Diffuse Horizontal
Irradiance, 𝑮𝒃 Direct Normal Irradiance or beam irradiance and 𝜃𝑧 the zenith angle. The
referred to types of photograph voltaic radiation are established in Figure I.1
The energy, 𝐸𝑝ℎ , of each photon is directly related to the wavelength λ by the relationship
5
Chapter I State of the Art
ℎ𝐶
𝐸𝑝ℎ = (I.2)
λ
The installation of any solar energy system in a given site requires preliminary studies.
Indeed, sizing and simulation are essential to ensure optimal operation. To carry out such
tasks, reliable measurements over relatively long periods of certain meteorological variables,
and especially those of solar radiation, are essential. The lack of a long series of data or data
series of low quality (discontinuity and unreliability) can combine errors in the design, sizing,
and prediction of solar system performance, which hurts investment. Unfortunately,
measurements of solar radiation are generally inaccurate and rare worldwide; Especially in
Algeria, due to the high price of measuring devices. There are only a small number of solar
radiation stations, which is why there is a lack of solar radiation measurements in large areas
on the one hand. On the other hand, where these data exist, there are generally periods of
failure due to failures or low monitoring, since the majority of these stations belong to
establishments that do not benefit economically from these data.
The world has developed, and energy needs are growing to support both economic
development and the requirements in terms of comfort and consumption of populations. At the
moment, we are coming to a critical moment in energy exploitation: we realize the fragility
6
Chapter I State of the Art
and inconsistency of our functioning. Indeed, the planet’s resources in fossil sediment are
depleting, and oil is rarefied, and, in addition to the economic consequences, it is clear that we
must either find alternatives to current energy sources or find an alternative to our mode of
civilization itself. Without energy, our daily life disappears. Also, the exploitation of fossil
fuels poses another problem: the impact on the environment is massive. If this has long been
ignored, the preservation of the environment becomes a global issue, again with significant
economic stakes. The environmental community has long been aware of this policy, and
recently public opinion. The ecological and industrial disasters plunge people into a grip of
awareness of the dangers generated by the impact of humanity on our planet. We shall see in
this section the evolution of mentalities and policies over the past thirty years in the face of
these difficulties and this ever-increasing demand for energy. We will then focus on the
implications of these various directives at the international and national levels [19].
[Link] Radiation
24 × 60 2𝜋𝑛
𝐺0 = 𝐺𝑠𝑐 [1 + 0.034 cos ( 365 )] [𝜔0 . sin(𝜑) . sin(𝛿 ) + cos(𝜑) . cos(𝛿 ) . sin(𝜔0 )] (I.3)
𝜋
The radiation received on the Earth’s atmosphere occupies only a small portion of the
spectrum of solar electromagnetic waves. Wavelengths between 0.2 and 2.5 µm characterize
it; it includes the domain of the visible (light waves from 0.4 to 0.8 µm).
7
Chapter I State of the Art
Solar energy collectors, which correspond to solar cells, will therefore have to be
compatible with these wavelengths to be able to trap the photons and return them in the form
of electrons [20].
There are two major movements of the earth: the revolution of the world around them
The Sun and the earth’s rotation around its polar axis. These two movements are
Important in solar energy applications. Generally, it is more convenient to study the apparent
motion of the Sun in the celestial vault.
The position of the Sun in the celestial vault is detected by two systems of classical
coordinates: the time system and the horizontal system.
The Earth rotates around the Sun in an eccentric elliptical trajectory 𝐸0 = 0.0167, as
shown in Figure (I.3). The annual periodicity of this motion makes it possible to understand
the phenomenon of the seasons. The distance earth-sun, therefore, varies during the year. On
average, the earth-sun distance is used as the basis for the “astronomical unit,” 1 ua
corresponding to 150.106 km (𝑟0 ). It reaches its maximum at the summer solstice (Aphélie;
1,017 ua or 152,106 km) and its minimum at the winter solstice (Perihelia; 0,983 ua or
8
Chapter I State of the Art
147,106 km). It depends on the day of the year number j, which varies from 1 to 365 (or 366
for leap years). The ground distance (r) is given by the equation (I-4) proposed by and which
gives a good precision [19].
𝑟
𝐸0 = ( 𝑟0 ) = 1.00110 + 0.034221𝑐𝑜𝑠𝛽 + 0.001280𝑠𝑖𝑛𝛽 + 0.000719𝑐𝑜𝑠𝛽
(I-4)
+0.000077𝑠𝑖𝑛2𝛽
2π(j−1)
With: 𝛽 = in radians
365
The solar declination represents the angle between the Earth-Sun direction and the
plane of the equator at noon dedicated solar time. This is a magnitude that often intervenes in
the various calculations related to solar radiation. It varies sinusoidally during the year
between -23.45 (winter solstice) and 23.45 (summer solstice) and cancels out at equinoxes.
Several mathematical formulae have been proposed to calculate approximate values of (𝜹).
Equation (I.5) is the one we adopted for our calculations; it gives a (𝜹) with great precision :
9
Chapter I State of the Art
2π(j−1)
Wit: 𝛽 = in radians
365
𝝎 is the angle between the meridian and the hour circle that contains the Sun. It is
positively counted westward from the meridian. It measures the Sun’s course in the sky. The
hourly angle varies from 15° per hour canceling at noon, dedicated solar time (TSV), it is
calculated from the real solar time by the relationship:
𝝅
𝝎= (𝑻𝑺𝑽 − 𝟏𝟐 ) (I.6)
𝟏𝟐
TSV represents dedicated solar time in hours based on earth rotation around its polar axis
and its revolution around the Sun. The duration of the solar day varies during the year
because:
The real solar time differs from the legal time (TL) of the site considered for three main
reasons:
The difference in longitude between the site under consideration and the longitude
used as the legal time reference (TL) is 4 minutes per degree of longitude.
Correction due to legal time changes between summer and winter.
The difference in solar time from one day to the next. Moreover, it is the correction of
the equation times. This correction varies during the year from -14.3mn to +16.4mn.
The approximate formula can calculate it:
The relationship between TSV and TL taking these corrections into account is:
24𝐿
𝑇𝑆𝑉 = 𝑇𝐿 + 𝐸𝑡 + +𝐶 (I.8)
2𝜋
10
Chapter I State of the Art
L is the longitude of the site considered about the meridian of Greenwich, affected by the sign
(+) for the longitudes East and the sign (-) for the longitudes West.
2𝜋 24𝐿
𝜔= (𝑇𝐿 + 𝐸𝑡 + + 𝐶 − 24 ) (I.9)
24 2𝜋
We also calculate the hourly angle of sunset and sunrise by the relationship following :
𝜔𝑠 = 𝑐𝑜𝑠 −1 (− tan(𝜑) . tan(𝛿 )) where 𝜔 is the latitude and 𝛿 the angle of declination [22].
The horizontal system is most convenient for common applications. Two angular sizes
determine the position of the Sun: The height of the Sun (h) and azimuth ( 𝑨𝒛 )(Figure I. 5).
[Link] (𝑨𝒛 )
It is the horizontal angle of the direction of the Sun with the direction of the south. The
knowledge of the azimuth.’ 𝑨𝒛 ’ makes it possible to calculate the angle of incidence of the
rays on a non-horizontal surface [19].
11
Chapter I State of the Art
[Link] Height
It is the vertical angle of the direction of the Sun with the horizontal plane. Sometimes
one talks about the zenith angle, which is the complement to h, such as Az = 90°-h. The height
angular of the sun h measures the angular distance of the Sun from the horizon[19].
Solar photovoltaic energy has been of great interest in recent years. It is non-polluting
energy and provides real solutions to the various problems that The European Union and its
Member States are currently discussing climate change and the energy crisis [20].
[Link] Effect
[Link] System
The photovoltaic system consists of a field of modules and a set of components that
adapt the electricity produced by the modules to the specifications of receptors[21].
The following figure represents the synoptic diagram of an autonomous photovoltaic system
12
Chapter I State of the Art
[Link] Cell
[Link]
It consists of the stacking of two layers of silicon previously exposed to ion beams, one
to phosphorus(-) ions and the other to boron(+) ions. The first layer has an excess of electron
and the other a deficit. They are said respectively doped N and doped P. This process is called
the «doping» and serves to create an electric field between the two zones where is created a
junction called PN, and directed of the zone (P) to the zone (N).
The area (N) is covered by a metal grid that acts as a cathode K while a metal plate A
covers the other side of the crystal and acts as an anode. A light beam that strikes the device
can penetrate the crystal through the grid and cause a voltage to appear between the cathode
and anode.
13
Chapter I State of the Art
[Link] Principle
When the two doped layers are introduced into contact, the extra electrons in the doped
N cloth diffuse into the doped P material. The in the beginning doped N vicinity will become
positively charged, and the at first doped P region is negatively charged. An electric field is
thus created between them, which tends to push the electrons back into the N zone and the
holes towards the P zone; a junction known as PN has been formed. By adding metal contacts
to the N and P zones, a diode is obtained.
When this diode is illuminated, photons having an energy (ℎ𝑣 ) greater than or equal to
the bandwidth of the forbidden band 𝐸𝑔 Excite the silicon atoms and create positive and
negative charges; thus, the electrons and holes created in the P and N regions respectively
diffuse and reach the space charge zone, accelerated by the internal electric field, they cross
the transition zone. The N region receives electrons and charges negatively. The P region
accepts holes and charges positively.
If a charge is placed at the cell terminals, the electrons in the N-zone join the holes in
the P-zone via the external connection, creating an electrical current [21].
The function of a cell can be represented by the curve I=f (V), which indicates the
evolution of the current generated by the photovoltaic cell according to the voltage at these
terminals from the short circuit to the open circuit.
14
Chapter I State of the Art
the short circuit current (𝐼𝑐𝑐 ) corresponding to the current supplied by the cell
when the voltage at its terminals is zero.
the circuit voltage (𝑉𝑐𝑜 ) corresponding to the voltage that appears at the
terminals of the cell when the current flow is zero.
Between these two values, there is an optimum, at a voltage known as the maximum
voltage 𝑉𝑚 and a maximum current I'm, giving the greatest power (𝑃𝑚𝑝𝑝 ) or peak power[21].
The pattern of the current-voltage characteristic (Figure (I.9) , Figure (I.10)) varies
according to the environmental conditions (illumination and temperature).
15
Chapter I State of the Art
temperature. Therefore the maximum power (Figure (I.11)) delivered by the photovoltaic cell
decreases.
[Link]
Apart from the specific efficiency of each cell type (depending on the properties of the
material used), the final efficiency depends on the energy captured on the cell surface. This
depends on the solar irradiation arriving on the surface of the cell, which, in addition to the
16
Chapter I State of the Art
factors mentioned above (latitude, declination, solar angle, etc.), depends on the angle of
incidence.
The maximum efficiency, in the same place, is obtained when the solar radiation is
perpendicular to the catchment area, that is to say, the angle of incidence of the radiation on
the cell is 90° [22].
Several technologies are currently being developed. They are classified into three
categories (generations) whose detailed description is widely discussed in the literature
specialized
Developments around the third generation target yields ranging from 30-70% and significant
cost reductions[22].
17
Chapter I State of the Art
[Link] Battery
The fact that solar energy is not available for operation of the powered system requires
the use of batteries in the installations autonomous to store energy.
[Link]
[Link]
[Link](Users)
There are two types of devices powered by the system, the direct current such as
telecommunications equipment, water pumping, and AC in the case of domestic use. This case
requires an inverter. The use of photovoltaic energy must be thought in terms of power. It is,
therefore, more advantageous to look for consumers running continuously rather than adding
an inverter and a 220 𝑉𝑎𝑐 consumer.
V. Conclusion
In this chapter, we have defined the majority of meteorological parameters, which are
included in the calculations of solar radiation, whether out of the atmosphere or soil. We have
18
Chapter I State of the Art
recalled the most important empirical models of solar irradiation. Finally, we presented the
experimental data, which are necessary for forming a database for our prediction models
throughout this study. There are several ways to process solar radiation data. The next chapter
is devoted to the state of the art of solar radiation prediction as that a time series.
19
CHAPTER II
techniques
Chapter II Machine learning and deep learning techniques
[Link]
Artificial intelligence (AI) has over the past decade become a popular topic both within
and outside the scientific community; a wealth of articles in technology and non-technological
journals have the subject of machine learning (ML), Deep Learning (DL) and AI. Still there is
confusion about AI, ML and DL. The terms are strongly interconnected, but not
interchangeable. In this review we do not (try) use technical jargon to better explain these
concepts to a clinical audience.
In 1956, a group of computer scientists suggested that computers could be
programmed to think and argue, "that every aspect of learning or any other property of
intelligence [could] be described in such a precise way that a machine [could] be induced to
simulate it." They described this principle as "artificial intelligence." Simply put, AI is an area
focused on automating the intellectual tasks normally performed by humans, and ML and DL
are specific methods to achieve this goal. This means that they are in the field of AI, but AI
includes approaches that do not include any form of "learning." Thus, the sub-field known as
symbolic AI focuses on hardcoding rules (i.e., explicit writing) for each possible scenario in a
particular area of interest. These rules, written by people, come from a priori knowledge of the
particular topic and task to be completed. For example, if you would program an algorithm to
modulate the room temperature of an office, he or she probably already know what
temperatures are convenient for people to work at and would cool the room, when
temperatures rise above a certain threshold and heat when they fall below a lower threshold to
program. Although symbolic AI can solve clearly defined logical problems, it often fails in
tasks that require higher pattern recognition, such as speech recognition or image
classification. These more complicated tasks are where ML and DL methods perform well.
This report summarizes machine learning and deep learning methodology for the audience
without extensive technical computer programming background [23].
[Link] LEARNING
20
Chapter II Machine learning and deep learning techniques
Machine learning has been a real success in recent years. With the growth exponential number
of digital data available, we need to use new analysis methods, and so-called machine learning
methods correspond to this need. In this paper, we will study several types of machine
learning algorithms, and in this part, we will explain different generalities common to all.
We’ll start with a quick introduction of the other learning families, and then the general
operation of a machine learning algorithm. Then we will list the main advantages and
disadvantages, with a closer look at over-learning. Finally, we will explain the simplest
machine learning model, as well as its use in this memory [24].
[Link] Principle
[Link] Learning
Supervised learning is intended to create a predictive function for one of the variables
Of our database as a function of others. The variable we want to predict is called
"variable to be explained," the other variables used to guide the prediction are the "variables
Explanatory". The algorithm tries to learn, by browsing the database available,
The different causal links between the explanatory variables and the variable to be
explained. Once our model is created, it matches each possible combination of explanatory
variables, a prediction of the corresponding variable to be explained. For this, he must group
the individuals of the base into subgroups by maximizing the homogeneity of the variable to
be defined. Still, the groups must be separated only according to their explanatory variables.
There are two categories of supervised learning:
Regression algorithms, when the variable to be explained, is quantitative. In this case,
the prediction is a value.
Classification algorithms, when the variable to be explained, is qualitative. In this case,
the prediction is often a probability of belonging to the different modalities of the
variable in question [25].
21
Chapter II Machine learning and deep learning techniques
[Link] Learning
Non-served learning occurs when there is no response variable to predict. These
algorithms are also used on a database, but in this case, its purpose is to determine the
structure present in that database.
To do this, it must, like the supervised algorithms, group individuals into the most
homogeneous subgroups possible, however here we no longer have variables to explain, so
homogeneity must be done on all variables.
What interests us here is not the prediction of new data, but instead, how groups are
determined and what commonalities exist between the individuals of each subgroup.
Of the two learning families discussed above, we will only use supervised learning
algorithms in this brief. We aim to estimate a variable to be explained, the S/P ratio of
contracts, based on the information available about these contracts, which are the explanatory
variables. Supervised learning methods, therefore, appear to be adapted to our problems, such
as the CART algorithm (Classification And Regression Trees), bagging, or random forest. The
purpose of this section is to present the operation and the main characteristics common to
these methods [25].
[Link] Learning
In this type of study, each marked and unlabeled statistic may want to be utilized to
shape the indispensable knowledge. The framework receives a reward for each good or
incorrect forecast. Depending on the reward, the subsequent forecast should be generated. At
the factor when new information is given to the framework, the framework will pastime to
stumble on the excellent execution way or be part of a couple of execution pathways for
forecasting and pause for the reward. When the received reward takes place to be greatest with
recognition to the previous rewards for equal input, at that point, this pathway flips out to be
agreeable Reinforcement gaining knowledge of is utilized in net-primarily based games, for
example, Chess [25].
22
Chapter II Machine learning and deep learning techniques
II. [Link]
Regression is used when predicting a continuous variable, which can therefore take
any value. The classes representing the note change have therefore were considered constant
and the regression algorithm allowed to obtain decimal values that we finally rounded to get a
vector of predicted levels [26].
23
Chapter II Machine learning and deep learning techniques
it is achieved that
𝐵
1
Ŷ𝑖 = ∑ ŶRF(m)i,k (II.1)
𝐵 𝑘=1
24
Chapter II Machine learning and deep learning techniques
The formula is, therefore, the same as for a bootstrap, the difference in how to build
the tree after creating the bootstrap sample [24].
As in the previous part, this method is explained in a prominent and concise manner
with a diagram:
initial database
Ŷ1 Ŷ2 Ŷ𝐵
𝐵
1
the final result is the average : ∑ Ŷ𝑗
𝐵 𝑗=1
25
Chapter II Machine learning and deep learning techniques
As with the bagging model, the first step in creating our random model forest is the
determination of the parameters. There are two of them: B, the number of samples bootstrap,
and mtry the number of variables selected at each separation. To stay consistent with the
previous part, we will note a 𝐵𝑜𝑝𝑡𝑖 and 𝑚𝑡𝑟𝑦𝑜𝑝𝑡𝑖 the chosen parameters.
For parameter B, we will operate in the same way as for the bagging model. We are
therefore going to test a random forest model with a large number of trees, here 1000. Then
we will find from what number of trees the OOB error, defined in the previous part, is
stabilized.
To create our random forest model with 1000 trees to determine 𝐵𝑜𝑝𝑡𝑖 We had to set a
value for mtry. As a reminder, a good indication of this parameter is mtry ≈ √𝑝 , with p the
number of variables that can potentially be selected for each draw.
Here, p = 29, this number is more significant than that mentioned in the section on
elimination variables. Indeed, in a random forest, for the category variable, for example, the
model can draw separately each of the six variables created during the aggregation of the
database. So if we seek to determine the number of variables, the category of the vehicle
counts for six here. In contrast, it counted for only one in the selection of variables, which
explains this difference.
For now, we set the mtry parameter as follows: mtry ≈ √29 ≈ 5.4, we will take so for
this test the rounding to the nearest integer, that is to say, mtry = 5. Of course, we will then
determine the value of 𝑚𝑡𝑟𝑦𝑜𝑝𝑡𝑖 using a more sophisticated method. We set it to 5 only the
time of the tests to determine 𝐵𝑜𝑝𝑡𝑖 [26].
Although random forests perform well out-of-the-box, there are several tunable
hyperparameters that we should consider when training a model. The main hyperparameters to
consider include:
26
Chapter II Machine learning and deep learning techniques
min_samples_split = min number of data points placed in a node before the node is split
We can see on this graph the evolution of the OOB error as a function of the number of
bootstrap samples, and therefore of the number of trees. We can see that the error stabilizes
later as for the bagging model. Here we will now fix 𝐵𝑜𝑝𝑡𝑖 = 400 [26].
[Link]. [Link] error depending on the number of trees in the random forest model.
Now that we have determined the first parameter, Bopti, let's move on to the choice of
the second, 𝑚𝑡𝑟𝑦𝑜𝑝𝑡𝑖 .
For the mtry parameter, we will create models with different values for mtry, and then
we will test these models on our validation basis to calculate the MSE error. We will choose
the mtry that gave the model with the lowest MSE.
27
Chapter II Machine learning and deep learning techniques
We first had to choose a range of mtry values that we are going to test. As the
indication mtry ≈ √𝑝 gave us a result between 5 and 6, we took a range around these values,
so we have arbitrarily set it between 2 and 12. If we subsequently find that this range is not
sufficient to make our choice, we give ourselves the option to change it later [26].
Finally, we specify that all models tested on this range are tested with 400 bootstrap
samples, as we have chosen previously.
We can see that the indication mtry ≈ √𝑝 looks good because models with mtry
values of 5 and 6 are better than others. We now fix 𝑚𝑡𝑟𝑦𝑜𝑝𝑡𝑖 = 6, whose model got a slightly
smaller error than the one with mtry = 5.
Thus, the estimate made in this part will be composed of 12 random forest models,
each having 400 bootstrap samples and a value of 6 for the parameter mtry. The models thus
created obtained an MSE on the test basis of 0.87 by predicting the S / P and 0.74 by the
number of claims. These are the best scores so far [24].
28
Chapter II Machine learning and deep learning techniques
The most usual approach of boosting is likely the one developed via Freund and
Schapire in 1995: the AdaBoost (for Adaptive Boosting). However, it will now not be precise
in this brief. However, the reader is referred to Article A Decision-Theoretic Generalization of
On-Line Learning and an Application to Boosting for more critical data about this method.
Like different boosting methods, the Gradient Boosting optimizes the overall performance of a
set of so-called “low” prediction fashions by means of assembling them into a remaining
model. This is known as a “low” prediction model, a classification or regression approach that
is simply barely greater nice than a random draw. The low prediction mannequin normally
used with a Gradient Boosting is a CART choice tree. More concretely, this approach of the
Gradient Tree Boosting then consists in making a succession of choice bushes the place every
mannequin is constructed on the residual error of the preceding one [25].
29
Chapter II Machine learning and deep learning techniques
This concept of impurity reduction can be seen in (Figure II.5), where the removal of
impurity is more critical in the illustration on the left than in the illustration on the right.
The problem of maximizing impurity ∆𝑖 defined in equation (II.4) can also be written:
where 𝑃𝑔 and 𝑃𝑑 are the probabilities of the left and right nodes, respectively. This
problem translates the way the algorithm works: it scans all possible values for all variables 𝑥𝑖
for j = 1,...,p to find the best cutout that will maximize the variation of impurity ∆𝑖 . If the node
variance represents the impurity function, then the maximization problem 3.2 can be rewritten:
For more details on the machine learning method CART, the reader can refer to
Roman Timofeev, Classification, And Regression Trees[27].
The Gradient Boosting is an iterative algorithm that initially distributes weights equal
to all predictions and then adapts them to each step, so that bad predictions are over-weighted
to the next step so that the “low” prediction model Pay more attention to it. Let’s go back to
Friedman’s original article by using his notes to develop the idea of Gradient Boosting in more
detail. We have a sample size (n) consisting of (p) explanatory variables. Note 𝑥𝑖 = (𝑥𝑖,1 ,...
𝑥𝑖,𝑝 ) ∈𝑅𝑝 , The vector of the explanatory (p) variables corresponding to the observation i. The
objective is to find an approximation 𝐹approx (𝑥) of a function F(x) linking the explanatory
variables x to the response variable y and minimizing the expectation of a certain loss function
30
Chapter II Machine learning and deep learning techniques
L(y, F(x)) on the attached distribution of x and y [28]. Mathematically, the function F(x) is
then defined by:
For 𝑚 = 1 to M do :
o Step 1. Compute the negative gradient
𝜕𝐿(𝑦𝑖 , 𝐹 (𝑥𝑖 ))
ȳ𝑖 = −
𝜕𝐹𝑥𝑖
o Step 2. Fit a model
𝑁
To avoid over-learning of the data and to degrade future predictions, the Gradient
Boosting has several hyper-parameters that allow optimizing its use by limiting this over-
learning phenomenon [29]:
A-Min_Samples_Split
Defines the minimal range of samples (or observations) which are required in a node to
be viewed for splitting.
31
Chapter II Machine learning and deep learning techniques
It is used to manage over-fitting. Higher values stop a model from mastering members
of the family, which would possibly be noticeably unique to the specific pattern chosen
for a tree.
Too extreme values can lead to under-fitting; hence, it ought to be tuned to the use of
CV.
B- Min_Samples_Leaf
Defines the minimal samples (or observations) required in a terminal node or leaf.
It is used to manage over-fitting comparable to min_samples_split.
Generally, decrease values ought to be chosen for imbalanced type troubles due to the
fact that the areas in which the minority classification will be in the majority will be
minimal.
C- Max_Depth
D-Max_Features
The range of facets to reflect on consideration while looking for an exceptional split.
These will be randomly selected.
As a thumb-rule, the rectangular root of the complete quantity of facets works
excellent. However, we ought to take a look at up to 30-40% of the total variety of
features.
Higher values can lead to over-fitting; however, relies upon case to case.
E- Learning_Rate
He determines the have impact of every tree on the closing effect. GBM works by way
of beginning with a preliminary estimate, which is up to date the usage of the output of
32
Chapter II Machine learning and deep learning techniques
every tree. The gaining knowledge of parameter controls the magnitude of this
alternate in the forecast.
Lower values are commonly favored as they make the mannequin strong to the unique
traits of the tree and, as a consequence permitting it to generalize well.
Lower values would require a greater wide variety of timber to mannequin all the
family members and will be computationally expensive.
F- N_estimators
G- Loss
33
Chapter II Machine learning and deep learning techniques
from the training dataset. DL, on the contrary, could be considered as establishing both
representation learning and machine learning together. DL pursuits to together learn essential
features along with multiple levels of cumulative intricacy and abstraction and the concluding
prediction. Figure (II.6) illustrates the fundamental difference between ML and DL, where
traditional ML involves manual feature selection, and on the contrary, DL employs automated
feature selection.
[Link] LEARNING
Deep Learning is a new area of ML research, which has been introduced to bring the
ML closer to its primary objective: artificial intelligence. It concerns algorithms inspired by
the structure and functioning of the brain. They can learn multiple levels of representation to
model complex relationships between data.
[Link]. [Link] relationship between artificial intelligence, ML, and deep learning.
Deep Learning is based on the idea of artificial neural networks and is designed to
manage large amounts of data by adding layers to the network. A deep learning model can
extract characteristics from raw data through multiple layers of processing consisting of
multiple linear and non-linear transformations and learn about these characteristics step by
34
Chapter II Machine learning and deep learning techniques
step across each layer with minimal human intervention. Over the past five years, deep
learning has moved from a niche market where only a handful of researchers were interested
in the area most sought by researchers. Research related to deep learning now appears in top
journals like Science, Nature, and Nature Methods, to name a few. Deep learning has
conquered the GO, learned to drive a car, diagnosed cancer and autism, and even become an
artist. The term "Deep Learning" The time period "Deep Learning" used to be first added to
ML via Dechter (1986) and artificial neural networks by using Aizenberg et al. (2000) [32].
This paper presents a history of deep learning from Aristotle to the present. The
various milestones are summarized in this table
35
Chapter II Machine learning and deep learning techniques
ML algorithms work well for a wide variety of problems. However, they have failed to
solve some significant issues of AI, such as voice recognition and object recognition. The
development of deep learning was motivated in part by the failure of traditional algorithms in
such a task of AI. But it wasn’t until more data was made available thanks to Big Data and
connected objects and computing machines became more powerful that the real potential of
Deep Learning was understood.
36
Chapter II Machine learning and deep learning techniques
[Link]. 8. The performance difference between Deep Learning and most ML algorithms
depends on the amount of data.
Voice recognition
Automatic tagging of music pieces
Advanced speech synthesis
The design of new pharmaceutical molecules
There are a large number of variations of deep architectures. Most of them are derived
from some original parenting architectures. It is not always possible to compare the
performance of all architectures, as they are not all evaluated on the same datasets. Deep
Learning is a rapidly growing field, and new architectures, variants, or algorithms appear
every week.
37
Chapter II Machine learning and deep learning techniques
and a mathematician Walter Pitts in 1943. They mentioned how neurons would possibly work.
Neural networks consist of the layer of entering neurons or alerts, which can be different
characteristic values, an output layer the place the result of the community obtained, and a
quantity of a range of hidden layers between the enter and output layers. Also, each layer has a
few or various neurons.
Input alerts are surpassed thru the network, layer by way of layer, by way of the usage
of the weighted connections to, in the end attain the output layer. At some neurons, a nonlinear
feature can be triggered. The purpose of gaining knowledge of the method is discovering
weights that would make the neural community display favored behavior. This is the instance
of Multilayer Perceptron (MLP), which is additionally referred to as Feedforward Neural
Networks (FNN). An accepted structure of artificial neural networks is proven in Figure
(II.9).
Despite the reality that the feedforward neural networks have been efficaciously
applied and bought higher overall performance in many tasks, it does now not think about the
transient prospect that characterizes sequential data. That is, they are no longer moderately
proper when it comes to the statistics that rely on the preceding data. To remedy such a variety
of troubles, the neural networks have advanced to so-referred to as recurrent neural networks.
In the subsequent subsections [43].
38
Chapter II Machine learning and deep learning techniques
[Link]
The fundamental unit of neural networks is one of the types of artificial neurons, called
a perceptron. A perceptron takes a fixed number of inputs and produces a single output. The
way of computing the result essentially has several steps, such as taking an input, calculating
the weighted sum by introducing weights, bias term, and applying activation function. The
process described above can be formalized as follows:
The perceptron has n inputs represented as an input vector x = (𝑥1 , 𝑥2 , . . . , 𝑥𝑛 ). Each input
has an assigned weight that defined by a vector of weights w = (𝑤1 , 𝑤2 , . . . , 𝑤𝑛 ).
Consequently, the weighted input values are combined, which gives the weighted sum:
ε = 𝑤 · 𝑥 = ∑𝑛𝑖=1 𝑤𝑖 · 𝑥𝑖 (II.8)
At this step, the activation function is applied to the weighted sum to calculate the output, and
the weighted sum is compared with a threshold θ to produce an output y that is either 0 or 1,
depending on whether or not it exceeds the threshold. Thus,
1, ε ≥ 0
𝑦 = σ(ε) = { (II.9)
0, ε < 0
[Link]-layer Perceptron
39
Chapter II Machine learning and deep learning techniques
neuron of the identical layer or of the neuron of preceding layers (this is the case for recurrent
neural networks)
nodes that are no goal of any connection are referred to as entering neurons. An MLP
that must be utilized to enter patterns of dimension n needs to have n enter neurons, one for
every dimension. Input neurons are usually enumerated as neuron 1, neuron 2, neuron three
nodes that are no supply of any connection are referred to as output neurons. An MLP
can have extra than one output neuron. The range of output neurons relies upon on the
way the goal values (desired values) of the education patterns are described.
all nodes that neither enter neurons nor output neurons are referred to as hidden
neurons.
40
Chapter II Machine learning and deep learning techniques
Sigmoid
1
𝑓 (𝑥) = 1+𝑒 −𝑥 𝑂𝑟 𝑥 ∈ R (II.9)
The goal is to convert the input value into a probability of 1 if it is a very large
positive number, and conversely, to 0 if the input is a very large negative number
Softmax
𝑧
𝑒 𝑗
𝑓(𝑥)𝑗 = 𝐾 for all j ∈ {1, … , K}, x = {x1, … , xk} ou k ∈ R (II.10)
∑𝐾=1 𝑒 𝑧𝑘
The function gives an output vector of K strictly positive real numbers and sum 1.
Tanh
𝑒 𝑥 − 𝑒 −𝑥
𝑓 (𝑥 ) = tanh(𝑥) = (II.11)
𝑒 𝑥 + 𝑒 −𝑥
41
Chapter II Machine learning and deep learning techniques
ReLu
The ReLu linear grinding unit function is the simplest function, defined by:
0 𝑓𝑜𝑟 𝑥 < 0
𝑓 (𝑥 ) = { (II.12)
𝑥 𝑓𝑜𝑟 𝑥 ≥ 0
[Link] functions
SGD
42
Chapter II Machine learning and deep learning techniques
iteration SGD (Stochastic Gradient Descent), which updates the parameters for each example
of the 𝑥𝑖 dataset and 𝑦𝑖 label (label).
η: modulates the correction (η too low, slow convergence; η too high, oscillation)
𝛻(𝑥𝑡): The gradient at point 𝑥𝑡 shows the direction and importance of the slope in the
vicinity of 𝑥𝑡.
This method is usually faster. However, due to frequent updates, convergence becomes more
difficult (find the minimum objective function).
Adam
“Adam Adaptive Moment optimization” is one of the latest and most effective
algorithms for gradient descent optimization. Adam calculates the exponential mean of the
gradient as well as the squares of the gradient for each parameter. The learning rate is then
multiplied by the mean of the gradient and divided by the square root of the exponential mean
of the gradients. Then the update is added.
𝑣𝑡 = 𝛽1 × 𝑣𝑡−1 − (1 − 𝛽1 ) × 𝑔𝑡 (II.13)
𝑣𝑡
∆w𝑡 = −ƞ × 𝑔𝑡 (II.15)
√𝑠𝑡 +ϵ
43
Chapter II Machine learning and deep learning techniques
Adamax
𝑝 𝑝 𝑝
𝑠𝑡 = β2 × 𝑠𝑡−1 − (1 − β2 ) × 𝑔𝑡 (II.18)
Ou p = ∞.
Adadelta
[Link] Regularization
Deep neural network (DNN) fashions have several parameters and can model
relatively composite functions. This capability is a boon and a bane. Such prototypes would
often overfit on the training-set and would drop accuracy and generalizability over the test-set.
Regularization in ANN terminology speaks of the technique of regulating neural network
layers for stopping the over-fitting. Dropout (also recognized as dropout chance or dropout
rate) is the most extensively utilized regularization approach in DL. During the learning
manner, the hidden layer(s) neurons are chosen randomly and are discarded, relying on the
dropout rate. Precisely, randomly chosen neurons are dropped-out, i.e., dropped out neurons
could not replace weights anymore, accordingly supporting the learning manner to avert the
problem of overfitting [38].
Humans don’t start thinking from scratch every second. By reading this essay, you
understand each word according to your understanding of the preceding words. You don’t
throw everything away and start thinking from scratch again. Your thoughts have
perseverance.
44
Chapter II Machine learning and deep learning techniques
Traditional neural networks (TNNs) can’t do that, and that seems to be a significant
gap. For example, imagine that you want to classify the type of event that occurs at each stage
of a movie. It is not known how a traditional neural network could use its reasoning on
previous events in the film to inform them later.
These loops make recurring neural networks a bit mysterious. However, if you think a
little more, it turns out that they are not all different from a standard neural network. A
network of recurrent neurons can be considered as multiple copies of the same network, each
transmitting a message to a successor, as shown in Figure II.16. Consider what happens if we
unravel the loop:
45
Chapter II Machine learning and deep learning techniques
This chain nature reveals that recurring neural networks are intimately linked to
sequences and lists. They are the natural architecture of the neural network to be used for such
data.
In recent years, there has been an incredible success in applying NRNs to a variety of
issues: speech recognition, language modeling, translation, captioning of images, etc [39].
One of the appeals of RNNs is the concept that they may be capable of joining
preceding records to the current task, such as the use of remaining video frames may inform
the appreciation of the contemporary structure. If RNNs ought to do this, they’d be
instrumental. But can they? It depends [39].
Sometimes, we solely want to seem at current statistics to operate the current task. For
example, reflect on consideration on a language mannequin is making an attempt to predict the
subsequent phrase-based totally on the preceding ones. If we are making an attempt to
envision the ultimate phrase in “the clouds are in the sky,” we don’t want any in addition
context – it’s rather apparent the subsequent expression is going to be the sky. In such cases,
the place the hole between the applicable records and the region that it’s wanted is small,
RNNs can examine to use the previous data.
46
Chapter II Machine learning and deep learning techniques
France, from in addition back. It’s absolutely feasible for the hole between the applicable
statistics and the factor the place it is required to emerge as very large.
Unfortunately, as that gap grows, RNNs become unable to learn to connect the information.
[Link]
47
Chapter II Machine learning and deep learning techniques
[Link] Principle
LSTMs have internal mechanisms called gates, which are used to regulate the flow of
information by learning what information is useful to keep or forget in time. In doing so, the
network will be able to transmit relevant data, thanks to the state of the cell, acting as a
transport route that transfers related information throughout the sequence chain. As the status
continues, information is added or deleted through three main doors. Initially, the input is the
output predicted previously (ℎ𝑡−1) as well as the current input 𝑋𝑡). This data will be broken
down into three streams. The objective is to update the status of the cell from (𝐶𝑡−1 to 𝐶𝑡.) [40].
A- Input Gate
The Input Gate updates the status of the cell. The activation function is Sigmoid. For each
input, the gate will provide an output value between 0 and 1, and decide which value to update
(0 means not necessary and 1 means important). This result is then multiplied by the current
state
With:
48
Chapter II Machine learning and deep learning techniques
𝑊𝑡 : Weight
σ: Sigmoid function
B- Forget Gate
The Forget Gate forgets, decides what information should be discarded or kept. The
information of the previous state is passed to the Tanh function giving values between -1 and
1. Then, to select the important characteristics, a sigmoid layer will decide which value will be
updated
These two results will be multiplied and added to the state. At this stage, the state of the cell
is:
49
Chapter II Machine learning and deep learning techniques
𝐶𝑡 = 𝑓𝑡 × 𝐶𝑡−1 + 𝑓𝑐 × 𝐼𝑡 (II.23)
C- Output Gate
The output Gate will pass the previous hidden state and the current entry in a sigmoid
function
ℎ𝑡 = 𝑜𝑡 × tanh(𝐶𝑡) (II.25)
50
Chapter II Machine learning and deep learning techniques
[Link] Algorithm
Algorithm: Operating algorithm of the doors according to three algorithms "if-then otherwise."
If Value Input " 0, then
Input Gate: Transmission of external information inside the block
If Value is Forgotten " 0, then
The previous state of the CEC affects the calculation of the present state
If Value Output Gate " 0, then
Input Gate: Transmission of Inside Information to the Outside of the Block
otherwise
Output Gate: No information provided on the exit
End if
otherwise
the cell forgets its past
end if
otherwise
Output Gate: Information blocked
End if
Table. II. [Link] Algorithm [41].
51
Chapter II Machine learning and deep learning techniques
[Link] of LSTM
The constant error backpropagation within memory cells results in LSTM's ability to
bridge very long time lags in case of problems similar to those discussed above
For long time lag problems such as those discussed in this paper, LSTM can handle
noise, distributed representations, and continuous values. In contrast to nite state
automata or hidden Markov models, LSTM does not require an a priori choice of a nite
number of states. In principle, it can deal with unlimited state numbers.
For troubles mentioned in this paper, LSTM generalizes nicely even if the positions of
broadly separated, applicable inputs in the enter sequence do now not matter. Unlike
preceding approaches, ours shortly learns to distinguish between two or greater
extensively isolated occurrences of a specific component in an input sequence, besides
relying on excellent quick time lag training exemplars. There appears to be no need for
parameter no tuning. LSTM works well over a broad range of parameters such as
learning rate, input gate bias, and output gate bias. For instance, to some readers, the
learning rates used in our experiments may seem extensive. However, a large learning
rate pushes the output gates towards zero, thus automatically countermanding its
adverse effects [42].
52
Chapter II Machine learning and deep learning techniques
Both the output of the ahead and backward layers are calculated by way of potential of
the preferred LSTM equations, Equations (II.27) - (II.28). The BLSTM layer produces an
output vector, 𝑦𝑡 , which is calculated via the equation:
Where 𝜎 function combines each the output sequences, the 𝜎 function could be of 4
kinds: concatenating, summation, common and multiplication function, and incorporating
BRNNs with LSTM neurons outcomes a bidirectional LSTM recurrent neural network
(BLSTM RNN). The BLSTM RNN is successful in gaining access to long term context
statistics in each backward and ahead direction. The aggregate of both the forward and
backward LSTM layers is regarded as a single BLSTM layer. It has been proven that the
bidirectional fashions are appreciably higher than regular unidirectional models in some
domains, like phoneme classification and speech recognition [42]. (Figure II.24) illustrates a
bidirectional LSTM shape with three consecutive time steps.
[Link]. [Link] BLSTM RNN structure with three consecutive time steps[42].
53
Chapter II Machine learning and deep learning techniques
[Link]
In this chapter we have presented the basic concepts of Machine learning and deep
learning including its operating principle, its main components and also its limitations. Then,
regarding machine learning, we talked about Random forest (RF) and Gradient boosting
Machine (GBM), about these methods and their use, and we mentioned the pros and cons of
both types.
Then we described a new variant of neural network called Deep Learning (DL). This
technique is characterized by its ability to solve the problem of the complexity of training
(NN) as well as its power to represent the forms (inputs) in a powerful, automatic and
discriminating way.
Finally, we discussed the different models of Deep Learning namely the CNN, the
LSTM, and BILSTM. The latter was very detailed because it will be the subject of several
experiments in the next chapter.
54
CHAPTER III
[Link]
All algorithms based on machine learning follow a predictive model that estimates a
certain type of data with high accuracy. A large data set is essential for the learning algorithm
to understand the behavior of the system. The first step for machine learning is data
acquisition. The collected data were shared by various interested parties and summarized in
useful information. The steps included in this process are data purification and data
delimitation. The data were separated into two disjoint sets, training, testing . The training
dataset was used for model training and testing. The dataset was used for model optimization
and evaluation.
Fig. III. [Link] matrix to identify the essential characteristics of the whole.
55
Chapter III Results and Discussion
Testing Dataset: The sample of data used to provide an unbiased evaluation of a final model
fit on the training dataset.
Dataset: A dataset consists of about two components, the two components are rows and
columns. In addition, a main feature of a record is that it is organized in such a way that each
row contains an observation.
NASA HI-SEAS missions serve as a testbed and training ground for humans as we
develop the capability to explore Mars. A recent NASA Space Apps Challenge hackathon
asked participants to use the data collected on the HI-SEAS site to predict solar radiation
given a set of measurable weather conditions. Knowing when the conditions are most
favorable to solar radiation incident is crucial in deciding when and where to deploy solar
energy recovery equipment, especially for settlers or astronauts on the surface of Mars.
These datasets are four-month weather data from the HI-SEAS weather station (September
to December 2016). For each dataset, the fields are:
A-line number (1-n) is useful to sort the results of this export The UNIX date time_t
(seconds since January 1, 1970). Useful for sorting the results of this export with other export
results Date in yyyy-mm-dd format Local time in hh: mm: ss Format 24 hours Digital data, if
applicable (maybe an empty string) Text data, if functional (can be an empty string).[44]
56
Chapter III Results and Discussion
Solar radiation: watts per meter^2 (W/m2) Temperature: degrees Fahrenheit (°F)
Humidity: percentage (%) Barometric pressure: (Hg)
Wind direction: degrees (°) Wind speed: miles per hour (mph)
Time at sunrise: Hawaii time Time at sunset: Hawaii time
Table. III. [Link] units of each dataset.
II .[Link]
At each timestamp of each day, there are values for all other variables. No other
variables affect time or date values. Therefore, the date and time of day are independent
variables.
For each date, there is a value for "Time at Sunrise" and "Time at Sunset." The
difference in these values gives the length of a given day, which is directly related to the
date. Further exploration of the dataset is required to determine whether the size of a given
date outweighs the amount of useful information it provides.
Temperature, pressure, and humidity do not directly affect each other significantly, but
since they are all properties that describe the local atmosphere, they do not vary independently
of each other. Similarly, these three variables have a strong relationship with the time of day.
[Link] Matrix
Fig. III. [Link] matrix to identify the essential characteristics of the whole.
57
Chapter III Results and Discussion
[Link] Charts
Then, to better understand the data, hourly, and monthly averages of several variables
were viewed using bar charts.
Data preprocessing is a data mining technique in which raw data is converted into an
understandable format. The real data is often incomplete: missing attribute values, missing
specific attributes of interest, data preprocessing is a proven method of solving such
problems[45].
58
Chapter III Results and Discussion
Splitting the data set into test set and training set.
Feature Scaling.
𝑥𝑖 −𝑚𝑖𝑛(𝑥)
𝑥𝑖𝑛 = (III.1)
𝑚𝑎𝑥(𝑥) −𝑚𝑖𝑛(𝑥)
Avec :
[Link] EVALUTION
[Link] ENVIRONMENT
[Link] Colab
Google Colab is a free cloud service that now supports free GPUs, improving coding
skills in the Python programming language. Develop in-depth learning applications using
popular libraries such as Keras, TensorFlow, PyTorch. The most important feature that
distinguishes Colab from other free cloud computing services is: Colab provides a GPU and is
entirely free [46].
59
Chapter III Results and Discussion
[Link]
[Link] Notebook
The Jupyter Notebook is an open-source net utility that lets you create and share files
that comprise stay code, equations, visualizations, and narrative textual content. Uses include
data cleaning and transformation, numerical simulation, statistical modeling, data
visualization, machine learning, and much more [48].
60
Chapter III Results and Discussion
Is a source code editor developed by Microsoft for Windows, Linux and macOS. It
includes built-in Git and support for debugging, syntax highlighting, smart code completion,
snippets, and code refactoring. It is exceptionally customizable, permitting customers to
exchange the theme, keyboard shortcuts, preferences, and set up extensions that add extra
features. The supply code is free and open-source, launched beneath the permissive MIT
license. Compiled binaries are freeware for any purpose [49].
[Link] Presentation
[Link] Software
61
Chapter III Results and Discussion
Tensorflow
TensorFlow was created by the Google Brain team to research ML and Deep Learning. It
is considered a modern version of Theano.
• Advantages
supported by Google
A very large community
Multi-GPU support
• Disadvantages
Keras
The highest level, the most user-friendly framework on the list. It allows users to choose
whether the models they build are running on Theano or TensorFlow.
• Advantages
Python
The perfect backend for Theano or TensorFlow
High-level, intuitive interface
• Disadvantages
[Link]
Deep Learning and machine learning is an area with intense computational
requirements, and the availability of resources (especially in GPU) dedicated to this task will
fundamentally influence the user experience because, without its resources, it will take too
long to learn from one’s mistakes what can be discouraging. The experiments were all carried
62
Chapter III Results and Discussion
out on a machine that offers acceptable performance, the characteristics of which are as
follows:
To make the analysis as rigorous as possible, the tests must be performed under the
same conditions whenever possible. All experiments were taken on the same computer in
order not to compromise the performance analysis. Each test was repeated four times. This is
because some results were not equal with each run of a simulation. In both methods, the
simulation time presents small related variations to obtain the performance error of each
prediction and to analyze each parameter and better, and it was necessary to use some
mathematical tools, pervasive in this case type of studies:
𝟏
𝐌𝐀𝐄 = 𝒏 ∑𝐧𝐣=𝟏|𝐲 − 𝐲𝐢 | (III.2)
63
Chapter III Results and Discussion
𝟏
𝑴𝑺𝑬 = 𝒏 ∑𝒏𝒋=𝟏(𝒚 −𝒚𝒊 ) ² (III.3)
𝟏
√ ∑𝒏
𝒋=𝟏(𝒚−𝒚𝒊 )²
𝒏
RRMSE = 𝟏 𝒏 × 𝟏𝟎𝟎 (III.5)
∑ 𝒚
𝒏 𝒋=𝟏
Different ranges of RRMSE can be defined to show the capability of the models, so that
model accuracy is:
64
Chapter III Results and Discussion
The correlation coefficient measures how close the predicted values are to the actual values.
Clearly, the value of the correlation coefficient more comparable to the unit implies a better
prediction.
𝒄𝒐𝒗(𝒚𝒊, 𝒚)
R= (III.6)
𝝈𝒚𝒊 𝝈𝒚
F- Coefficient Of Determination (R ²)
𝟏
∑𝒏
𝒋=𝟏(𝒚−𝒚𝒊 )²
𝑹𝟐 = 𝟏 − 𝒏 𝟏 𝒏 (III.7)
∑ 𝒚
𝒏 𝒋=𝟏
[Link] Forest
The hyperparameters of the random forest technique in 4 learning and testing time
65
Chapter III Results and Discussion
Random forest
𝟏𝒆𝒓 CAS 13.99 1.99 0.999 1.603 336 268.90 9.64 0.995 12.23 6.12
2é𝑚𝑒 CAS 88.69 4.92 0.999 4.177 149 364.81 11.22 0.993 14.25 1.98
3é𝑚𝑒 CAS 19.75 2.45 0.999 1.956 118 527.83 15.23 0.993 13.80 7.70
4é𝑚𝑒 CAS 13.69 2.08 0.999 1.62 351 315.35 11.34 0.995 10.66 15.64
Table. III. 4. Results were obtained from the random forest technical basis.
Next, we compared the actual outputs with the RF regression outputs in the learning and
testing base. We obtained the best results in the 1𝑠𝑡 the case according to the evaluation
criteria (R² =0.995, RRMSE=12.23, MAE = 9.64), Figure III.10-11 reports the correspondence
between the two test outputs in the 1𝑠𝑡 case
Fig. III. [Link] and calculated outputs for the random forest test.
66
Chapter III Results and Discussion
Fig. III. [Link] and calculated outputs for learning with random forest.
Fig. III. [Link] and calculated outputs for learning and testing with random forest.
67
Chapter III Results and Discussion
1𝑒𝑟 CAS 0.79 0.71 0.999 0.39 1.34 368.98 12.27 0.993 14.33 45.36
2é𝑚𝑒 CAS 0.78 0.48 0.999 0.39 0.30 195.69 8.45 0.996 10.44 18.00
𝟑é𝒎𝒆 CAS 12.09 0.97 0.999 1.54 0.40 183.41 8.04 0.996 10.10 49.19
4é𝑚𝑒 CAS 15.14 1.04 0.999 1.71 0.33 269.61 9.83 0.996 9.86 41.69
Table. III. 6. Results were obtained from the technical basis of the gradient boosting
regression.
Then we compared the actual outputs with the GBR regression outputs in the learning and
testing base. We obtained the best results in the 3𝑟𝑑 the case according to the evaluation
criteria (R²=0.996, RRMSE=10.10, MAE= 8.04) Figure III.13-14 reports the correspondence
between the two test outputs in the 3𝑟𝑑 case
Fig. III. [Link] and calculated outputs for the test with Gradient boosting regression.
68
Chapter III Results and Discussion
Fig. III. [Link] and calculated outputs for learning with Gradient boosting regression.
Fig. III. [Link] and calculated outputs for learning and testing with GBR.
69
Chapter III Results and Discussion
BI-LSTM
1𝑒𝑟 CAS 2.88 1.53 0.999 0.757 3h.45 1.53 0.98 0.999 1.13 5.03
m
𝟐é𝒎𝒆 CAS 0.61 0.59 0.999 0.34 1h.24 0.85 0.74 0.999 0.69 3.13
m
3é𝑚𝑒 CAS 3.05 1.38 0.999 0.763 37m.6 2.39 1.17 0.999 0.92 3.05
s
4é𝑚𝑒 CAS 8.89 1.93 0.999 1.31 1h.35 17.93 2.88 0.999 2.54 11.96
m
Table. III. [Link] were obtained from the Bi-LSTM technical basis.
Then we compared the actual outputs with the Bi-Directional Long Short Term Memory
regression outputs in the learning and testing base. We obtained the best results in the 2𝑛𝑑 the
case according to the evaluation criteria (R²=0.999, RRMSE=0.67, MAE = 1.24) Figure
III.19-20 reports the correspondence between the two test outputs in the 2𝑛𝑑 case.
Fig. III. [Link] and calculated outputs for the test with BI-LSTM.
70
Chapter III Results and Discussion
Fig. III. [Link] and calculated outputs for learning with BI-LSTM.
Fig. III. [Link] and calculated outputs for learning and testing with BI-LSTM.
Table. III. [Link] hyperparameters of the Deep Neural Network technical dataset.
71
Chapter III Results and Discussion
1𝑒𝑟 CAS 20.326 3.48 0.999 1.99 175.2 16.13 3.05 0.999 2.99 1.88
2é𝑚𝑒 CAS 241.01 10.00 0.997 6.88 126.3 104.76 6.37 0.998 7.63 1.36
3é𝑚𝑒 CAS 24.635 3.852 0.999 2.20 256.3 16.10 3.62 0.999 2.99 2.35
𝟒é𝒎𝒆 CAS 8.00 2.24 0.999 1.25 235.3 6.01 1.99 0.999 1.83 2.23
Table. III. 10. Results were obtained from the Deep Neural Network technical basis.
Then we compared the actual outputs with the Deep Neural Network regression outputs in
the learning and testing base. We obtained the best results in the 4𝑡ℎ the case according to the
evaluation criteria (R²=0.999, RRMSE=2.21, MAE = 8.77) Figure III.22-23 reports the
correspondence between the two test outputs in the 4𝑡ℎ case.
Fig. III. [Link] and calculated outputs for the test with Deep Neural Network.
72
Chapter III Results and Discussion
Fig. III. [Link] and calculated outputs for learning with Deep Neural Network.
Fig. III. [Link] and calculated outputs for learning and testing with Deep Neural Network.
Table. III. [Link] hyperparameters of the Long Short Term Memory technical dataset.
73
Chapter III Results and Discussion
Technique: LSTM
1𝑒𝑟 CAS 187.47 7.94 0.998 6.07 49m.6 213.61 8.68 0.995 12.19 5.23
s
2é𝑚𝑒 CAS 144.67 7.72 0.998 5.33 1h.49 165.51 7.95 0.997 12.86 7.98
m
𝟑é𝒎𝒆 CAS 6.40 1.66 0.999 1.12 3h.22 9.11 2.17 0.999 2.25 7.82
m
4é𝑚𝑒 CAS 33.99 4.15 0.999 2.58 2h.02 37.65 4.42 0.999 4.57 5.78
m
Table. III. [Link] obtained from the Long Short Term Memory technical basis.
Next, we compared the actual outputs with the Long Short Term Memory regression outputs
in the learning and testing base. We obtained the best results in the 3𝑟𝑑 the case according to
the evaluation criteria (R²=0.999, RRMSE=2.65, MAE = 12.62) Figure III.25-26 reports the
correspondence between the two test outputs in the 3𝑟𝑑 case
Fig. III. [Link] and calculated outputs for the LSTM test.
74
Chapter III Results and Discussion
Fig. III. [Link] and calculated outputs for learning with LSTM.
Fig. III. [Link] and calculated outputs for learning and testing with LSTM.
[Link] OF RESULTS
The table below makes a comparison of in terms of Training accuracy and Validation
accuracy with the same NASA HI-SEAS dataset.
𝐿𝑆𝑇𝑀 6.40 1.66 0.999 1.12 3h.2m 9.11 2.17 0.999 2.25 7.82
𝐷𝑁𝑁 8.00 2.24 0.999 1.25 235.3 6.01 2.24 0.999 1.25 2.23
𝑩𝑰. 𝑳𝑺𝑻𝑴 0.61 0.59 0.999 0.34 1h.2m 0.85 0.74 0.999 0.69 3.13
𝐺𝐵𝑅 12.09 0.97 0.999 1.54 0.40 183.41 8.04 0.996 10.10 49.19
RF 13.99 1.99 0.999 1.603 336 268.90 9.64 0.995 12.23 6.12
75
Chapter III Results and Discussion
Based on the results obtained from Tables III.14, the following can be noted.
One of these results is that the BI-LSTM method One of the best ways used to predict solar
radiation using artificial intelligence
[Link]
Throughout this work, it was possible to learn more about artificial intelligence techniques,
in particular about the ML and DL models and how these models could be applied to solar
irradiance Prediction.
After this, two different algorithms were developed using the studied LSTM and DNN
methods to perform solar irradiance predictions. Before the final tests, the models were
adjusted to perform the best predictions.
76
Conclusion
Conclusion
As we have seen in the introduction to this thesis , the Solar system sizing requires reliable,
comprehensive and long-term solar radiation data. To overcome the lack of long series of
reliable solar data, necessary for the optimization and optimal sizing of solar systems, an
approach to predict hourly global solar radiation from cheaper meteorological data was presented
in this manuscript. Ten different associations of five meteorological variables were used for
develop two types of ANN models. The model that gives the best performance is a neuronal
autoregressive with external BI-LSTM inputs with eight inputs. It was used to predict hourly
solar radiation from more available and cheaper weather data (Solar radiation, Humidity, Wind
direction, Time at sunrise, Temperature ,Barometric pressure ,Wind speed ,Time at sunset). The
prediction accuracy of the proposed model is about 0.64% with a R² value of 0.999. In addition,
different learning samples were used, the results showed that the proposed model required
concrete training beforehand in order to give an acceptable predictive accuracy.
Future research could include a model that re-evaluates the times for each region and updates the
input masks daily; such a model could improve the accuracy of the morning and dusk time
region, as changes in season and obstacles affect dawn and dusk.
In conclusion, our approach provides radiation series synthetic solar to be used in the optimal
sizing and planning of solar energy systems.
68
BIBLIOGRAPHY
BIBLIOGRAPHY
[2] Xiangyang Yea, Qing T. Zenga,b , Julio C. Facellia , Diana I. Brixnerc , Mike Conwaya ,
Bruce E. Bray Predicting Optimal Hypertension Treatment Pathways Using Recurrent Neural
Networks. International Journal of Medical Informatics 2020.
[3] Chunchun Chena , Pu Zhang b , Yuan Liuc , Jun Liud, Financial quantitative investment
using convolutional neural network and deep learning technology, Neurocomputing ,2019
[4] Mohammad Mehedi Hassan , Abdu Gumaei , Ahmed Alsanad , Majed Alrubaian , Giancarlo
Fortino , A Hybrid Deep Learning Model for Efficient Intrusion Detection in Big Data
Environment , Information Sciences,2019
[5] A. Khosravi* , R.N.N. Koury, L. Machado, J.J.G. Pabon , Prediction of hourly solar radiation
in Abu Musa Island using machine learning algorithms , Journal of Cleaner Production ,2018
[6] Junho Lee, Wu Wang, Fouzi Harrou⁎ , Ying Sun , Reliable solar irradiance prediction using
ensemble learning-based models: A comparative study, Energy Conversion and Management,
2020
[7] Jianwu Zeng, Wei Qiao , Short-term solar power prediction using a support vector machine ,
Renewable Energy,2013
[8] Nor Azuana Ramli, Mohd Fairuz Abdul Hamid, Nurul Hanis Azhan, and Muhammad Alif
As-Siddiq Ishak , Solar power generation prediction by using k-nearest neighbor method , AIP
Conference Proceedings 2019
[9] Chao-Rong Chen and Unit Three Kartini , k-Nearest Neighbor Neural Network Models for
Very Short-Term Global Solar Irradiance Forecasting Based on Meteorological Data , 2017
[10] J Liu1 , M Y Cao2 , D Bai1 and R Zhang, Solar radiation prediction based on random forest
of feature-extraction ,2019
[11] JürgenSchmidhuber, Deep learning in neural networks: An overview, Neural Network ,2015
[12] S. Ghimire, R. C. Deo, N. Raj, J. Mi. Deep solar radiation forecasting with convolutional
neural network and long short-term memory network algorithms. Appl. Energy
2019;253:113541.
[13] Z. Pang, F. Niu, Z. O'Neill. Solar radiation prediction using recurrent neural network and
arti_cial neural network: A case study with comparisons. Renew. Energy 2020;156:279-289.
[14] D. Guijo-Rubio et al. Evolutionary artificial neural networks for accurate solar radiation
prediction. Energy 2020
[15] Qing X, Niu Y. Hourly day-ahead solar irradiance prediction using weather forecasts by
LSTM. Energy 2018;148:461–8.
[16] Kaba K, Sarيgül M, Avc يM, Kandيrmaz HM. Estimation of daily global solar radiation
using deep learning model. Energy 2018;162:126–35.
[17] Kira M. Sargent « Supporting Renewable Energy: Lessons from the Deer Island Treatment
Plant » Doctoral thesis, Washington University in St. Louis Environmental Studies Program,
Spring 2010 St. Louis, Missouri
[19] M.A. Atwater, J.T. Ball, “A numerical solar radiation model based on standard
meteorological observations”, Sol. Energy 21, pp. 163–170, 1978
[20] TRAHI Fatiha « Prédiction de l’irradiation solaire globale pour la région de Tizi-Ouzou
par les réseaux de neurones artificiels. Application pour le dimensionnement d’une installation
photovoltaïque pour l’alimentation du laboratoire de recherche LAMPA. » MEMOIRE DE
MAGISTER EN ELECTRONIQUE , Université Mouloud Mammeri de Tizi-Ouzou
[21] Lila Croci « Gestion de l’énergie dans un système multi-sources photovoltaïque et éolien
avec stockage hybride batteries/supercondensateurs » Doctoral thesis, L’UNIVERSITE DE
POITIERS, Submitted on 7 Feb 2014
[22] Saad Motahhir, Abdelaziz El Ghzizal, Aziz Derouich « Modélisation et commande d’un
panneau photovoltaïque dans l’environnement PSIM » Doctoral thesis, Submitted on 19 Apr
2018
[23] Rene Y. Choi; Aaron S. Coyner; Jayashree Kalpathy-Cramer; Michael F. Chiang; J. Peter
Campbell, Introduction to Machine Learning, Neural Networks, and Deep Learning,
Translational Vision Science & Technology , February 2020
[24] Adrien BELLEVILLE « Prédiction des S/P individuels sur le produit flottes automobiles »
Mémoire présenté le : pour l’obtention du Diplôme Universitaire d’actuariat de l’ISFA et
l’admission à l’Institut des Actuaires
[25] Shalev-Shwartz, S. & Ben-David, S., “Understanding Machine Learning: From Theory to
Algorithms”, Cambridge University Press, 2014
[26] Morgane LAUR « Anticipation des changements de notes des obligations du portefeuille
d’un assureur par méthode de machine learning » Mémoire présenté devant l’Université Paris
Dauphine pour l’obtention du diplôme du Master Actuariat et l’admission à l’Institut des
Actuaires
[27] Classification And Regression Trees for Machine Learning. Jason Brownlee. April 8, 2016
URL:[Link]
learning/
[28] Alexey Natekin, Alois Knoll, Gradient boosting machines, a tutorial, fortiss GmbH,
Munich, Germany, NEUROROBOTICS, Gradient boosting machines, a tutorial,2013
[29] Complete Machine Learning Guide to Parameter Tuning in Gradient Boosting (GBM) in
Python, AARSHAY JAIN, FEBRUARY 21, 2016, URL :
[Link]
boosting-gbm-python/
[30] Sepp Hochreiter Fakultat “LONG SHORT-TERM MEMORY” Technische Universitat
Munchen 1997
[31] Mike Schuster and Kuldip K Paliwal, “Bidirectional recurrent neural networks,” Signal
Processing, IEEE Transactions, vol. 45, no. 11, pp. 2673–2681, 1997
[32] Moualek Djaloul Youcef « Deep Learning pour la classification des images » Mémoire de
fin d’études pour l’obtention du diplôme de Master en Informatique 2016
[33] Haohan Wang ,Bhiksha Raj “On the Origin of Deep Learning” Article · February 2017
[34] HAWKINS Jeff, BLAKESLEE Sandra. On intelligence : how a new understanding of the
brain will lead to the creation of truly intelligent machines. 2007
[35] Meriem Bahi , Mohamed Batouche , Deep Learning for Ligand-Based Virtual Screening in
Drug Discovery , University Constantine-2 Abdelhamid Mehri Constantine, Algeria ,2018
[36] Chigozie Enyinna Nwankpa, Winifred Ijomah, Anthony Gachagan, and Stephen Marshall ,
Activation Functions: Comparison of Trends in Practice and Research for Deep Learning ,2018
[37] Adam — latest trends in deep learning optimization. Vitaly Bushaev. Oct,22,2018. URL :
[Link]
[40] Understanding LSTM Networks. Oinkina . Posted on August 27, 2015 . URL :
[Link]
[41] Huimei Han , Xingquan Zhu ,Ying Li , Generalizing Long Short-Term Memory Network
for Deep Learning from Generic Data , 2020
[42] Z. Cui, S. Member, R. Ke, S. Member, and Y. Wang, “Deep Stacked Bidirectional and
Unidirectional LSTM Recurrent Neural Network for Network-wide Traffic Speed Prediction,”
pp. 1–12, 2018.
[43] Daniel Durstewitz , Georgia Koppe, Andreas Meyer, Lindenberg Deep neural networks in
psychiatry,2019
[46] Google Colab Free GPU Tutorial .Fuat . Jan 26, 2018 .URL : [Link]
learning-turkey/google-colab-free-gpu-tutorial-e113627b9f5d
This thesis presents a prediction study of the different components of the solar radiation using
artificial neural networks (ANN). The results of this study are crucial for the design and sizing of
any solar energy system. A series of experimental hourly measurements of year variables were
available for this study. Models (ANN) with different structures, in particular, different
combinations of inputs as well as other numbers of hidden neurons, have been set up. To
evaluate these models, the regression coefficient (R2) and the error estimators Relative Root
Mean Square Error (RRMSE) and Mean Square Error (MSE) were used. Random Forest (RF)
and Gradient Boosting Machine (GBM) and Deep Neural Network (DNN), Long Short Term
Memory (LSTM) were compared with Bidirectional LSTM to generate horizontal hourly global
solar radiation from less expensive exogenous variables. The results show BI-LSTM superiority
with 8 entries. The test of this model to produce accurate forecasts offers good accuracy
(R²=0.999, MSE = 0.85 and RRMSE = 0.69%). Using different sizes of the learning sample we
showed that from one year of data, our model gives satisfactory results. A comparison of our
results with the literature confirmed that our models (ANN) exceed other estimation methods and
that the proposed models ensure an authentic prediction of the different components of hourly
solar irradiation from endogenous and exogenous variables that are more available and less
expensive.
Keywords: ANN, solar power, forecasting, renewable energy, machine learning, deep learning.
Résumé
Cette mémoire présente une étude prédictive des différentes composantes de la rayonnement
solaire utilisant des réseaux neuronaux artificiels (RNN). Les résultats de cette étude sont
cruciaux pour la conception et le dimensionnement de tout système d’énergie solaire. Une série
de mesures horaires expérimentales de variables de huit ans était disponible pour cette étude. Des
modèles (RNN) avec différentes structures, en particulier, différentes combinaisons d’entrées
ainsi que différents nombres de neurones cachés ont été mis en place. Pour évaluer ces modèles,
on a utilisé le coefficient de régression (R2) et les estimateurs d’erreur Erreur quadratique
moyenne relative (RRMSE) et Erreur quadratique moyenne (MSE). On a comparé Random
Forest (RF) et Gradient Boosting Machine (GBM) et Deep neural Network (DNN), Long Short
Term Memory (LSTM) avec Bidirectional LSTM pour générer un rayonnement solaire mondial
horaire horizontal à partir de variables exogènes moins coûteuses. Les résultats montrent une
supériorité BI-LSTM avec 8 entrées. Le test de ce modèle pour produire des prévisions
authentiques montre une bonne précision (R²=0,999, MSE = 0,85 et RRMSE = 0,69 %). En
utilisant différentes tailles de l’échantillon d’apprentissage, nous avons montré qu’à partir d’une
année de données, notre modèle donne des résultats satisfaisants. Une comparaison de nos
résultats avec la littérature a confirmé que nos modèles (RNN) dépasser les autres méthodes
d’estimation et que les modèles proposés assurent une prédiction authentique des différentes
composantes de l’irradiation solaire horaire à partir de variables endogènes et exogènes plus
disponibles et moins coûteuses.
ملخص
تقدم هذه الرسالة دراسة تنبؤ للمكونات المختلفة لإلشعاع الشمسي باستخدام شبكات عصبية اصطناعية (RNA).إن نتائج هذه
الدراسة تشكل أهمية حاسمة في تصميم أي نظام للطاقة الشمسية وتحجيم حجمه .وقد توفرت لهذه الدراسة سلسلة من القياسات
التجريبية بالساعة لمتغيرات ثماني سنوات .وتم إعداد نماذج ) (RNAذات هياكل مختلفة ،وخاصة تركيبات مختلفة من
المدخالت فضالً عن أعداد مختلفة من الخاليا العصبية المخفية .لتقييم هذه النماذج ،تم استخدام معامل االنحدار ) (R2ومقدرات
األخطاء النسبة للخطأ المتوسط الحسابي للجذر المتوسط ) (RMSEوخطأ مربع المتوسط (MSE).تمت مقارنة الغابة
العشوائية ) (RFوماكينة تعزيز التدرج ) (GBMوالشبكة العصبية العميقة) ، (DNNالذاكرة طويلة األجل ) (LSTMمع
تقنية LSTMثنائية االتجاه لتوليد إشعاع شمسي عالمي بالساعة أفقي من المتغيرات الخارجية األقل تكلفة .وتظهر النتائج تفوق
بي-إل تي سي مع 8مداخل .إن اختبار هذا النموذج إلنتاج توقعات حقيقية يبين دقة جيدة،MSE = 0.85 ، (R²=0.999
RMSE = 0.69%).باستخدام أحجام مختلفة من عينة التعلم التي أظهرنا أنه من عام واحد من البيانات يقدم نموذجنا نتائج
مرضية .وقد أكدت مقارنة النتائج التي تحققناها بالمؤلفات أن نماذجنا تتجاوز أساليب التقدير األخرى وأن النماذج المقترحة
تضمن التنبؤ الحقيقي بالمكونات المختلفة لإلشعاع الشمسي بالساعة من المتغيرات الداخلية والخارجية األكثر توافرا ً وأقل تكلفة .
الكلمات المفتاحية :شبكة العصبية اسطناعية ,الطاقة الشمسية ,الطاقات المتجددة ,تعلم اآللة ,تعلم العميق
Hyperparameter tuning is crucial in optimizing the performance of machine learning models like Random Forest and Gradient Boosting Machines for solar irradiance prediction. This process involves adjusting model parameters such as the number of estimators, maximum depth of trees, and learning rates, among others, to improve the model's predictive power. Proper tuning ensures that the model does not overfit or underfit the data, enabling it to generalize well to unseen data. For instance, tuning the learning rate can balance the model's convergence speed with accuracy. Therefore, hyperparameter tuning is critical for achieving the best possible performance from these models in solar irradiance prediction tasks .
Data availability and quality are crucial in the successful modeling of solar radiation predictions using machine learning. High-quality datasets provide the necessary information for training and validating models, ensuring their predictions are accurate and reliable. In the field of solar radiation prediction, comprehensive datasets that include accurate weather-related variables over long periods are vital for capturing the temporal dynamics needed by models like LSTM and BI-LSTM to make precise forecasts. Poor quality or insufficient data can lead to inaccurate models that fail to generalize well, highlighting the importance of robust data preprocessing and selection in the modeling process .
The Long Short-Term Memory (LSTM) network is particularly suitable for forecasting tasks like solar irradiance prediction due to its ability to capture and retain temporal dependencies in time-series data. LSTMs are designed to overcome the limitations of traditional RNNs by preventing long-term dependency problems, enabling the network to learn patterns across various time frames within the data. This ability makes LSTMs ideal for capturing the cyclical and seasonal patterns often found in solar irradiance, leading to more accurate and reliable forecasts .
The primary challenges associated with the unpredictability of solar irradiance include its variability due to changing weather conditions and atmospheric factors, which directly impact solar energy production. Machine learning assists in overcoming these challenges by employing sophisticated models that can analyze large volumes of weather data to identify patterns and predict future solar irradiance more accurately. Techniques such as neural networks, including LSTM and BI-LSTM, leverage temporal patterns in the data to improve the accuracy of predictions, allowing for better management and integration of solar energy into power systems and reducing operational costs .
Integrating multiple machine learning models can significantly enhance the accuracy of solar irradiance predictions. By combining or comparing results from different models, such as Random Forest, Gradient Boosting Machines, and deep learning models like LSTM and CNN, it's possible to leverage the strengths of each approach. Ensemble methods can help mitigate individual model weaknesses, thus yielding more accurate and robust predictions. For instance, combining the spatial recognition capabilities of CNNs with the temporal sequence learning of LSTMs provides a well-rounded forecasting tool, leading to improved prediction performance .
Machine learning has significantly improved solar radiation prediction by providing advanced models like support vector machines, k-nearest neighbors, and deep learning models, including convolutional neural networks and Long Short-Term Memory networks. These models enable more accurate and reliable forecasts of solar irradiance, which is critical for optimizing the operation and integration of solar power within energy grids. For example, deep learning models such as LSTM are designed to capture time dependencies in data, leading to more precise day-ahead solar irradiance forecasts. This enhanced capability allows better planning and operation of power delivery, minimizing operating costs and improving the integration of renewable energy into the grid .
R² (coefficient of determination) and MAE (Mean Absolute Error) are significant metrics for evaluating solar radiation prediction models. R² measures how well the model's predictions match the actual data, with a value closer to 1 indicating better conformity. A high R² value signifies that the model captures the variability within the data well. MAE, on the other hand, indicates the average magnitude of errors in the predictions, with lower values representing more precise predictions. Together, these metrics provide a comprehensive view of a model's accuracy and reliability; high R² and low MAE values are desired outcomes in model evaluations .
Deep learning techniques differ from traditional machine learning methods primarily in their ability to handle complex and high-dimensional data by using multi-layer neural networks. These networks are capable of automatically extracting features from raw data, allowing for greater accuracy in tasks like solar radiation prediction. Whereas traditional methods like Random Forest and Gradient Boosting Machines rely on manually engineered features and simpler models, deep learning approaches such as CNNs and LSTMs can model non-linear relationships and capture dependencies over time without requiring extensive preprocessing. This allows deep learning models to achieve higher prediction accuracy and performance in solar radiation forecasting, adapting better to varying data conditions .
Bidirectional LSTM (BI-LSTM) offers the advantage of accessing context information in both forward and backward directions, making it superior to traditional unidirectional LSTM, which processes data in a single direction. This allows BI-LSTM to construct a more comprehensive understanding of the time dependencies present in the solar irradiance data, improving the accuracy of predictions. Studies have shown that bidirectional models perform significantly better in various domains, including solar irradiance prediction .
Convolutional neural networks (CNNs) contribute to the two-stage deep learning approach for solar radiation prediction by serving as the initial stage where they are utilized to extract relevant features from the input data. This feature extraction process is critical as it reduces the dimensionality and highlights the key aspects of the data that are necessary for accurate predictions. Once the CNN has identified these features, a subsequent network such as a Long Short-Term Memory (LSTM) network takes over to make the final predictions. This combination allows for a robust model structure that first understands the input data's spatial features and then captures temporal dependencies for precise solar radiation forecasting .