Remote Sensing for Landscape Management
Remote Sensing for Landscape Management
MOOC COURSE
Contents
M3T1. Remote sensing Observation ............................................................................................. 3
M3T2. Climate modelling .............................................................................................................. 9
M3T3. Land Use and Land Cover Change modelling ................................................................... 13
M3T4. Hydrological characterisation and models ...................................................................... 21
M3T5. Erosion models................................................................................................................. 36
M3T6. Water temperature models ............................................................................................. 42
M3T7. Water quality models....................................................................................................... 48
2
Module 3. Databases and modelling
M3T1.
Remote sensing Observation
Teacher: Ana Silio Calzada
Introduction
In this section, we will explore the fundamentals of optical remote sensing, and its application
in monitoring the environment. We will also discuss the advantages and limitations of these
techniques and how they can be applied to address various research questions.
The first aspect that needs to be addressed is what we understand by optical remote sensing,
which is no other than the science of acquiring information about an object, an area or
phenomena without physical contact, through the analysis of electromagnetic radiation on the
visible and near-infrared regions of the spectrum. It relies on sensors mounted on satellites,
airborne or other platforms that detect and measure the amount and quality of light reflected
or emitted from the Earth's surface.
First, optical sensors capture the electromagnetic radiation, the energy that travels through
space in the form of electromagnetic waves. When electromagnetic radiation interacts with the
Earth's surface, it can be reflected or absorbed. Reflectance refers to the fraction of incident
radiation that is reflected back, while absorption occurs when radiation is absorbed by the
surface materials.
3
Module 3. Databases and modelling
Different materials and features interact with light in unique ways, leading to variations in the
measured signals, with a distinctive pattern of reflectance or emission at different wavelengths,
known as Spectral Signatures.
The temporal variability of spectral signatures is known as spectral behaviour, a concept that
includes as well their variation as a function of weather conditions, seasons, and illumination
conditions.
The analysis of such information, or spectral analysis, helps identify and differentiate land cover
types, vegetation species, and other environmental components.
Thus, to be sure that these spectral signatures are properly collected, remote optical sensors
must be designed accordingly. In this sense, not all sensors are suitable for observing everything,
but have a primary mission involving sensors with specific characteristics that have an impact
4
Module 3. Databases and modelling
on data collection and analysis. Sensor parameters such as spatial resolution, spectral resolution
and radiometric resolution influence the level of detail and accuracy with which surface features
are captured. Understanding sensor characteristics is essential for selecting appropriate data
sources and interpreting the acquired data.
Last but not least, the approaches followed for image processing and analysis are fundamental,
as this series of pre-processing steps are crucial to obtain reliable and comparable information
over different time periods and study areas. These steps include:
Data is processed and later analysed using image processing techniques, spectral indices, and
classification algorithms. These methods help extract information about land cover, vegetation
health, water quality, and other environmental parameters.
5
Module 3. Databases and modelling
To further extract meaningful information from optical remote sensing data, advanced data
integration and analysis techniques are employed. This involves combining multispectral,
hyperspectral, and LiDAR data with other datasets, such as climate data, and ground-based
measurements.
Data fusion, machine learning algorithms, and spatial modelling techniques are used to analyse
and interpret these complex datasets, providing the tools to make informed decisions in terms
of management and conservation.
You will see more about this in the following sections of this online course.
6
Module 3. Databases and modelling
Focussing on the work addressed in ALICE project, we are going to go through some examples
about vegetation-related mapping, which plays a crucial role in catchment ecosystems,
influencing the hydrological cycle, biodiversity, and overall ecosystem health.
Land use/land cover mapping represents one of the most common examples of applied optical
remote sensing in ecosystem monitoring.
The combined analysis of spectral signatures and indexes, such as the Normalized Difference
Vegetation Index (NDVI) and the Enhanced Vegetation Index (EVI), is frequently used to identify
and classify different land cover types, such as forests, grasslands, wetlands, and urban areas.
Furthermore, the detection of land use/land cover changes over time is crucial for
understanding landscape dynamics, and the environmental impact of climate change and
human activities on the landscape, while providing the tools to support sustainable land
management practices.
Linked to this, but going a step further, biodiversity assessment is another important application
of optical remote sensing in continental ecosystem monitoring. Spectral signature analysis,
together with vegetation patterns, climate data, and ground-truthing training data, allow to
estimate species richness, identify biodiversity hotspots, and monitor changes in habitat quality.
These assessments contribute to conservation planning, biodiversity management and the
identification of areas requiring protection or restoration efforts.
7
Module 3. Databases and modelling
Corolarium
To sum up, optical remote sensing techniques offer several advantages over traditional ground-
based methods. These techniques can cover large areas quickly and efficiently, providing
spatially and temporal continuous data. They can also provide information on hard-reach areas,
and inaccessible terrains.
Nevertheless, despite these advantages, there are challenges to address. These include
atmospheric effects on data quality, the need for high spatial and temporal resolution imagery,
the integration of multi-sensor data, and the need for a clear line of sight to the Earth’s surface,
which can be challenging in areas with dense vegetation or steep topography.
Now, we encourage you to continue to explore and utilize remote sensing techniques for a
better understanding of the environment.
8
Module 3. Databases and modelling
M3T2.
Climate modelling
Teacher: André Fonseca
Introduction
In this lesson I will explain climate modeling concept based on statistical downscaling and bias
correction of data to project future climate data.
IPCC
The Intergovernmental Panel on Climate change most commonly known as IPCC, is a scientific-
political organization created in 1988 within the framework of the United Nations (UN) on the
initiative of the United Nations Environment Program (UNEP) and the World Meteorological
Organization (WMO).
Its main objective is to synthesize and disseminate the most advanced knowledge about the
climate changes that affect the world today, specifically global warming, pointing out its causes,
effects and risks to humanity and the environment.
There are several factors that affect climate change from Forcing agents to Actors. Major forcing
agents are: Greenhouse gases, Aerosols, change in Land Use and Land Cover, Solar and Volcano
activity. While actors can be considered as People, Industry, Agriculture, Urbanization and
Vehicles.
9
Module 3. Databases and modelling
Globally, there has been an increase of around 1.1°C since the pre-industrial era.
Shared Socioeconomic Pathways (SSPs). In the lowest emissions scenario, the light blue line,
temperatures may reach 1.4 Celsius above baseline levels at the end of the century, whereas
they may climb 4.4 Celsius under worst scenario, the SSP5 which is the dark red line.
10
Module 3. Databases and modelling
In this context of regional scale changes to assess change in climate we cannot simply rely on
General Circulation Models but couple them with Regional Climate Models.
To develop future climate data that represent more accurately local conditions, an ensemble of
couple General Circulation Models with Regional Climate Models is commonly used to reduce
model’s uncertainty. These models available from the IPCC will need further tailoring. By
performing statistical downscaling techniques and a bias-correction with respect to local
observed data, allows the increase of resolution to a finer scale. The new high-resolution data is
then ready to be used to perform studies on a local scale such as: hydrological modelling, climate
change impacts and future assessments that provide stakeholders with tools for decision
support and adaptation.
For ALICE Project, these datasets were developed for all 4 case studies, that includes Portugal,
Spain, France and Northern Ireland and United Kingdom.
This example shows the particular case study of Paiva River in Portugal. First statistical
downscaling techniques are applied to an Observed Gridded Dataset to increase the resolution
from approximately 10 km to approximately 1 km spatial resolution. To further assess future
data, bias correction is performed to the previously mentioned coupled General Circulation
Models and Regional Climate Models based on the newly created data and thus project climate
data for the future. In this example for two Representative Concentration Pathways 4.5 and 8.5.
As predictable due to the ongoing climate change, mean temperature is expected to increase in
the Paiva River region. While historical data shows a range between 10 and 15 degrees Celsius,
future assessment shows and increase up to 18 degrees Celsius in the region.
11
Module 3. Databases and modelling
Image credits
Earth right now – How much warmer will your future be.
[Link]
12
Module 3. Databases and modelling
M3T3.
Land Use and Land Cover Change modelling
Teacher: Thomas Houet
Introduction
In this section, we are going to talk about modelling land use and land cover changes.
Context: Why?
Because it allows to explore the future and better anticipate some undesirable changes and plan
today more sustainable land strategies.
For example, we can account for one Blue and Green Infrastructure strategy and imagine
different evolutions of the agriculture. Then we can evaluate the impacts of the resulting land
use changes on different variables such as landscape connectivity and water quality.
13
Module 3. Databases and modelling
Implicating stakeholders and local actors in the co-design of scenarios allows to account for local
specificities and improve simulations’ realism.
Through an iterative process, stakeholders provide key elements that will be used as parameters
settings of the modelling, they validate simulations which are then used to evaluate potential
environmental impacts before disseminating all results.
To simulate land use and land covers two main categories of models exist: the cellular automata
and the agent-based models.
The former are using rules to modify the colour of cells representing land uses or land covers.
The latter add agents that can interact between them and with their environment complexifying
rules that determines land use changes.
There is a wide diversity of models that combine both approaches and any others modelling
tools corresponding to a so-called hybrid category.
14
Module 3. Databases and modelling
Models can also be categorized according to the scale they are dealing with. Fine-scale models
aim to simulate land use and land cover changes with geographical elementary units such as
parcels that allow a realistic landscape representation. They often do not cover wide areas.
Large-scale models aim at simulating land use and land cover changes over large areas, but their
resolution is often quite coarse with cells of few hundred meters or kilometers width.
Lastly, neutral models are based on simple and unrealistic landscape where the effect of specific
land use and land covers rules and interaction can be analysed.
15
Module 3. Databases and modelling
In ALICE, we developed an innovative hybrid fine scale model called FORESCEM that can basically
simulate business-as-usual scenarios thanks to a calibration phase and suitability maps like most
of existing LUCC models. It can also consider interactions between land uses and land covers
changes, bifurcations, new land planning strategies while preserving landscape patterns.
16
Module 3. Databases and modelling
It also uses maps defining where some land use and land covers changes are not possible which
are commonly called, excluded maps. For instance, these maps groups all inundated and natural
reserve areas.
17
Module 3. Databases and modelling
- the quantity of changes, also called the land demand for each land cover that could be
a linear projection over the future based on past changes or non-linear based on expert
knowledge or any external database;
- the future land planning strategies that are translated into excluded or regulated areas
and when they will take place.
18
Module 3. Databases and modelling
Last, but not least, it simulates at each time step for one land use (for instance agriculture) the
possible land cover changes. For agricultural land use change, the model simulates the crops
rotation at each time step accounting for constrains on the crops’ duration and succession.
For example, if at time 1 a land abandonment has been simulated on one pixel, the model will
simulate shrublands on that pixel for the next 7-time steps (or more, according to the modeler
set rules) before turning into a forest. After a specific duration, a shrubland would evolve into a
forest.
FORESCEM thus integrates all interactions between agricultural, urban, forestry and even
abandoned land uses.
Simulations
The final example applied on the French case study shows the simulation of land use and land
cover changes where agricultural crops rotations evolve under the assumption of an
intensification of dairy production inducing a slight increase of maize (in orange) and wheat and
rapeseeds (in yellow), and where urban sprawling does not affect Blue and Green Infrastructure.
Conclusion
19
Module 3. Databases and modelling
To conclude, we can emphasize that modelling land use and land cover changes is a critical step
for integrated assessment studies as it constitutes the input for evaluating the possible impacts
of anthropogenic changes.
Depending of the objectives of the study simple or complex models can be used:
- Simple models can facilitate the stakeholders’ engagement throughout the provision of
new knowledge.
- Simulating more realistic projections needs more complex models which are time-
consuming to develop, describe and use but they are effective to evaluate
environmental impacts considering local specificities.
20
Module 3. Databases and modelling
M3T4.
Hydrological characterisation and models
Teacher: Cristina Prieto Sierra
Introduction
In this section I will provide a flavor about hydrological modeling set up.
I will focus in lumped conceptual hydrological models that work at catchment scale.
Motivation
However, we only have a number of streamflow measurements that are limited in space and
time.
Therefore, we need hydrological models that simulate hydrological responses as a surrogate for
the data.
21
Module 3. Databases and modelling
Furthermore, decision making in a context of fluctuating weather patterns from year to year and
increasing demands on water resources throughout the world requires improved hydrological
models.
In a scientific context, the aim is to identify a model structure reflecting the relevant hydrological
processes of interest in a catchment.
22
Module 3. Databases and modelling
In an operational context, the aim is to provide reliable predictions of streamflow sometime into
the future (where measurements are not possible), for example, from minutes/hours (during
floods) to weeks/months (water supply); or in ungauged catchments (where measurements are
not available).
23
Module 3. Databases and modelling
In the next slides I will reflect the general steps in the modeling procedure.
The current slide provides a summary of the steps and next slides provide further explanation.
24
Module 3. Databases and modelling
25
Module 3. Databases and modelling
26
Module 3. Databases and modelling
27
Module 3. Databases and modelling
3. Procedural model are the techniques of numerical analysis to define the code that will
run in a computer.
The transformations from the equations of the conceptual model to the code of the
procedural model has the potential to add error to the true solution of the original
equations.
28
Module 3. Databases and modelling
All the models used in hydrology have equations that involve a variety of different input
and state variables.
There are variables that define the time-variable boundary conditions during a
simulation, such as the rainfall.
There are the state variables, such as soil water storage that change during a simulation
as a result of the model calculations.
There are the initial values of the state variables that define the state of the catchment
at the start of a simulation.
Finally, there are the model parameters that define the characteristics of the catchment
area or flow domain. For example, the mean residence time in the saturated zone. They
are usually considered constant during the period of a simulation.
It is not easy to specify the values of the parameters for a particular catchment a priori.
Therefore, parameter calibration is needed.
29
Module 3. Databases and modelling
30
Module 3. Databases and modelling
31
Module 3. Databases and modelling
32
Module 3. Databases and modelling
6. Performance metrics
The last step is evaluating the performance of the predictions. We can evaluate the
performance via different metrics. A widely used set of metrics are:
33
Module 3. Databases and modelling
34
Module 3. Databases and modelling
Conclusions
Certainly, it should not be assumed that there is one model representative of all catchments
(that is, the “one size fits all”) but the “uniqueness of the place”. A hydrological model should
be tested for its adequacy in each catchment.
Image credits
Figures in slides 2 and 3: 11thWorld Congress on Water Resources and Environment (EWRA) 25-
29 June 2019, Managing Water Resources for a Sustainable Future: Prieto, C., Le Vine, N.,
Kavetski, D., Álvarez, A., and Medina, R. Towards reducing model error in flow predictions in
ungauged basins via a Bayesian approach. Oral contribution.
Figures in slide 4 and 5: EGU 2022, Prieto, C., Le Vine, N., Kavetski, D., Fenicia, F., Scheidegger,
A., and Vitolo, C.: An exploration of Bayesian identification of dominant hydrological
mechanisms in ungauged catchments, EGU General Assembly 2022, Vienna, Austria, 23–27 May
2022, EGU22-6212, doi: 10.5194/egusphere-egu22-6212, 2022. Oral contribution. EGU
Highlighted (public interest)
Figures in slides 6, 8, 10, 12-16 2: Figure 1.2. A schematic outline of the steps in the modelling
processes from Beven, K. J. (2011). Rainfall-runoff modelling: the primer. John Wiley & Sons.
Figures in slides 9, 11, 14: Clark, M. P., Slater, A. G., Rupp, D. E., Woods, R. A., Vrugt, J. A., Gupta,
H. V., Wagener, T., and Hay, L. E. (2008), Framework for Understanding Structural Errors (FUSE):
A modular framework to diagnose differences between hydrological models, Water Resour.
Res., 44, W00B02, doi:10.1029/2007WR006735.
Figure in slide 14: addapted from Fenicia, F., Kavetski, D., and Savenije, H. H. G. (2011), Elements
of a flexible approach for conceptual hydrological modeling: 1. Motivation and theoretical
development, Water Resour. Res., 47, W11510, doi:10.1029/2010WR010174.
35
Module 3. Databases and modelling
M3T5.
Erosion models
Teacher: Ignacio Pérez Silos
Introduction
In this lesson, we are going to explore and exemplify how we can model erosion in the river
catchment context.
Erosion models
As we saw in module two, water erosion encompasses seven different types of erosion
processes. Nowadays there are a multitude of erosion models, however, no single model
considers all of them holistically. Most models focus, at best, on solving together those types of
erosion that are related to each other by the nature of the physical process. In this block we
focus on those processes that are modelled most frequently: rill erosion, gully erosion and slope
failures.
Currently available models range from simplistic expressions to highly articulated models
capable of considering a large number of interacting factors and physical relationships. Among
the first, the Universal Soil Loss Equation and its modifications are by far the most widely used
models. They propose statistical relationships between empirical observations of soil loss,
rainfall erosivity and soil type, corrected using information on slope and vegetation cover
properties. Its derived semi-empirical models attempt to integrate simple equations describing
other processes such us sediment transport. Finally, physical based models consist on algorithms
derived from theoretical principles that aim to comprehensively represent from soil erosion to
transport and deposition processes. They seek to better represent spatio-temporal distributions
of soil loss, ground conditions and improve their transferability to a wider set of environments.
36
Module 3. Databases and modelling
Now we are going to show how we have modelled erosion in the ALICE project.
In this respect, we base our methodology on the NETMAP software. NETMAP includes an index
for modelling erosion using three main inputs: mean annual precipitation, a digital elevation
model and vegetation maps.
However, the most advantageous property of NETMAP is the ability to create a virtual
watershed: a digital representation of the characteristics of the river basin and its channel
network in which to simulate the spatio-temporal interactions between the terrestrial and
fluvial environments.
37
Module 3. Databases and modelling
According to our approach, we first model the sediment potentially mobilised at the source. In
this sense, the GEP index allows us to obtain for each pixel its potential susceptibility to erosion,
which considers only abiotic criteria such as slope and topographic convergence. NETMAP also
determines the probability that the sediment generated in each pixel will be delivered to the
fluvial network. We integrate this information into a single layer over which we intersect the
vegetation map. In this way we can determine in which pixels the presence of forest is protecting
the soil from erosion (and its degree of importance) and in which pixels there is a greater risk of
erosion because the vegetation cover is not as stabilising as the forest
38
Module 3. Databases and modelling
Secondly, we consider the final delivery of mobilised sediment to the river network. For this,
NETMAP aggregates the delivered GEP index value of all pixels without forest protection and
draining to the same river reach. This value is transferred to each river reach of the river
network. In parallel, we also transfer to each river reach the percentage of forest contained in
the riparian area. Thus, while the value of the delivered GEP index indicates the probability of
potential sediment delivery to the river in reaches without riparian forest, in reaches with
riparian forest we obtain a quantification of the amount of sediment that the forest is potentially
filtering and preventing it from reaching the river.
39
Module 3. Databases and modelling
Finally, we are going to comment some of the main results of erosion modelling in three study
catchments in northern Spain.
As we can see in the figure, most of the sediment entering the river network comes from the
headwaters, especially the most deforested ones, where the protective and stabilising function
of the forest on the slopes has been lost. Our methodology allows us not only to identify which
of these slopes are most vulnerable to erosion, but also which drain into river reaches without
riparian forest that have lost their sediment filtering capacity and would therefore be a higher
priority for restoration. On the other hand, it makes it possible to identify where the forest
ecosystems that are contributing most to reducing erosion at source and in the delivery of
sediment to the river network are located. All of this is spatially explicit at the level of detail of
our vegetation mapping and digital elevation models (typically between twenty and thirty
metres resolution).
40
Module 3. Databases and modelling
41
Module 3. Databases and modelling
M3T6.
Water temperature models
Teacher: Laura Concostrina Zubiri
Introduction
In this lesson, I will describe the methods and tools to develop water temperature models
Traditionally, the water temperature has been predicted using surrogates such as altitude,
latitude, catchment area and air temperature. However, temperature surrogates may not
accurately represent the thermal environments experienced by the biota. In other cases, the
data they provide are too coarse to calculate local scale processes, such as stream metabolism.
Moreover, we lack temperature data for most streams and stream temperature data in the
absence of human activity is particularly scarce. Therefore, modelling is the most suitable way
to estimate it. There are several approaches to modelling water temperature although most
models can be classified into one of three groups: (i) regression models; (ii) stochastic models;
and (iii) deterministic models.
Regression models
Regression models consist of simple linear regression, multiple regression or logistic regression.
Simple linear regression models predict water temperature using only air temperature as the
input parameter and are quite effective at the weekly and monthly time scales. Multiple
regression models include explanatory variables other than the air temperature, such as river
discharge and time lag data. Logistic regression models are used on the basis that the air/water
temperature relationships are not necessarily linear. This may be due to influences by
42
Module 3. Databases and modelling
groundwater at low air temperatures and to evaporative cooling at high air temperatures,
among others. However, these models may not perform well on a daily scale.
Stochastic and deterministic models are most often used when the water temperature is
modelled at a daily scale. Stochastic models are simpler because they require only air
temperature as the input parameter, whereas deterministic models use all relevant
meteorological data to calculate energy components. In stochastic models, temperatures are
predicted at specific sites only (0D), while deterministic models can be carried out at different
spatial scales (1D, 2D). Therefore, deterministic models can quantify the different heat fluxes
acting on the river environment and at the same time consider the impact of different scenarios.
For example, they can predict water temperature in the absence of streamside vegetation or
under modified river discharge.
43
Module 3. Databases and modelling
The method and tools for developing the models should be selected according to the question
to be answered
In general, selecting a particular water temperature model depends on the modelling objective
as well as the data requirements. Depending on the spatial and temporal scale at which the
water temperature has to be modelled, the desired number of parameters and complexity to be
considered, and data availability, we should choose between regression, stochastic and
deterministic models. After choosing the most appropriate model, water and air temperature,
together with other potential additional data need to be collected. Water and air temperature
data can be obtained from stations or data loggers. Ideally, time series at the hourly or daily
scale and longer than one year are preferred to calibrate the model.
44
Module 3. Databases and modelling
In the ALICE project, we developed a linear model based on equilibrium principles to predict
water temperatures at a daily scale using average catchment temperatures. The model was
calibrated using air and water temperature data from hydrological stations. The model was able
to predict river temperatures with an average error in the order of +-0.5 ºC. The influence of hill-
side forests on water temperature was previously evaluated and then considered in the model
by modifying air temperature data in forest areas with forest using linear regression. The water
temperature data was crucial to calculate accurate river metabolism rates across fluvial
networks.
45
Module 3. Databases and modelling
Water temperature models are essential to better understanding and managing river
ecosystems.
Temperature is a critical water quality parameter because it regulates the maximum dissolved
oxygen concentration of the water, and influences the rate of chemical and biological reactions.
Therefore, water temperature determines the spatial distribution of stream biota and controls
biological processes. Water temperature models allow us to predict natural stream temperature
or expected changes in stream temperature under different environmental scenarios. In the
context of climate, land use and land cover change, this information is essential to develop
sound river management and restoration strategies.
46
Module 3. Databases and modelling
Image credits
47
Module 3. Databases and modelling
M3T7.
Water quality models
Teacher: Tamara Rodríguez Castillo
Introduction
Water quality modelling refers to the application of techniques and methods to assess and
predict the water quality in different water bodies, such as rivers or lakes. These models help
understand the physical, chemical, and biological processes that affect water quality and predict
how they may change in response to different conditions or human activities.
- First, with physical models, which use physical and mathematical principles to simulate
hydrodynamic and substance transport processes in water.
- Second, with chemical models that focus on the chemical processes that occur in water,
such as equilibrium reactions or degradation of contaminants.
- Third, with biological models that are used to understand and predict the biological
processes, such as algae growth or nutrient cycles.
- And, finally, with integrated models that combine elements of physical, chemical, and
biological models to provide a more comprehensive view of water quality.
1. The first step is to define the objectives of the study and to determine the specific water
quality parameters to model.
48
Module 3. Databases and modelling
2. Then, the required data is collected from field campaigns, research studies, or public
databases.
3. Thirdly, it’s necessary to identify the key factors (both natural and anthropogenic) that
affect water quality in your ecosystem.
4. Afterward, you have to select an appropriate modelling technique based on the
objectives and available data. In the field of water quality, both process-based models
and statistical models are commonly used to understand and predict water quality
parameters.
5. The fifth step is to calibrate and validate the selected model. We calibrate adjusting the
model parameters to match the observed data with the modelled data, and validate by
comparing the model predictions with independent data sets not used in the calibration
process.
6. And finally, the last step is to run the calibrated model to simulate different water quality
conditions and to analyse the results to understand the spatial and temporal variations,
identify hotspots of poor water quality, assess the impacts of different scenarios, and
evaluate the effectiveness of potential management strategies.
Modelling techniques
Well, as mentioned above, the most common modelling techniques in water quality are process-
based models and statistical models.
1. Process-Based models
Definition
Process-based models simulate the physical, chemical, and biological processes that
occur in water systems. They are based on mathematical equations and principles that
represent the underlying mechanisms governing water quality. So, they require detailed
knowledge of the system's characteristics.
49
Module 3. Databases and modelling
However, these models (1) typically require extensive data, which can be challenging to
collect and validate; (2) can be complex and computationally demanding; and (3) there
can be uncertainties associated with model structure, parameters, and input data.
Example (Delft3D)
50
Module 3. Databases and modelling
2. Statistical models
Definition
On the other hand, statistical models, also known as data-driven models, are developed
based on statistical relationships between water quality parameters and relevant
variables. They rely on observed data to establish correlations and make predictions.
Statistical models (1) are often simpler and easier to implement than process-based
models, (2) can capture patterns and relationships in the data without considering
underlying processes and (3) can be more flexible in handling missing data.
However, (1) it may not provide insights into the underlying processes driving water
quality changes, (2) it may not perform well when extrapolating beyond the range of
observed data, and (3) the selection of relevant variables can be subjective and may
overlook important factors affecting water quality.
51
Module 3. Databases and modelling
One of the most widely used statistical modelling techniques today is machine learning,
a branch of artificial intelligence that develops algorithms capable of learning patterns
and relationships directly from observed data, without incorporating the underlying
processes.
There are a multitude of machine learning techniques, such as Random Forest. Random
Forest is an ensemble learning algorithm used for classification and regression tasks. As
we can see in the slide, it creates multiple decision trees, each trained on random
subsets of data and features. When making predictions, it combines the outputs of all
the trees through voting (for classification) or averaging (for regression), resulting in a
more accurate and robust model that avoids overfitting and performs well on large
datasets.
As an example of the Random Forest application, in the slide I show the results of the
study conducted by Álvarez-Cabria et al. (2016). They modelled the spatial and seasonal
variability of three key water quality variables (water temperature and concentration of
nitrates and phosphates) for entire river networks in a large area in northern Spain. The
results provide a large-scale continuous picture of water quality, which could help
identify the main sources of change in water quality and assist in the prioritization of
river reaches for restoration projects.
52
Module 3. Databases and modelling
Image credits
4. Random Forest example: Álvarez-Cabria, M., Barquín, J., Peñas, F.J., 2016. Modelling the
spatial and seasonal variability of water quality for entire river networks: relationships
with natural and anthropogenic factors. Sci. Total Environ. 545–546, pp. 152-162.
[Link]
53