Reinforcement Learning in Process Control
Reinforcement Learning in Process Control
Review
Where Reinforcement Learning Meets Process Control: Review
and Guidelines
Ruan de Rezende Faria 1, * , Bruno Didier Olivier Capron 1 , Argimiro Resende Secchi 2
and Maurício B. de Souza, Jr. 1,2
1 Escola de Química, EPQB, Universidade Federal do Rio de Janeiro, Rio de Janeiro 21941-909, Brazil
2 Programa de Engenharia Química, PEQ/COPPE, Universidade Federal do Rio de Janeiro,
Rio de Janeiro 21941-972, Brazil
* Correspondence: rrfaria@[Link]
Abstract: This paper presents a literature review of reinforcement learning (RL) and its applications
to process control and optimization. These applications were evaluated from a new perspective
on simulation-based offline training and process demonstrations, policy deployment with transfer
learning (TL) and the challenges of integrating it by proposing a feasible approach to online process
control. The study elucidates how learning from demonstrations can be accomplished through
imitation learning (IL) and reinforcement learning, and presents a hyperparameter-optimization
framework to obtain a feasible algorithm and deep neural network (DNN). The study details a batch
process control experiment using the deep-deterministic-policy-gradient (DDPG) algorithm modified
with adversarial imitation learning.
Keywords: Markov decision process; imitation learning; transfer learning; process optimization
2. Reinforcement Learning
Four distinct but complementary phases mark the historical evolution of the RL
methodology; in chronological order, they are [1,33]:
1. Definition of the term from animal psychology;
2. Analysis for optimal control theory and machine learning;
3. Evolution of training procedures and pattern recognition;
4. Development of DNN, powerful hardware, data availability and more stable algorithms.
References that are important for the evolution of RL theory or have influenced the
current literature are listed in Table 1. Difficulties in defining the learning elements, as
well as creating and implementing algorithms, were the main challenges during Phase 2.
Howard [34] was the first to propose a viable RL algorithm based on dynamic programming
(DP). In Phase 3, there was a return of interest in the field of AI in general [1], where devel-
opments for RL remain current, specifically with Q-learning and REINFORCE algorithms
and stochastic RL theory. Phase 4 is when RL algorithms incorporated DNN, in addition
to methodologies to store information in memory (i.e., buffer replay), in order to stabilize
training and improve convergence.
Processes 2022, 10, 2311 3 of 31
2.1. Basics of RL
ML technologies are divided into three core classes: supervised learning, unsupervised
learning and reinforcement learning [25,43]. In supervised learning, knowledge about the
case study results from labeling its input and output features. An example of this is
classification and regression tasks where only cause (X) and effect (Y) are known; their
relationship is approximated (Y = f ( X )) using some data-driven technique (e.g., neural
networks) that has adequate generalization and accuracy when applied to the case study.
On the other hand, unsupervised learning obtains information about the case study only
through knowledge of the cause (X), and it is not necessary to include knowledge in the
form of a “teacher” to label its effect (Y). A typical application is clustering [44].
RL differs significantly from those technologies discussed above. According to Sutton
and Barto [1], this is because RL does not require external interference in the form of
a “teacher” to explore its environment, and is also not limited to learning how input
data is distributed (p( X )). Instead, it learns the best way to keep up with a given task
through repeated interactions with its environment. This makes the task complex, as more
elements have to be studied to characterize the problem. However, it can make machine
learning automated.
Defining an RL problem depends on understanding its essential elements: agent, envi-
ronment and reward. A widely studied example is the bandit problem. This is a problem
in which a fixed limited set of resources must be allocated between competing choices
(i.e., actions (at )) in a way that maximizes their expected gain. In other words, the agent’s
objective is to maximize the rewards sum along the trajectory h = ( a0 , a1 , · · · , a T ) [45].
Although the definition of the RL problem seems to be simple, selecting which actions
are the best without any a priori knowledge about the environment is a challenging task.
Sutton and Barto [1] detailed that it is necessary to estimate the value of the actions (At ) so
that it is possible to predict the most valuable actions in the long term. However, through
repeated executions of the bandit problem (episode), just selecting the k actions with the
highest value (exploitation) will likely lead to a sub-optimal expected sum of rewards
(Equation (1)), since RL does not explore enough combinations of actions throughout
the episodes, and neither approaches the expectation of Equation (1). The exploration–
exploitation trade-off dilemma is a classic problem in the RL literature. An alternative
would be to select some actions randomly at a specific rate, thus preventing the selection of
greedy actions and furthering exploration of the environment.
q∗ ( a) = E[ R(h)| At = at ] (1)
Processes 2022, 10, 2311 4 of 31
r t = r ( s t , a t , s t +1 ) ∈ R t (2)
Agent
For the case of continuous variables, the stochastic elements of this process include the
initial condition of the state s, in the form of Equation (3). In this equation, the probability
density function returns probability values for all possible states belonging to the set S.
p(s) ≥ 0, ∀s ∈ St
Z
(3)
p(s)ds = 1
s ∈ St
The probability transition from state s to state s0 when action a is taken (state transition)
defines the conditional probability density function p(s0 |s, a) (Equation (4)).
The decision made by the agent is determined by a policy (π), which aims to map
states to actions. In other words, it is a rule for deciding what to do given the current state
of the environment, and is robust enough to specify what to do in any situation and for the
entire state space. When using a deterministic policy, the action to be taken in each state is
unique, as shown in Equation (5).
π ( s ) ∈ A t , ∀ s ∈ St (5)
On the other hand, when the action to be taken in each state is stochastic, such a policy
is a function of the conditional probability density of taking action a in-state s, as shown in
Equation (6) [1,33,40].
π ( a|s) ≥ 0, ∀s ∈ St , ∀ a ∈ At
Z
(6)
π ( a|s)da = 1, ∀s ∈ St
a∈ At
Processes 2022, 10, 2311 5 of 31
T
R(h) = ∑ γ t −1 r ( s t , a t , s t +1 ) (7)
t =1
Initial state
s1 ∼ p ( s )
Agent Transition
at ∼ π ( at |st ) s t +1 ∼ p ( s t +1 | s t , a t )
t = t + 1
If t = T
Trajectory
h = [ s 1 , a 1 , · · · , s T , a T , s T +1 ]
In this equation, E pπ (h) denotes the expectation about the trajectory h extracted from
pπ (h), and pπ (h) denotes the probability density of observing the trajectory h under policy
π (Equation (9)).
T
p π ( h ) = p ( s 1 ) ∏ p ( s t +1 | s t , a t ) π ( a t | s t ) (9)
t =1
The procedure for computing the optimal policy π ∗ is not an obvious task due to the
problem of sequentially defining actions that can result in delayed rewards. Because of
this difficulty in determining which actions are right or wrong, it can be challenging to
decide which changes have to be made to improve the suboptimal policy. Therefore, it is
necessary to have efficient ways of discovering changes to the employed policy so that it is
improved [40,42,46].
2.2.3. Algorithms
Approaches employing value function are a traditional way to learn optimal policy.
The overall objective of value functions is to approximate the return value for all possible
trajectories (h) to improve the employed policy (π). When only the state value is taken for
the computation of the value function, it results in Equation (10), while if the action value
is also included, Equation (11) is obtained.
π ∗ ( a|s) = δ( a − aπ (s))
aπ (s) = argmax E pπ (h) [r (s, a, s0 ) + γV π (s0 )] (13)
a∈ At
The other option is to use the value function of the state–action pair, such that the
optimal policy depends on Equations (14) and (15), respectively.
T
Qπ (s, a, θ ) = ∑ φi (s, a)θi = φT (s, a)θ (16)
t =1
Processes 2022, 10, 2311 7 of 31
The problem now is to estimate θ to minimize Equation (17), using Monte Carlo (MC)
or temporal difference estimates for Qπ (s, a) [31].
θ ∗ = argmax J (θ ) (19)
θ
Z
J (θ ) = E pπ (h|θ ) [ R(h)] = p(h|θ ) R(h)dh (20)
T
p ( h | θ ) = p ( s 1 ) ∏ p ( s t +1 | s t , a t ) π ( a t | s t , θ ) (21)
t =1
indicated by the limit on the number of updates of all states and actions that can be
efficiently performed individually. On the other hand, with parametric approximators,
the update of the value function for each state–action pair individually also influences the
estimate of the value function for the other state–action pairs, thus giving it a capacity
for generalization that makes the learning process more powerful, but potentially more
challenging to manage and understand [1,25,46].
Within this perspective, in DNNs, updating the value function of each state–action
pair is conducted following Equations (22) and (23). The TD technique is used for the value
function evaluation (i.e., TD (0)) and defines the loss function (Lt ), with backpropagation
and stochastic descending-gradient algorithms to update the neural-network parameters
(θ). As a result, Sutton and Barto [1] did not expect to find a value function with zero error
for all states. But an approximation that balances the errors in different states.
2
Lt = [rt + γ max a0 Q(st+1 , a0 , θt )] − Q(st , at , θt ) (22)
δQ(st , at , θt )
θt+1 = θt + α [rt + γ max a0 Q(st+1 , a0 , θt )] − Q(st , at , θt ) (23)
δθt
Figure 3 presents a simplified scheme describing the procedure for training the critic
network (i.e., until each episode of size T ends). It is an iterative procedure in which the
sampled initial state is restarted after the completion of each episode of size T, depending
on how the environment was constructed and the rewards were formulated (e.g., setting
a threshold a priori for accumulated rewards, or declaring state transitions considered
∧
impossible or unfeasible) [25]. In addition, in the terminal state, the target value y t = rt .
Agent Transition
Initial state
Softmax
-greedy
until
Figure 3. Simplified diagram describing the training procedure of a critic (Q-learning) parameterized
by a DNN.
In its original form, the procedure described above may result in inadequate learning
∧
because the target y t is never exactly approximated in the forward step. It can result
∧
in a critic with divergent policy for overestimating y t , leaving some amount of residual
∧
TD-error (δt = y t − yt ). In Equation (24), the dispersion of the critic estimates depends on
its complexity, variance of future rewards and the TD-error, accentuated by the value of
γ ∼ 1 [48–50].
The alternatives proposed in the literature to alleviate the mentioned problems are:
Processes 2022, 10, 2311 9 of 31
• Controlling the exploratory component of the critic model used by the agent (e.g.,
softmax or e-greedy [46])) will contribute to adequately exploring enough transitions of
state, avoiding obtaining sub-optimal policies;
• Using experience replay to reduce the effect of temporal correlations between transi-
tions uniformly sampled at random from the buffer, which allows estimating θ with
important dynamic information [10];
∧
• Updating from the target value y t with delayed (or filtered) copies of the original
DNN (i.e., θt+1 = κθt + (1 − κ )θt+1 ) [48].
θ t +1 = θ t + α ∇ θ J ( θ t ) (25)
Z
∇θ J (θ ) = ∇θ p(h|θ ) R(h)dh
Z
= p(h|θ )∇θ log p(h|θ ) R(h)dh (26)
Z T
= p(h|θ ) ∑ ∇θ log π ( at |st , θ ) R(h)dh
t =1
T
∇θ J ( θ ) = E pπ (h|θ ) ∑ ∇θ log π (at |st , θ ) R(h) (27)
t =1
N T
1
∇θ J (θ ) =
N ∑ ∑ ∇θ log π (at,n |st,n , θ ) R(hn ) (28)
n =1 t =1
The actor categorizes actions instead of their value π ( at |st , θ ), and there is no way to
compute ∇ J directly using gradient-based optimization techniques. Because of this, the
Monte Carlo simulation approximates the value of Equation (28) (i.e., random sampling
N episodes of size T). As a result, the actor is updated by Equation (25), where R(hn ) is
directly proportional to ∇θ J, and log π ( at,n |st,n , θ ) is inversely proportional to the policy
adopted (i.e., it follows from the identity ∇ ln x = ∇xx ), as actions sampled with frequency
are chosen even if they do not produce the highest expected return [1,13,25,33,54].
Processes 2022, 10, 2311 10 of 31
For example, training with the REINFORCE [42] algorithm requires the execution
of a very large number of cycles, as represented in the diagram shown in Figure 4
(i.e., when dim( T ) ∼ inf), which is computationally unfeasible for some online appli-
cations [1,25]. Therefore, it uses a baseline to reduce variance, similar to the temporal-
difference-learning methodology. For example, Williams [42] used a constant to represent
it (e.g., mean reward). However, it is more appropriate to employ a state-specific option,
with the value taken immediately before the actor samples the action.
Actor
Initial state Transition
until
Monte Carlo
Step backward
Equation (28)
Equation (25)
Policy Value
This paradigm for DNN is shown in Figure 6, with the actor selecting the action at .
After this step, the critic weighs the descending gradient in the algorithm for updating
θ a (t + 1) (e.g., Q-AC, TD-AC, A-AC). In the backward step, Equation (30) characterizes
the updating of the actor’s parameters. Additionally, the critic updates the values of
such actions (Equation (31)). At the end, both parameterized models must achieve good
generalization so that the actor does not becomes stuck around sub-optimal solutions and
the critic minimizes the residual TD-error, as seen in the form of Equation (24) [1,13,25].
Actor
Initial state Transition
until
Step Forward
Step backward
Q-AC, TD-AC
Equations (30) and (31)
Figure 6. Simplified diagram describing the training procedure of a critic and actor parameterized by
a DNN.
∇ J (µθ ) = E pπ (h|µθ ) ∇θ µθ (s) ∇ a Qµ (s, a)| a=µθ (s) ; (32)
• Deep deterministic policy gradient (DDPG) (i.e., actor–critic and off-policy) [58]. This
is an updated version of the DPG algorithm regarding the use of DNN, replay buffer,
target networks and batch normalization, in addition to the possibility of handling the
exploration problem independent of the learning algorithm used;
• Proximal policy optimization (PPO) [59]. Contrary to the algorithms above (i.e.,
off-policy), PPO is an algorithm that learns while interacting with the environment
over different episodes (i.e., on-policy). Methodologically, this property comes from
another similar algorithm considered more complex (trust region policy optimiza-
tion (TRPO)), addressing the Kullback–Leibler (KL) divergence effect and surrogate
objective functions;
• Soft actor–critic (SAC) (i.e., actor–critic and off-policy) [56]. This algorithm is com-
posed of an actor and a critic, and includes a smooth value function, which is responsi-
ble for stabilizing the training of the actor and the critic. In addition, it also has similar
properties to the DDPG algorithm; however, it adds an entropy value to compose
the buffer.
rollout data
buffer (D)
Simulation
update
rollout data
buffer (D)
Simulation
buffer (E)
data-driven
update
learn
training phase
data-driven
Deployment
Figure 8. Diagram describing the modules for offline agent training (Modules (1) e (2)) and process-
line deployment (Module (3)).
online MDP by transferring information from the offline MDP and then starting to learn
from a condition that is at least sub-optimal.
applied in the offline training of the RL agent. This allows obtaining an RL agent with
information about the actual process, including process constraints and control objectives.
In general, all control algorithms are parameterized by DNNs for the actor and critic,
as they have already proven to be efficient in handling large and complex data [10]. Further-
more, only Ramanathan et al. [67] and Hwangbo and Sin [68] have employed value-based
methods, while the other authors used algorithms derived from actor–critic methods.
The explanation for this is the benefit of integrating an actor to decide the control actions
weighted by a critic, accelerating the learning process while reducing the variance of actions
selected by the actor resulting from TD methods.
Processes 2022, 10, 2311 16 of 31
Except for Shah and Gopal [28] and Kim et al. [30], other authors cited employed
off-policy learning, with the index (1) referring to training carried out in a simulated envi-
ronment. The practicality of the offline training step depends on a simulated environment
that is trustworthy, allowing the testing of various process conditions. Furthermore, TD
methods are preferable to MC methods precisely because they combine MC and DP ideas
to obtain process estimates and can be improved with eligibility traces (λ), which is a way
of weighting between TD(0) “targets” and Monte-Carlo “returns”.
Another important consideration is the agents used by Dogru et al. [20], Chen et al. [69]
and Oh et al. [70], who developed control structures with agents learning asynchronously.
The A2C and A3C algorithms are variations of the actor–critic algorithm with agents
learning asynchronously, with two or three agents in parallel [71]. In multi-agent DDPG
(i.e., MADDPG with two agents), the training is decentralized regarding the control actions
taken by each actor, while it is centralized by only one critic to evaluate the actions taken by
each actor. Thus, such approaches allow agents with competitive and cooperative control
objectives to improve offline learning and facilitate deployment in the entire process, as
seen in Chen et al. [69].
Powell et al. [18] innovated by proposing the first algorithm for RTO. They employed
an agent with a deep actor–critic algorithm and replaced the standard descending gra-
dient optimization algorithm with the particle swarm method, justifying it as a global
optimization method.
The most employed TL methodology for the RL-based control of batch and contin-
uous processes is the policy-transfer technique, which is similar to the approach used in
Peng et al. [8] (e.g., see [16–18,80]). It depends on a reliable process approximation for
extensive offline learning before application to the real process.
For example, Petsagkourakis et al. [14] implemented this TL methodology to develop
an RL-based controller for a biochemical batch process. First, they obtained the optimal
policy for the simulated environment. Then, they applied it online by allowing the output
layer of the DNN to learn from consecutive batches while freezing the remaining layers.
Mowbray et al. [27] developed a TL framework employing an inverse RL technique, which
allowed them to analyze historical process data for synchronous identification of a reward
function and then obtain the control policy in a step before its application in the online
process. According to the work’s objective, this was proven necessary, as the proposed RL
agent started the learning process following the policy already obtained in the previous step.
The other forms of TL in RL (e.g., RS, LD, IM) have still not been significantly explored
because they combine all forms of learning and obtaining a stable policy therefore becomes
challenging. An example of these difficulties is the inclusion of the process constraints.
However, this is the way forward in chemical processes in the future, as it will result in a
more efficient and faster control agent to adapt to changes in the process dynamics [12,32].
4.1. Overview
First, implementing a new methodology for optimal process control depends on
designing an MDP based on the process dynamics. Through the sections already described
in this review, the available knowledge about the process dynamics will define the stages of
offline training and TL. In other words, the objective is to extract useful information from
already established policies in the process and simulation data.
Processes 2022, 10, 2311 18 of 31
The focus has been on batch and continuous processes, in which the inclusion of the
transfer-learning process is complex. At this point, state-of-the-art development is still
being consolidated, as the inclusion of process constraints remains to be addressed. Some
exceptions are the works of Pan et al. [26] and Mowbray et al. [27,64].
As long as the process states can be observed, the configuration of the offline training
step is directly related to the training of the learning agent itself, namely:
1. The choice of the algorithm;
2. The exploration–exploitation trade-off dilemma;
3. Hyperparameter optimization.
The choice of algorithm depends on the complexity of the process. There is a preference
for off-policy algorithms for batch and continuous processes, as they allow the storage of
information from other policies. However, the exploration–exploitation trade-off dilemma
analysis indicates the most appropriate policy to approximate the MDP. Being stochastic,
it deals better with the uncertainties inherent in the process. Finally, hyperparameter
optimization is essential to obtain a feasible algorithm and DNN.
The transferring of knowledge obtained in the previous phases is also necessary to
implement the agent for optimal process control. In Section 3.4, common approaches to
robot control were detailed. However, for chemical processes, a conservative approach must
consider process constraints. Thus, learning from demonstrations, transferring actor and
critic policies, adapting rewards (i.e., reward shaping) and observing the correspondence
between offline and online environments are essential to fulfilling this purpose.
Lastly, process-line maintenance is another critical issue to consider. At this point, the
control dynamics and complexity of the process indicate how often measurements are avail-
able. This allows establishing when to retrain the obtained policy. Generally, the preference
is to directly apply an alternative in the online environment without performing a backup.
However, the current state of the art is still embryonic. For example, Wang and Ye [83]
proposed consciousness-driven reinforcement learning, which showed superior results to
standard algorithms (i.e., DDPG, PPO) for simplified examples only (i.e., OpenAIGym).
Hence, the current state of the art in chemical processes is exploring RL techniques that
include information from demonstration to discover unknown a priori information and
limit the search space of the chosen algorithm.
rollout data
training phase
policy reuse
reward shaping
Deployment
Figure 9. Diagram containing the modules for offline agent training (Modules (1) and (2)) and
process-line deployment (Module (3)) adapted for chemical processes.
a t = µ ( a t , θ a ) + ηt (34)
1
∇ J (µθ ) =
K ∑ ∇θ µ(st , µ(st+1 , θ a0 ) ∇ a Q(st , at , θ a )| a=µ(st ,θa ) (38)
1
Ld =
2K ∑ yd log(xd ) + (1 − yd ) log(1 − xd ) (39)
Generator
rollout data
buffer (D)
Simulation Discriminator
update
D E
Expert
buffer (E)
data-driven
learn
Offline learning
Figure 10. Diagram describing the modules for offline training of a generator and expert with a
discriminator.
Domain limit
(hyperparameters) Optimization scenario
• The algorithm is specific to MDP where the sampled state–action pairs are continuous;
• The simplicity of the algorithm, which makes writing the source code easier, the
proposed updates to the algorithm (e.g., prioritized buffer replay, inverting gradient)
and distributed optimization;
• There are applications for optimal process control (e.g., for DDPG, see Ma et al. [17]
and Spielberg et al. [80]);
• The algorithm combines RL and adversarial imitation learning to learn from
demonstrations.
ds1 a s
= −( a1 + 0.5a21 )s1 + 0.5 2 2 (40)
dt s1 + s2
ds2
= a1 s1 − 0.7a2 s1 (41)
dt
ds1
= −( a1 + 0.5a21 )s1 + a2 (42)
dt
ds2
= a1 s1 − a2 s1 (43)
dt
The optimization objective is to maximize E[s2 ( T )] given the limits of the control
actions (i.e., a1 ∈ (0, 5) and a2 ∈ (0, 5)) for batches of ten time intervals, with the reward for
each time interval taken from the simulation and discriminated with the expert reward, as
shown in Equation (44). In this equation, a standard DDPG is derived from β = 1 and an
indirect-imitation-learning approach from β = 0. In addition, the initial condition is fixed
at the s1 = 1 and s2 = 0 (s0 = (1, 0)).
Figure 12. Optimization scheme with Optuna to validate the online control experiment with domain
limit according to Table 4.
Table 4 summarizes the essential hyperparameters for optimization with the TPE
algorithm and the search space for the domain containing the hyperparameters. β is the
main hyperparameter to validate the proposed algorithm against the standard DDPG
algorithm [58] and sub-optimal demonstrations from the actual process (E), which were
obtained with REINFORCE algorithm [14]. In addition, the hyperparameters N, D, the
activation function and the optimizer (i.e., Adam [101]) do not influence the optimization of
E[s2 ( T )]. Then, a limit value considered sufficient for the control experiment was defined.
When running the online control experiment according to the algorithmic implemen-
tation details described above, Figure 13 shows that the TPE algorithm took 18 trials to
obtain policies with an objective function value comparable to that obtained by
Petsagkourakis et al. [15] (i.e., 0.583) against 0.575 e 0.55 with the DDPG algorithm and
Processes 2022, 10, 2311 25 of 31
MPC, respectively. The TPE algorithm obtained some erroneous policies until trial 33,
where the best value was obtained (i.e., 0.595). At the end, the TPE algorithm showed dis-
persion of the objective function values precisely because it was not able to find a solution
that exceeded the best value, as detailed in Bergstra et al. [95].
Figure 13. Optimization history plot for modified DDPG considering 100 trials with the
TPE algorithm.
Figure 14 shows the hyperparameters that had more influence on the objective func-
tion. The highlight here is the hyperparameter β, which weighs the importance of the
discriminator for the objective value; according to [79] it is crucial to properly know the
dynamics of the environment to specify this hyperparameter correctly. This can be seen
in Figure 15, where values of β important to the objective function are dispersed around
its optimal value (i.e., β = 0.8), indicating that including demonstrations in the process of
offline learning results in appropriate online control policies.
Figure 14. Hyperparameter importance plot for modified DDPG considering 100 trials with the
TPE algorithm.
Processes 2022, 10, 2311 26 of 31
Figure 15. Hyperparameter importance plot considering 100 trials of the TPE algorithm for data
distribution and sampling.
Regarding the structures of the actor, critic and discriminator networks regarding
the number of layers and neurons and the learning rate, the critic network has the most
significant contribution to the objective function (e.g., see αc and Nc0 ). More details can be
found in Figure 16, where a shallow architecture (i.e., Lac = 2) is sufficient to approximate
the value function, which depends on a significant number of neurons only in the first
layer for local approximation of the input features (i.e., state–action pair), as explained in
Das et al. [102]. Furthermore, training with the gradient-descent algorithm is less complex
compared to the actor and discriminator. The former depends on a deeper architecture in
general (i.e., L a = 4) and training with a conservative GD optimization algorithm due to its
greater complexity and nonlinearity to approximate the deterministic gradient (Figure 17).
The latter also depends on a deep architecture and a conservative GD optimization
(Figure 18), which is due to the complexity of the objective function resulting from genera-
tive and discriminative networks (i.e., by employing binary cross-entropy loss) [87].
Figure 16. Hyperparameter importance slice plot for modified DDPG considering 100 trials of the
TPE algorithm for the critic network.
Processes 2022, 10, 2311 27 of 31
Figure 17. Hyperparameter importance slice plot for modified DDPG considering 100 trials of the
TPE algorithm for the actor network.
Figure 18. Hyperparameter importance slice plot for modified DDPG considering 100 trials of the
TPE algorithm for the discriminative network.
Processes 2022, 10, 2311 28 of 31
6. Conclusions
These final considerations result from an in-depth state-of-the-art review of artificial
intelligence and optimal process control. The objective was to develop a complete guide for
hyperparameter optimization, imitation learning and transfer learning, since there are no
articles covering such subjects together. Thus, the conclusions made answer several claims
about the integration of reinforcement learning and process control:
• State-of-the-art technologies are still embryonic;
• Batch and continuous processes require different learning structures;
• Developing state-of-the-art offline training technologies is essential;
• Transfer learning has a broad meaning in RL, since it can encompass learning from
demonstration, reward shaping, policy transfer and inter-task mapping;
• The proposed modified DDPG algorithm with an off-policy discriminator confirmed
the hypothesis that information from process demonstrations improves the perfor-
mance of the standard DDPG algorithm, as detailed in Section 5.
Finally, the control experiment carried out in Section 5 provides some guidelines for
extending this approach to more complex systems (batch and continuous). These results
can support the study of hyperparameter optimization for other case studies. Furthermore,
the main challenge is developing new control structures appropriate to process constraints.
To achieve this, it remains necessary to improve Modules 1, 2 and 3, shown in Figure 9.
Author Contributions: R.d.R.F. participated in all steps of the research method: conceptualization,
methodology, writing—original draft preparation. For review and editing, all authors participated.
Conceptualization and supervision, B.D.O.C., A.R.S. and M.B.d.S.J. All authors have read and agreed
to the published version of the manuscript.
Funding: This study was financed in part by the Coordenação de Aperfeiçoamento de Pessoal de
Nível Superior—Brasil (CAPES)—Finance Code 001. Maurício B. de Souza Jr. is grateful for financial
support from CNPq (Grant No. 311153/2021-6) and Fundação Carlos Chagas Filho de Amparo à
Pesquisa do Estado do Rio de Janeiro (FAPERJ) (Grant No. E-26/201.148/2022).
Data Availability Statement: Not applicable.
Conflicts of Interest: The authors declare no conflict of interest.
References
1. Sutton, R.S.; Barto, A.G. Reinforcement Learning: An Introduction; MIT Press: Cambridge, MA, USA, 2018.
2. Bellman, R. Dynamic Programming; Princeton University Press: Princeton, NJ, USA, 1957; Volume 95.
3. Bellman, R. A Markovian decision process. J. Math. Mech. 1957, 6, 679–684. [CrossRef]
4. Hoskins, J.; Himmelblau, D. Process control via artificial neural networks and reinforcement learning. Comput. Chem. Eng. 1992,
16, 241–251. [CrossRef]
5. Hinton, G.; Srivastava, N.; Swersky, K. Neural networks for machine learning lecture 6a overview of mini-batch gradient descent.
Cited On 2012, 14, 2.
6. Hinton, G.E.; Srivastava, N.; Krizhevsky, A.; Sutskever, I.; Salakhutdinov, R.R. Improving neural networks by preventing
co-adaptation of feature detectors. arXiv 2012, arXiv:1207.0580.
7. Wulfmeier, M.; Posner, I.; Abbeel, P. Mutual alignment transfer learning. In Proceedings of the Conference on Robot Learning
(PMLR), Mountain View, CA, USA, 13–15 November 2017; pp. 281–290.
8. Peng, X.B.; Andrychowicz, M.; Zaremba, W.; Abbeel, P. Sim-to-real transfer of robotic control with dynamics randomization. In
Proceedings of the 2018 IEEE International Conference on Robotics and Automation (ICRA), Brisbane, Australia, 21–25 May 2018;
pp. 3803–3810.
9. Silver, D.; Schrittwieser, J.; Simonyan, K.; Antonoglou, I.; Huang, A.; Guez, A.; Hubert, T.; Baker, L.; Lai, M.; Bolton, A.; et al.
Mastering the game of go without human knowledge. Nature 2017, 550, 354–359. [CrossRef]
10. Mnih, V.; Kavukcuoglu, K.; Silver, D.; Rusu, A.A.; Veness, J.; Bellemare, M.G.; 38 Graves, A.; Riedmiller, M.; Fidjeland, A.K.;
Ostrovski, G.; et al. Human-level control through deep reinforcement learning. Nature 2015, 518, 529–533. [CrossRef]
11. Vinyals, O.; Babuschkin, I.; Czarnecki, W.M.; Mathieu, M.; Dudzik, A.; Chung, J.; Choi, D.H.; Powell, R.; Ewalds, T.; Georgiev, P.;
et al. Grandmaster level in StarCraft II using multi-agent reinforcement learning. Nature 2019, 575, 350–354. [CrossRef]
12. Nian, R.; Liu, J.; Huang, B. A review on reinforcement learning: Introduction and applications in industrial process control.
Comput. Chem. Eng. 2020, 139, 106886. [CrossRef]
Processes 2022, 10, 2311 29 of 31
13. Buşoniu, L.; de Bruin, T.; Tolić, D.; Kober, J.; Palunko, I. Reinforcement learning for control: Performance, stability, and deep
approximators. Annu. Rev. Control 2018, 46, 8–28. [CrossRef]
14. Petsagkourakis, P.; Sandoval, I.O.; Bradford, E.; Zhang, D.; del Rio-Chanona, E.A. Reinforcement learning for batch bioprocess
optimization. Comput. Chem. Eng. 2020, 133, 106649. [CrossRef]
15. Petsagkourakis, P.; Sandoval, I.O.; Bradford, E.; Zhang, D.; del Rio-Chanona, E.A. Reinforcement learning for batch-to-batch
bioprocess optimisation. In Computer Aided Chemical Engineering; Elsevier: Amsterdam, The Netherlands, 2019; Volume 46,
pp. 919–924.
16. Yoo, H.; Kim, B.; Kim, J.W.; Lee, J.H. Reinforcement learning based optimal control of batch processes using Monte-Carlo deep
deterministic policy gradient with phase segmentation. Comput. Chem. Eng. 2021, 144, 107133. [CrossRef]
17. Ma, Y.; Zhu, W.; Benton, M.G.; Romagnoli, J. Continuous control of a polymerization system with deep reinforcement learning. J.
Process Control 2019, 75, 40–47. [CrossRef]
18. Powell, K.M.; Machalek, D.; Quah, T. Real-time optimization using reinforcement learning. Comput. Chem. Eng. 2020, 143, 107077.
[CrossRef]
19. Nikita, S.; Tiwari, A.; Sonawat, D.; Kodamana, H.; Rathore, A.S. Reinforcement learning based optimization of process
chromatography for continuous processing of biopharmaceuticals. Chem. Eng. Sci. 2021, 230, 116171. [CrossRef]
20. Dogru, O.; Wieczorek, N.; Velswamy, K.; Ibrahim, F.; Huang, B. Online reinforcement learning for a continuous space system
with experimental validation. J. Process Control 2021, 104, 86–100. [CrossRef]
21. Ławryńczuk, M.; Marusak, P.M.; Tatjewski, P. Cooperation of model predictive control with steady-state economic optimisation.
Control Cybern. 2008, 37, 133–158.
22. Skogestad, S. Control structure design for complete chemical plants. Comput. Chem. Eng. 2004, 28, 219–234. [CrossRef]
23. Backx, T.; Bosgra, O.; Marquardt, W. Integration of model predictive control and optimization of processes: Enabling technology
for market driven process operation. IFAC Proc. Vol. 2000, 33, 249–260. [CrossRef]
24. Adetola, V.; Guay, M. Integration of real-time optimization and model predictive control. J. Process Control 2010, 20, 125–133.
[CrossRef]
25. Aggarwal, C.C. Neural Networks and Deep Learning; Springer: Berlin/Heidelberg, Germany, 2018; Volume 10, pp. 978–983.
26. Pan, E.; Petsagkourakis, P.; Mowbray, M.; Zhang, D.; del Rio-Chanona, E.A. Constrained model-free reinforcement learning for
process optimization. Comput. Chem. Eng. 2021, 154, 107462. [CrossRef]
27. Mowbray, M.; Smith, R.; Del Rio-Chanona, E.A.; Zhang, D. Using process data to generate an optimal control policy via
apprenticeship and reinforcement learning. AIChE J. 2021, 67, e17306. [CrossRef]
28. Shah, H.; Gopal, M. Model-free predictive control of nonlinear processes based on reinforcement learning. IFAC-PapersOnLine
2016, 49, 89–94. [CrossRef]
29. Alhazmi, K.; Albalawi, F.; Sarathy, S.M. A reinforcement learning-based economic model predictive control framework for
autonomous operation of chemical reactors. Chem. Eng. J. 2022, 428, 130993. [CrossRef]
30. Kim, J.W.; Park, B.J.; Yoo, H.; Oh, T.H.; Lee, J.H.; Lee, J.M. A model-based deep reinforcement learning method applied to
finite-horizon optimal control of nonlinear control-affine system. J. Process Control 2020, 87, 166–178. [CrossRef]
31. Badgwell, T.A.; Lee, J.H.; Liu, K.H. Reinforcement learning–overview of recent progress and implications for process control. In
Computer Aided Chemical Engineering; Elsevier: Amsterdam, The Netherlands, 2018; Volume 44, pp. 71–85.
32. Görges, D. Relations between model predictive control and reinforcement learning. IFAC-PapersOnLine 2017, 50, 4920–4928.
[CrossRef]
33. Sugiyama, M. Statistical Reinforcement Learning: Modern Machine Learning Approaches; CRC Press: Boca Raton, FL, USA, 2015.
34. Howard, R.A. Dynamic Programming and Markov Processes; MITPL: Cambridge, MA, USA, 1960.
35. Thorndike, E.L. Animal intelligence: An experimental study of the associative processes in animals. Psychol. Rev. Monogr. Suppl.
1898, 2, 1.
36. Minsky, M. Neural Nets and the Brain-Model Problem. Doctoral Dissertation, Princeton University, Princeton, NJ, USA, 1954;
Unpublished.
37. Minsky, M. Steps toward artificial intelligence. Proc. IRE 1961, 49, 8–30. [CrossRef]
38. Barto, A.G.; Sutton, R.S.; Anderson, C.W. Neuronlike adaptive elements that can solve difficult learning control problems. IEEE
Trans. Syst. Man Cybern. 1983, SMC-13, 834–846. [CrossRef]
39. Sutton, R.S. Learning to predict by the methods of temporal differences. Mach. Learn. 1988, 3, 9–44. [CrossRef]
40. Watkins, C.J.C.H. Learning from Delayed Rewards; University of Cambridge: Cambridge, UK, 1989.
41. Gullapalli, V. A stochastic reinforcement learning algorithm for learning real-valued functions. Neural Netw. 1990, 3, 671–692.
[CrossRef]
42. Williams, R.J. Simple statistical gradient-following algorithms for connectionist reinforcement learning. Mach. Learn. 1992,
8, 229–256. [CrossRef]
43. Bishop, C.M. Pattern Recognition and Machine Learning; Springer: Berlin/Heidelberg, Germany, 2006.
44. LeCun, Y.; Bengio, Y.; Hinton, G. Deep learning. Nature 2015, 521, 436–444. [CrossRef] [PubMed]
45. Berry, D.A.; Fristedt, B. Bandit Problems: Sequential Allocation of Experiments (Monographs on Statistics and Applied Probability);
Chapman and Hall: London, UK, 1985; Volume 5, pp. 71–87.
46. Sutton, R.S.; Barto, A.G. Introduction to Reinforcement Learning; MIT Press Cambridge: Cambridge, MA, USA, 1998; Volume 135.
Processes 2022, 10, 2311 30 of 31
47. Shannon, C.E. A mathematical theory of communication. ACM SIGMOBILE Mob. Comput. Commun. Rev. 2001, 5, 3–55. [CrossRef]
48. Silver, D.; Lever, G.; Heess, N.; Degris, T.; Wierstra, D.; Riedmiller, M. Deterministic policy gradient algorithms. In Proceedings of
the International Conference on Machine Learning (PMLR), Bejing, China, 22–24 June 2014.
49. Thrun, S.; Schwartz, A. Issues in using function approximation for reinforcement learning. In Proceedings of the 1993 Connectionist
Models Summer School; Lawrence Erlbaum: Hillsdale, NJ, USA, 1993.
50. Fujimoto, S.; Van Hoof, H.; Meger, D. Addressing function approximation error in actor-critic methods. arXiv 2018,
arXiv:1802.09477.
51. Sutton, R.S.; McAllester, D.A.; Singh, S.P.; Mansour, Y. Policy gradient methods for reinforcement learning with function
approximation. In Proceedings of the Advances in Neural Information Processing Systems, Denver, CO, USA, 29 November–4
December 2000; pp. 1057–1063.
52. Gordon, G.J. Stable function approximation in dynamic programming. In Machine Learning Proceedings 1995; Elsevier: Amsterdam,
The Netherlands, 1995; pp. 261–268.
53. Tsitsiklis, J.N.; Van Roy, B. Feature-based methods for large scale dynamic programming. Mach. Learn. 1996, 22, 59–94. [CrossRef]
54. Grondman, I.; Busoniu, L.; Lopes, G.A.; Babuska, R. A survey of actor-critic reinforcement learning: Standard and natural policy
gradients. IEEE Trans. Syst. Man Cybern. Part C (Appl. Rev.) 2012, 42, 1291–1307. [CrossRef]
55. Ramicic, M.; Bonarini, A. Augmented Replay Memory in Reinforcement Learning With Continuous Control. arXiv 2019,
arXiv:1912.12719.
56. Haarnoja, T.; Zhou, A.; Abbeel, P.; Levine, S. Soft actor-critic: Off-policy maximum entropy deep reinforcement learning with a
stochastic actor. In Proceedings of the International Conference on Machine Learning (PMLR), Stockholm, Sweden, 10–15 July
2018; pp. 1861–1870.
57. Benhamou, E. Variance Reduction in Actor Critic Methods (ACM). arXiv 2019, arXiv:1907.09765.
58. Lillicrap, T.P.; Hunt, J.J.; Pritzel, A.; Heess, N.; Erez, T.; Tassa, Y.; Silver, D.; Wierstra, D. Continuous control with deep
reinforcement learning. arXiv 2015, arXiv:1509.02971.
59. Schulman, J.; Wolski, F.; Dhariwal, P.; Radford, A.; Klimov, O. Proximal policy optimization algorithms. arXiv 2017,
arXiv:1707.06347.
60. Kaelbling, L.P.; Littman, M.L.; Cassandra, A.R. Planning and acting in partially observable stochastic domains. Artif. Intell. 1998,
101, 99–134. [CrossRef]
61. Bonvin, D. Optimal operation of batch reactors—A personal view. J. Process Control 1998, 8, 355–368. [CrossRef]
62. Bonvin, D.; Srinivasan, B.; Ruppen, D. Dynamic Optimization in the Batch Chemical Industry; Technical Report; NTNU: Trondheim,
Norway, 2001.
63. Arpornwichanop, A.; Kittisupakorn, P.; Mujtaba, I. On-line dynamic optimization and control strategy for improving the
performance of batch reactors. Chem. Eng. Process. Process. Intensif. 2005, 44, 101–114. [CrossRef]
64. Mowbray, M.; Petsagkourakis, P.; Chanona, E.A.d.R.; Smith, R.; Zhang, D. Safe Chance Constrained Reinforcement Learning for
Batch Process Control. arXiv 2021, arXiv:2104.11706.
65. Oh, T.H.; Park, H.M.; Kim, J.W.; Lee, J.M. Integration of reinforcement learning and model predictive control to optimize
semi-batch bioreactor. AIChE J. 2022, 68, e17658. [CrossRef]
66. Ellis, M.; Durand, H.; Christofides, P.D. A tutorial review of economic model predictive control methods. J. Process Control 2014,
24, 1156–1178. [CrossRef]
67. Ramanathan, P.; Mangla, K.K.; Satpathy, S. Smart controller for conical tank system using reinforcement learning algorithm.
Measurement 2018, 116, 422–428. [CrossRef]
68. Hwangbo, S.; Sin, G. Design of control framework based on deep reinforcement learning and Monte-Carlo sampling in
downstream separation. Comput. Chem. Eng. 2020, 140, 106910. [CrossRef]
69. Chen, K.; Wang, H.; Valverde-Pérez, B.; Zhai, S.; Vezzaro, L.; Wang, A. Optimal control towards sustainable wastewater treatment
plants based on multi-agent reinforcement learning. Chemosphere 2021, 279, 130498. [CrossRef]
70. Oh, D.H.; Adams, D.; Vo, N.D.; Gbadago, D.Q.; Lee, C.H.; Oh, M. Actor-critic reinforcement learning to estimate the optimal
operating conditions of the hydrocracking process. Comput. Chem. Eng. 2021, 149, 107280. [CrossRef]
71. Mnih, V.; Badia, A.P.; Mirza, M.; Graves, A.; Lillicrap, T.; Harley, T.; Silver, D.; Kavukcuoglu, K. Asynchronous methods for deep
reinforcement learning. In Proceedings of the International Conference on Machine Learning, New York, NY, USA, 19–24 June
2016; pp. 1928–1937.
72. Tan, C.; Sun, F.; Kong, T.; Zhang, W.; Yang, C.; Liu, C. A survey on deep transfer learning. In Proceedings of the International
Conference on Artificial Neural Networks, Rhodes, Greece, 4–7 October 2018; Springer: Berlin/Heidelberg, Germany, 2018;
pp. 270–279.
73. Taylor, M.E.; Stone, P. Transfer learning for reinforcement learning domains: A survey. J. Mach. Learn. Res. 2009, 10, 1633–1685.
74. Peirelinck, T.; Kazmi, H.; Mbuwir, B.V.; Hermans, C.; Spiessens, F.; Suykens, J.; Deconinck, G. Transfer learning in demand
response: A review of algorithms for data-efficient modelling and control. Energy AI 2022, 7, 100126. [CrossRef]
75. Joshi, G.; Chowdhary, G. Cross-domain transfer in reinforcement learning using target apprentice. In Proceedings of the 2018
IEEE International Conference on Robotics and Automation (ICRA), Brisbane, Australia, 21–25 May 2018; pp. 7525–7532.
76. Zhu, Z.; Lin, K.; Dai, B.; Zhou, J. Learning sparse rewarded tasks from sub-optimal demonstrations. arXiv 2020, arXiv:2004.00530.
Processes 2022, 10, 2311 31 of 31
77. Yan, M.; Frosio, I.; Tyree, S.; Kautz, J. Sim-to-real transfer of accurate grasping with eye-in-hand observations and continuous
control. arXiv 2017, arXiv:1712.03303.
78. Christiano, P.; Shah, Z.; Mordatch, I.; Schneider, J.; Blackwell, T.; Tobin, J.; Abbeel, P.; Zaremba, W. Transfer from simulation to
real world through learning deep inverse dynamics model. arXiv 2016, arXiv:1610.03518.
79. Kostrikov, I.; Agrawal, K.K.; Dwibedi, D.; Levine, S.; Tompson, J. Discriminator-actor-critic: Addressing sample inefficiency and
reward bias in adversarial imitation learning. arXiv 2018, arXiv:1809.02925.
80. Spielberg, S.; Tulsyan, A.; Lawrence, N.P.; Loewen, P.D.; Bhushan Gopaluni, R. Toward self-driving processes: A deep
reinforcement learning approach to control. AIChE J. 2019, 65, e16689. [CrossRef]
81. Hausknecht, M.; Stone, P. Deep reinforcement learning in parameterized action space. arXiv 2015, arXiv:1511.04143.
82. Hou, Y.; Liu, L.; Wei, Q.; Xu, X.; Chen, C. A novel ddpg method with prioritized experience replay. In Proceedings of the 2017
IEEE international conference on systems, man, and cybernetics (SMC), Banff, AB, Canada, 5–8 October 2017; pp. 316–321.
83. Wang, X.; Ye, X. Consciousness-driven reinforcement learning: An online learning control framework. Int. J. Intell. Syst. 2022,
37, 770–798. [CrossRef]
84. Feise, H.J.; Schaer, E. Mastering digitized chemical engineering. Educ. Chem. Eng. 2021, 34, 78–86. [CrossRef]
85. Hua, J.; Zeng, L.; Li, G.; Ju, Z. Learning for a robot: Deep reinforcement learning, imitation learning, transfer learning. Sensors
2021, 21, 1278. [CrossRef]
86. Hussein, A.; Gaber, M.M.; Elyan, E.; Jayne, C. Imitation learning: A survey of learning methods. ACM Comput. Surv. (CSUR)
2017, 50, 1–35. [CrossRef]
87. Goodfellow, I.; Pouget-Abadie, J.; Mirza, M.; Xu, B.; Warde-Farley, D.; Ozair, S.; Courville, A.; Bengio, Y. Generative adversarial
nets. Adv. Neural Inf. Process. Syst. 2014, 27.
88. Hutter, F.; Hoos, H.H.; Leyton-Brown, K.; Stützle, T. ParamILS: An automatic algorithm configuration framework. J. Artif. Intell.
Res. 2009, 36, 267–306. [CrossRef]
89. Hutter, F. Automated Configuration of Algorithms for Solving Hard Computational Problems. Ph.D. Thesis, University of British
Columbia, Vancouver, BC, Canada, 2009.
90. Coates, A.; Ng, A.Y. The importance of encoding versus training with sparse coding and vector quantization. In Proceedings of
the 28th International Conference on Machine Learning (ICML), Washington, DC, USA, 28 June–2 July 2011.
91. Coates, A.; Ng, A.; Lee, H. An analysis of single-layer networks in unsupervised feature learning. In Proceedings of the
Fourteenth International Conference on Artificial Intelligence and Statistics, Fort Lauderdale, FL, USA, 11–13 April 2011;
pp. 215–223.
92. Akiba, T.; Sano, S.; Yanase, T.; Ohta, T.; Koyama, M. Optuna: A next-generation hyperparameter optimization framework. In
Proceedings of the 25th ACM SIGKDD International Conference on Knowledge Discovery & Data Mining, Anchorage, AK, USA,
4–8 August 2019; pp. 2623–2631.
93. Rapin, J.; Teytaud, O. Nevergrad—A Gradient-Free Optimization Platform. 2018. Available online: [Link]
FacebookResearch/Nevergrad (accessed on 10 September 2022).
94. Liaw, R.; Liang, E.; Nishihara, R.; Moritz, P.; Gonzalez, J.E.; Stoica, I. Tune: A Research Platform for Distributed Model Selection
and Training. arXiv 2018, arXiv:1807.05118.
95. Bergstra, J.S.; Bardenet, R.; Bengio, Y.; Kégl, B. Algorithms for hyper-parameter optimization. In Proceedings of the Advances in
Neural Information Processing Systems, Granada, Spain, 12–15 December 2011; pp. 2546–2554.
96. Snoek, J.; Larochelle, H.; Adams, R.P. Practical bayesian optimization of machine learning algorithms. In Proceedings of the
Advances in Neural Information Processing Systems, Lake Tahoe, NA, USA, 3–6 December 2012; pp. 2951–2959.
97. Li, L.; Jamieson, K.; Rostamizadeh, A.; Gonina, E.; Hardt, M.; Recht, B.; Talwalkar, A. Massively parallel hyperparameter tuning.
arXiv 2018, arXiv:1810.05934.
98. Li, L.; Jamieson, K.; DeSalvo, G.; Rostamizadeh, A.; Talwalkar, A. Hyperband: A novel bandit-based approach to hyperparameter
optimization. J. Mach. Learn. Res. 2017, 18, 6765–6816.
99. Jaderberg, M.; Dalibard, V.; Osindero, S.; Czarnecki, W.M.; Donahue, J.; Razavi, A.; Vinyals, O.; Green, T.; Dunning, I.; Simonyan,
K.; et al. Population based training of neural networks. arXiv 2017, arXiv:1711.09846.
100. Bergstra, J.; Bardenet, R.; Kégl, B.; Bengio, Y. Implementations of algorithms for hyper-parameter optimization. In Proceedings of
the NIPS Workshop on Bayesian Optimization, Sierra Nevada, Spain, 16–17 December 2011; p. 29.
101. Kingma, D.P.; Ba, J. Adam: A method for stochastic optimization. arXiv 2014, arXiv:1412.6980.
102. Das, L.; Sivaram, A.; Venkatasubramanian, V. Hidden representations in deep neural networks: Part 2. Regression problems.
Comput. Chem. Eng. 2020, 139, 106895. [CrossRef]