0% found this document useful (0 votes)
45 views42 pages

Adaptive Portfolio Trading with RRL

This manuscript presents an adaptive portfolio trading system that utilizes recurrent reinforcement learning (RRL) to optimize risk-return portfolio allocation with a focus on expected maximum drawdown (E(MDD)). The proposed system demonstrates superior performance compared to traditional methods, validating its effectiveness across various transaction costs and market conditions. The authors recommend an automated portfolio rebalancing system that adapts to market volatility and transaction costs for improved trading outcomes.
Copyright
© All Rights Reserved
We take content rights seriously. If you suspect this is your content, claim it here.
Available Formats
Download as PDF, TXT or read online on Scribd
0% found this document useful (0 votes)
45 views42 pages

Adaptive Portfolio Trading with RRL

This manuscript presents an adaptive portfolio trading system that utilizes recurrent reinforcement learning (RRL) to optimize risk-return portfolio allocation with a focus on expected maximum drawdown (E(MDD)). The proposed system demonstrates superior performance compared to traditional methods, validating its effectiveness across various transaction costs and market conditions. The authors recommend an automated portfolio rebalancing system that adapts to market volatility and transaction costs for improved trading outcomes.
Copyright
© All Rights Reserved
We take content rights seriously. If you suspect this is your content, claim it here.
Available Formats
Download as PDF, TXT or read online on Scribd

Accepted Manuscript

An adaptive portfolio trading system: A risk-return portfolio


optimization using recurrent reinforcement learning with expected
maximum drawdown

Saud Almahdi, Steve Y. Yang

PII: S0957-4174(17)30440-2
DOI: 10.1016/[Link].2017.06.023
Reference: ESWA 11394

To appear in: Expert Systems With Applications

Received date: 24 March 2017


Revised date: 29 May 2017
Accepted date: 14 June 2017

Please cite this article as: Saud Almahdi, Steve Y. Yang, An adaptive portfolio trading system: A risk-
return portfolio optimization using recurrent reinforcement learning with expected maximum drawdown,
Expert Systems With Applications (2017), doi: 10.1016/[Link].2017.06.023

This is a PDF file of an unedited manuscript that has been accepted for publication. As a service
to our customers we are providing this early version of the manuscript. The manuscript will undergo
copyediting, typesetting, and review of the resulting proof before it is published in its final form. Please
note that during the production process errors may be discovered which could affect the content, and
all legal disclaimers that apply to the journal pertain.
ACCEPTED MANUSCRIPT

Highlights

• A reinforcement learning trading algorithm with expected drawdown risk


is proposed.

T
• The expected maximum drawdown is shown to improve portfolio signal

IP
generation.

• The effectiveness of the method is validated using different transaction

CR
costs.

• An adaptive portfolio rebalancing system with automated retraining is


recommended.

US
AN
M
ED
PT
CE
AC

1
ACCEPTED MANUSCRIPT

An adaptive portfolio trading system: A risk-return


portfolio optimization using recurrent reinforcement
learning with expected maximum drawdown

Saud Almahdia , Steve Y. Yanga,∗

T
a Financial Engineering Program, School of Business, Stevens Institute of Technology, 1

IP
Castle Point on Hudson, Hoboken, NJ 07030, USA

CR
Abstract

US
Dynamic control theory has long been used in solving optimal asset allocation
problems, and a number of trading decision systems based on reinforcement
learning methods have been applied in asset allocation and portfolio rebalanc-
AN
ing. In this paper, we extend the existing work in recurrent reinforcement
learning (RRL) and build an optimal variable weight portfolio allocation under
a coherent downside risk measure, the expected maximum drawdown, E(MDD).
M

In particular, we propose a recurrent reinforcement learning method, with a co-


herent risk adjusted performance objective function, the Calmar ratio, to obtain
ED

both buy and sell signals and asset allocation weights. Using a portfolio con-
sisting of the most frequently traded exchange-traded funds, we show that the
expected maximum drawdown risk based objective function yields superior re-
PT

turn performance compared to previously proposed RRL objective functions


(i.e. the Sharpe ratio and the Sterling ratio), and that variable weight RRL
long/short portfolios outperform equal weight RRL long/short portfolios under
CE

different transaction cost scenarios. We further propose an adaptive E(MDD)


risk based RRL portfolio rebalancing decision system with a transaction cost
AC

and market condition stop-loss retraining mechanism, and we show that the

∗ Corresponding author: Steve Y. Yang, Postal address: School of Business, Stevens


Institute of Technology, 1 Castle Point on Hudson, Hoboken, NJ 07030 USA. Tel.: +1 201
216 3394 Fax: +1 201 216 5385
Email addresses: salmahdi@[Link] (Saud Almahdi), [Link]@[Link]
(Steve Y. Yang)

Preprint submitted to Expert Systems with Applications June 15, 2017


ACCEPTED MANUSCRIPT

proposed portfolio trading system responds to transaction cost effects better


and outperforms hedge fund benchmarks consistently.
Keywords: Recurrent reinforcement learning, Expected maximum drawdown,
Optimal portfolio rebalancing, Downside risk

T
IP
1. Introduction

In financial investing, a general goal is to dynamically allocate a set of as-

CR
sets to maximize the returns over time and minimize risk simultaneously. For
investors it is essential to be able to invest in a portfolio that can satisfy their

US
preset goals by building an optimal portfolio initially and subsequently rebal-
ancing it optimally. Portfolio theory began with mean-variance optimization
by Markowitz (1952) where he proposed portfolio selection by maximizing the
AN
expected return while minimizing risk in the form of covariance matrices. Re-
balancing a portfolio re-optimizes the weights of the portfolio over a predefined
time horizon. The application of dynamic asset allocation using dynamic pro-
M

gramming methods was originally introduced by Bertsekas (1995). Due to the


curse of dimensionality in dynamic programming, automated self learning al-
gorithms are normally applied by investors and scholars in designing optimal
ED

trading strategies instead. The reinforcement learning method is a type of ap-


proximate dynamic programming and a subcategory of machine learning intro-
PT

duced by Sutton & Barto (1998), and has been broadly applied by investors and
researchers in building strategic asset allocation decision systems (Gold, 2003a;
Dempster & Leemans, 2006; Tan et al., 2011; Feuerriegel & Prendinger, 2016).
CE

In this paper, we apply the recurrent reinforcement learning (RRL) method


with a statistically coherent downside risk adjusted performance objective func-
AC

tion to simultaneously generate both buy/sell signals and optimal asset alloca-
tion weights. Moody et al. (1998) introduced recurrent reinforcement learning in
building a trading system where they examined the performance effect between
using the Sharp ratio vs. several economic utility functions. They concluded
that the Sharpe ratio behaves like an adaptive utility function, and when maxi-

3
ACCEPTED MANUSCRIPT

mizing the differential Sharpe ratio as immediate rewards in an online learning


mode, the Sharpe ratio significantly outperforms the ones maximizing profits
directly. Most of the subsequent work (Gold, 2003b; Maringer & Ramtohul,
2010, 2012; Pu et al., 2016) focused on equally weighted portfolios. Although

T
both Moody et al. (1998) and Bertoluzzo & Corazza (2008) mentioned potential

IP
drawdown effects on RRL performance, neither thoroughly examined the actual
effects.

CR
Many practitioners tend to adjust the commonly accepted theoretical models
to apply them to their particular situations, or develop measures that focus on
their specific interests. They often neglect the theoretical aspects or assumptions

US
of their adjustments, such as in the safety-first risk measures (i.e. the Sharpe
ratio, the Sortino ratio, the Sterling ratio, and the Calmar ratio). Bhansali
(2007) and Zimmermann et al. (2003) noted that many risk measures based on
AN
estimation of covariance matrices using historical data failed notoriously when
they are needed the most. They agreed that the difference in volatility and
correlations between up and down market environments implies the risk reduc-
M

tion potential is limited leaving them incapable of foreseeing stress-type events.


We argue that large drawdowns usually lead to fund redemption, and hence
ED

they should lead to very different optimal decisions. In this paper, we extend
the variable weight RRL long only approach by Moody & Saffell (2001) to a
long-short approach and examine the expected maximum drawdown E(MDD)
PT

(Magdon-Ismail & Atiya, 2004) effect on portfolio performance with joint inter-
action of transaction costs. Magdon-Ismail et al. (2003) and Magdon-Ismail &
CE

Atiya (2004) provided a statistically coherent downside risk measure, the Cal-
mar ratio with the expected maximum drawdown, which provides a theoretical
base for us to apply this downside risk measure as a differentiable objective func-
AC

tion in RRL. This E(MDD) based Calmar ratio (Magdon-Ismail et al., 2003)
is distinctly different from the exponential moving average drawdown approach
used by Moody & Saffell (2001).

4
ACCEPTED MANUSCRIPT

More specifically, we compare the Calmar ratio1 with the Sharpe ratio where
the risk adjusted measure of performance is calculated by the standard deviation
of the returns over a predefined time horizon. Furthermore, we use the recurrent
reinforcement learning method with two different objective functions through

T
which we incorporate different risk considerations. We show that the recurrent

IP
reinforcement learning with variable weight asset allocation gives a superior
performance when applied to a set of highly liquid exchange-traded funds (ETF)

CR
with various transaction cost considerations over a 5 year period. We also
document that when the expected maximum drawdowns are considered, the
RRL can generate a superior portfolio to the ones generated by the average

US
deviation performance measure - the Sharpe ratio. This confirms the intuition
that a reasonably low MDD is critical to the success of any fund.
In addition, we propose a portfolio allocation and rebalancing system using
AN
RRL with E(MDD) as the performance measure, and this trading system jointly
considers transaction costs and market conditions to automatically retrain the
system parameters to achieve better performances. We show that a trading
M

system with the stop-loss based on market volatility regime is able to make
the portfolio endure higher transaction costs in that the stop-loss strategy will
ED

exit the market when the volatility is high and retrain the parameters of the
signal generating process and generate new signals to reenter the market. Such
a trading decision system is adaptive to the market conditions and is more
PT

resilient to transaction cost shocks.


The rest of the paper is organized as follows. In Section 2, we review existing
CE

work on dynamic portfolio optimization using reinforcement learning methods.


We introduce the expected maximum drawdown and its application to RRL in
Section 3. We apply the RRL based portfolio rebalancing approach to a set
AC

of ETFs to compare the cost effect of the Sharpe ratio vs. the Calmar ratio

1 While there exist multiple definitions of the Sterling ratio, it measures return over max-
imum drawdown-10%, versus the Calmar ratio, which is similar to the Sterling ratio but
normally applied to a 3 year period using a maximum drawdown.

5
ACCEPTED MANUSCRIPT

using RRL in Section 4. Section 5 conducts a final analysis comparing the


performance of the proposed risk-return portfolio optimization with that of two
hedge fund indices, and Section 6 concludes the study and identifies some future
work.

T
IP
2. Literature review

Machine learning algorithms are widely used for financial market prediction

CR
and portfolio constructions, especially for automated trading strategies. Sutton
et al. (1992) first introduced the reinforcement learning method (Q-learning)
and provided its analytically proven capabilities for one class of adaptive optimal

US
control problems. Recurrent reinforcement learning was introduced by Moody
et al. (1998) where it was applied to stock trading as a learning algorithm and
AN
they extended a single stock trading into a long only portfolio optimization
method using the recurrent reinforcement learning where they used the defer-
ential of the Sharpe ratio as the objective function. In Moody & Saffell (2001)
M

with a direct reinforcement alteration, the authors compared their method with
Q-learning and temporal difference algorithms using real data and showed that
the deferential Sharpe ratio recurrent reinforcement learning system outper-
ED

forms Q-learning. The researchers also proposed the deferential Sterling ratio
as the performance criterion. However, this version of Sterling ratio neutralizes
PT

the downside risk through exponential smoothing.


As a result, a number of trading strategies have been proposed based on re-
current reinforcement learning methods to address issues such as different asset
CE

classes, transaction cost and market regime change. Gold (2003b) discusses the
application of recurrent reinforcement learning in the foreign exchange market
and proposed a two layer network. He compared it with a one layer network
AC

and found that the one layer network outperformed the two layer network due
to noisy financial data. Others added different algorithms to the recurrent rein-
forcement learning. Maringer & Ramtohul (2010, 2012) added regime switching
to the recurrent reinforcement learning where regime switching captures the

6
ACCEPTED MANUSCRIPT

different movements of the stock price over time. They added an additional
regime switching model to the recurrent reinforcement learning to capture the
non-linearity of the financial market and proposed two different methods based
on the regime switching: threshold recurrent reinforcement learning (TRRL)

T
in Maringer & Ramtohul (2010) and the smooth transition recurrent reinforce-

IP
ment learning (STRRL) in Maringer & Ramtohul (2012). The authors compared
TRRL and STRRL with RRL and used different transaction costs to show the

CR
performance of the models. They concluded that the regime switching recurrent
reinforcement learning matches the normal recurrent reinforcement learning in
a dataset having a single regime but it outperforms the RRL when the dataset

US
has distinctly different regime characteristics.
In addition, another strand of literature combines the recurrent reinforce-
ment learning method with other machine learning methods, such as genetic
AN
programming and neural networks. Pu et al. (2016) used a genetic algorithm
to improve recurrent reinforcement learning for equity trading where the RRL
is population-based and the trading system consists of a group of simulation
M

traders. The genetic algorithm (GA) is the selector and the recurrent reinforce-
ment learning is the trading system; the goal is to achieve the optimal combi-
ED

nation where the chosen indicators are exported to the recurrent reinforcement
learning trading system. The authors find that the GA-RRL system is more
stable than the buy and hold strategy but it did not outperform the buy and
PT

hold strategy in terms of producing positive Sharpe ratio means. Gorse (2010)
worked on transforming the recurrent reinforcement learning into a stochastic
CE

learning process by using the stochastic gradient ascent in the optimization and
training technique. This improvement helped the process to be less expensive
computationally as there is no need to save the previous signals, but only the
AC

t − 1 signal. Hens & Wöhrmann Hens & Wöhrmann (2007) applied recurrent
reinforcement learning on a long-term equity and bond portfolio, assuming a
rational investor with a constant risk aversion and a power utility function.
Bertoluzzo & Corazza (2008) developed an artificial neural network based on
reinforcement learning algorithm using the reciprocal of the returns weighted

7
ACCEPTED MANUSCRIPT

direction symmetry index as the measure of profitability. They proposed a pro-


cedure for the management of drawdown like phenomena, and concluded that
one can take into account a drawdown like phenomenon in the learning process.
In general, most of the current studies on trading decision systems agree

T
that there are a number of critical components that need to be considered in

IP
developing such systems in addition to the core algorithms. If not adequately
designed, these factors can significantly compromise the advantages of any ad-

CR
vanced machine learning or artificial intelligence based trading decision systems.
Cavalcante et al. (2016) surveyed the computational intelligence methods, pro-
posed to solve financial market problems, from 2009 to 2015. They specify a

US
framework to be followed by most computational intelligence approaches, con-
sisting of the following major components: a) data preparation (input variables,
output variables, acquisition, prepossessing, normalization); b) algorithm def-
AN
inition (choose model, configure architecture); c) training (define algorithm,
adjust parameters, perform training); d) model evaluation (define metrics, eval-
uate accuracy); e) trading strategies; and f) money evaluation. In this survey,
M

the authors note that Chande (2001) identified three characteristics of success-
ful trading strategies: a) a rule set defining entering and exiting trades, b) a
ED

risk control method, and c) money management. Other characteristics include


taking into account real world constraints such as backtesting with real trans-
action costs and slippage. Martinez et al. (2009) developed a trading system
PT

based on forecasting where they used an artificial neural network (ANN) to


forecast the asset price, and then designed a trading system specifying the exit
CE

and entry rules based on the forecasts. A stop loss strategy based on placing
a threshold on negative returns was also incorporated to accommodate market
condition changes. In a paper by Beraldi et al. (2011), the authors proposed a
AC

trading support system to help investors solving the strategic asset allocation
problem. They focused on the system and its modules and stages rather than
the solution method. The modules in their system included data management,
statistical analysis, scenario simulation, a model generator, a solution kernel and
a solutions analysis module. Their decision system integrated many solution ap-

8
ACCEPTED MANUSCRIPT

proaches based on statistical and stochastic optimization with a Monte Carlo


simulation of the scenarios. Eilers et al. (2014) developed an automated trading
decision system where they combined an artificial neural network with reinforce-
ment learning (RL) and seasonality. The authors trained the ANN using the

T
value iteration method of the RL while only optimizing the immediate reward.

IP
This method simplifies the optimal value function to be only the immediate re-
ward. The three layers include input neurons, hidden neurons, and one output

CR
neuron; and the feed-forward ANN is trained by minimizing the mean squared
error using the back-propagation algorithm. Feuerriegel & Prendinger (2016)
proposed a trading system based on news disclosures where the authors design

US
trading strategies that utilize textual news to obtain profits on the basis of novel
information entering the market. They developed a system for automated deci-
sion making using supervised and reinforcement learning. The system contained
AN
two main components: news sentiment extraction and trading strategy execu-
tion. They concluded that a trading system can be improved with additional
novel market information.
M

3. Data and methodology


ED

3.1. Methodology

In this paper, we use the recurrent reinforcement method in portfolio op-


PT

timization with different risk considerations through two objective functions.


Following Moody et al. (1998), we use the differential Sharpe ratio for dynamic
optimization of trading system performance. We use performance functions
CE

both to increase the convergence of the learning process and to adapt to chang-
ing market conditions during live trading. During this process, the parameter
AC

updates can be done during each forward pass through the training data, and
the influence of the performance measure can be computed at any time point.
We also assume a small or medium investor who can take fixed or variable sizes
of shares of each asset with no price impact in the market but with a fixed

9
ACCEPTED MANUSCRIPT

transaction cost.
Ft = tanh(x0t θ) (1)

T
IP
CR
US
AN
Figure 1: Recurrent reinforcement learning with portfolio allocation signals.

Before describing the recurrent reinforcement learning approach, we start


by comparing the different safety-first risk measures (i.e. the Sharpe ratio, the
M

Sortino ratio, the Calmar ratio, and the Sterling ratio). The Sharpe ratio is
widely used and it was developed based on mean-variance optimization; there-
ED

fore, its risk measure is the standard deviation of returns. The Sortino ratio
uses the downside deviation and it includes a target for the investment return.
The Sterling ratio is used by Moody & Saffell (2001) where the authors differ-
PT

entiated the Sterling ratio using an exponential moving average. Our definition
of the Calmar ratio includes the expected maximum drawdown which is defined
CE

by Magdon-Ismail & Atiya (2004), while the basic definition of the Calmar ratio
is similar to the Sterling ratio in that both use the maximum drawdown. The
issue here is that the Sterling ratio is purely empirical, depending on the dataset
AC

it is applied to, and lacking the analytical properties needed in RRL. We choose
to use the Calmar ratio defined using the expected maximum drawdown because
it is consistent, coherent and differentiable by definition as it can be seen from
the definition of the expected maximum drawdown. The relation between the
Sharpe and Calmar ratios is shown in Figs. (2) and (3). While Fig. (2) shows

10
ACCEPTED MANUSCRIPT

the relation between the Sharpe ratio and the expected maximum drawdown
per standard deviation unit, Fig. (3) shows the relation between the Calmar
ratio with different Sharpe ratio values and a moving time step. From this
construction, the Calmar ratio we use here is consistent with Sharpe ratio in a

T
nonlinear way, and it is a statistically coherent risk measure as the Sharpe ratio.

IP
Its long-term difference from the Sharpe ratio is evident (see Fig. 3).

CR
US
AN
M

Figure 2: E(MDD) per unit σ vs. the Sharpe ratio.


ED

The Calmar ratio (CR) is similar to the Sharpe ratio (SR) in that it is also
a risk adjusted measure of performance. However, it is an MDD risk metric
PT

that measures the maximum cumulative loss from a peak to a following bottom.
When the downside losses are considered rather than the average deviation from
mean return, the trading decisions will certainly be different. Here we derive
CE

reward based on this E(MDD) risk based measure, the Calmar ratio. We use
the deferential of the Calmar ratio as our objective function:
AC

γT
CT = (2)
E(M DD)

11
ACCEPTED MANUSCRIPT

T
IP
CR
US
Figure 3: Scaling the Calmar ratio with time.


AN

 2σ 2 γ 2 T T →∞ σ 2

 γ Qp ( 2σ 2 ) → γ (0.63519




+0.5 log T + log γ ) ifγ > 0
σ
E(M DD) = √


M


1.2533σ T ifγ = 0




 −2σ2 Q ( γ 2 T ) T →∞
→ −γT − σ
2
ifγ < 0
γ n 2σ 2 γ

T
ED

2 Shrp2 T →∞ T Shrp2
CT = →
Qp ( T2 Shrp2 ) 0.63519 + 0.5 log T + log Shrp
PT

where CT is the Calmar ratio over the time horizon T , E(M DD) is the ex-
pected maximum drawdown. The functions Qn (x) and Qp (x) are complicated
integral expansions that do not have a convenient analytical form and they are
CE

independent from γ, σ and T . Their forms can be found in Magdon-Ismail et al.


(2003) and in Pratap (2004). γ is the mean of the returns and σ is the standard
γ
deviation of returns, Shrp = σ.
AC

Next, we construct a variable weight portfolio with a reinforcement learning


signal. Let Ft ∈ {−1, 1} be the trading signal with only two values. For Ft > 0,
the investor would take a long position, and we set Ft = 1. For Ft < 0, the
investor would then take a short position, and we set Ft = −1. θ are the

12
ACCEPTED MANUSCRIPT

parameters that we want to train θ ∈ <M +2 where M is the time series that
we want to trade in. xt is a vector where xt = [1; rt ...rt−M ; Ft−1 ], rt is the log
return rt = log(pricet ) − log(pricet−1 ). We calculate the return at time t of our
position in Eq. (3):

T
Rt = µ ∗ [Ft−1 · rt − δ|Ft − Ft−1 |] (3)

IP
µ is the number of shares that is a constant number and it can be the maximum

CR
number of shares one can trade. δ is the transaction cost and it is also a constant.
Using a risk adjusted measure we will maximize the Sharpe ratio in Eq. (4):

E[Rt ]
ST =

US σ

where E[Rt ] is the mean of the return and σ is the standard deviation of the
(4)
AN
returns. The objective function is the deferential of the Sharpe ratio and it is
calculated by using the average of return and standard deviation of the return.
Using the chain rule we get Eq. (5):
M

T
1X
A= Rt
T t=1
ED

T
1X 2
B= R
T t=1 t
T
X dST dA
dST dST dB dRt dFt dRt dFt−1
{ }·{ }
PT

= + + (5)
dθ t=1
dA dRt dB dRt dFt dθ dFt−1 dθ
where we have:
dRt
CE

= −µδ · sgn(Ft − Ft−1 )


dFt
dRt
= −µ · rt + µδ.sgn(Ft − Ft−1 )
dFt−1
AC

dFt dFt−1
= (1 − tanh(x0t θ)2 ) · (xt + θ M +2 )
dθ dθ
dFt
The above equations conclude that dθ is recurrent, and the weights are up-
dated by the gradient ascent θi+1 = θi + ρ. dS
dθ , and ρ is the learning rate.
T

13
ACCEPTED MANUSCRIPT

The recurrent reinforcement learning can be used to optimize a variable weight


portfolio. First we change Eq. (1) for each asset to Eq. (6):

fit = logsig(x0it θ i ) (6)

T
where fit is the action on asset i at time t, and logsig is the log-sigmoid transfer
function. This then leads to the following equation:

IP
dFt dFt−1
= (1 − logsig(x0t θ)2 ) · (xt + θ M +2 ) (7)
dθ dθ

CR
For a long only portfolio with variable weights, Eq. (8) needs to be applied for
each asset:

US
Fit = softmax(fit )

where the sof tmax function is applied to the actions on all the assets. It will
Pn
assign weights to each asset and it already includes the constraint i Fit = 1,
(8)
AN
where n is the number of assets in the portfolio.
The decisions from training the parameters θ with a logsig activation func-
tion will result in the choice of a number in the interval [0, 1] using a log-sigmoid
M

function. The sof tmax function is applied to all the assets decision at time t,
and it will distribute asset allocation weights and assign the highest weight to
ED

the largest asset decision and redistribute the weights from high to low accord-
ing to the decision obtained from the logsig activation function. In our case,
the portfolio will act closely to the equal weight portfolio except if a decision on
PT

an asset is close to zero, then the sof tmax will redistribute the weight among
the other assets so we are switching the weights between the assets as we move
CE

forward with decisions. The sof tmax function is recommended for a portfolio
as discussed in Moody et al. (1998). Here we are defining clearly fit and we
are using the sof tmax out of the training model. The training and the initial
AC

decision are based on the logsig function within the model. This model is sensi-
tive to the number of assets as with more assets this model will act as an asset
selector by assigning variable weights. Due to the similar statistical properties
as the Sharpe ratio, the Calmar ratio based recurrent reinforcement learning as
an objective function will be similar but substitute ST with CT in Eqs. (4) and

14
ACCEPTED MANUSCRIPT

(5) accordingly.

3.2. Portfolio constraints

T
In portfolio optimization, practitioners consider some of the real world con-
straints in their optimization process such as cardinality constraint, floor and

IP
ceiling constraint, round-lot constraint, pre-assignment constraint and class con-

CR
straint. The cardinality constraint is used to limit the asset selection in the
portfolio to a K number of assets. The floor and ceiling constraints limit the
weight of each asset allocation to certain boundaries. The pre-assignment con-

US
straint allows the investor to pre-select a desired asset in the portfolio. The
round-lot constraint restricts the number of any asset to an exact multiple of
the normal trading lots. The class constraint limits the proportion invested in
AN
assets with common characteristics. In Chang et al. (2000), the authors pro-
posed three meta-heuristic algorithms (genetic algorithm, simulated annealing,
and particle swarm) to solve the portfolio constraint problems.
M

Many scholars followed Chang et al. (2000) and developed meta-heuristic


methods to solve the constrained portfolio optimization problems. Lwin et al.
ED

(2014) developed a learning guided multi-objective evolutionary algorithm to


solve the portfolio optimization problem with the cardinality constraint, floor
and ceiling constraint, pre-assignment constraint and round-lot constraint. The
PT

authors discussed that these constraints are hard to satisfy at any time as the
cardinality constraint by-itself is a mixed quadratic integer NP-hard problem
and the portfolio selection with round-lot constraint is an NP-complete prob-
CE

lem. The authors compared the performance of their proposed algorithm with
four different well-known multi-objective evolutionary algorithms (the Non-
AC

dominated Sorting Genetic Algorithm(NSGA-II), the Strength Pareto Evolu-


tionary Algorithm(SPEA-2), Pareto Envelope-based Selection Algorithm(PESA-
II), Pareto Archived Evolution Strategy(PAES)). The proposed method is com-
putationally efficient and yields a better result over all the four algorithms. In
a paper by Silva et al. (2015), the authors solved the constrained portfolio op-

15
ACCEPTED MANUSCRIPT

timization problem with cardinality constraint, quantity constraint, long only


constraint and transaction cost constraint by combining multi-objective evolu-
tionary (MOEA) algorithm with technical indicators, where the indicators are
determined by the MOEA algorithm and the selection method is adaptive as

T
the stocks selected are changing with time. The stocks are selected using funda-

IP
mental indicators, and the trading decisions are based on technical indicators.
During the testing phase on the S&P 500 stocks, the authors included a 2% of

CR
the stock value as a transaction cost. The proposed method outperforms the in-
dex in terms of returns and variance. Liagkouras & Metaxiotis (2016) suggested
that due to the intrinsic multi-objective nature of the constrained portfolio op-

US
timization problem, the multi-objective evolutionary algorithms proved to be
very useful and effective in handling the difficulties imposed by the problem in
a reasonable time. Chen et al. (2017) discussed the exact algorithms for solv-
AN
ing the constrained portfolio optimization problem where they mentioned that
the disadvantage of the exact algorithm is that it always needs more compu-
tational time and can find an optimal solution only in a specified time. The
M

authors then presented a heuristic approach which is an extension to the Non-


dominated Sorting and Local Search (NSLS) based multi-objective evolutionary
ED

framework and called it (e-NSLS) in order to solve the cardinality constrained


portfolio optimization problem. They compared their method with five different
algorithms (NSGA-II, SPEA-2, MOEA/D-DE, ABC-FC, GRASP-QUAD) and
PT

showed that the proposed method outperforms the other five algorithms in com-
putational results. Moreover, the authors used the Wilcoxon signed ranks test
CE

analysis to statistically test the significant performance of e-NSLS with the other
algorithms where the results show that the proposed method outperformed the
other algorithms.
AC

Although these constrained portfolio optimization problems are complex and


hard to solve (Chang et al., 2000; Moral-Escudero et al., 2006), there exist a
number of heuristic search based approaches in the current literature to help
practitioners to address their specific needs. In this paper, we primarily focus
on developing an effective RRL trading strategy and a trading system using dif-

16
ACCEPTED MANUSCRIPT

ferent objective functions in a dynamic portfolio optimization setting. We will


direct our attention in an unconstrained problem setting in the present study,
and yet we do not foresee major difficulties to combine our proposed approach
with the exiting heuristic portfolio constraint methods to address specific prac-

T
tical requirements. In fact, one could replace the gradient ascent search in the

IP
current RRL optimization with an evolutionary algorithm using a desirable ob-
jective function (e.g. Sharpe ratio or Calmar ratio) as the fitness function to

CR
optimize portfolio weights and constraints simultaneously. For future work, we
will combine evolutionary algorithms with our proposed Calmar RRL model to
investigate the benefit of introducing various portfolio constraints.

3.3. Data collection


US
In this study, we construct a five asset portfolio using five of the most com-
AN
monly traded exchange-traded funds from different asset categories. These as-
sets (identified by their ticker symbols and fund names) are as follows:
M

• IWD: iShares Russell 1000 Value

• IWC: iShares Micro-Cap


ED

• SPY: SPDR S&P 500 ETF

• DEM: WisdomTree Emerging Markets High Dividend


PT

• CLY: iShares 10+ Year Credit Bond

IWD ETF is an equity fund that holds mid and large-cap US stocks. This ETF
CE

tracks the performance of the Russell 1000 value index. IWC ETF is a fund
that seeks to correspond to the performance of the Russell micro-cap index. It
consists of small cap US based companies. SPY ETF tracks the S&P500 index.
AC

It represents all 500 stocks in the index and pays dividends on a quarterly basis.
DEM ETF is a fund that tracks the price and yield of the WisdomTree emerging
markets equity income index, and has an international geographical focus. The
CLY ETF is a fixed income asset class that follows the investment results of an

17
ACCEPTED MANUSCRIPT

index consisting of long-term US corporate bonds and dollar dominated bonds


with remaining maturities more than ten years. We extract the weekly closing
prices for each of five assets from Yahoo Finance using the fetch function in
MATLAB. The dates are from January 01, 2011 to December 31, 2015. We use

T
three years of weekly returns for training and two years for testing. Table (1)

IP
shows some statistical features of the assets selected over the total time horizon
of five years.

CR
Asset Mean Maximum Average
Name of returns drawdown volume

IWD
IWC
SPY
0.0015
0.0014
0.0018
US 0.1996
0.2714
0.1744
2,355,230
55,670
78,605,495
AN
DEM -0.0024 0.5199 315,405
CLY 0.0002 0.1602 185,953
M

Table 1: Statistical features of the ETFs


ED

4. Trading algorithms comparison

In this section, we first compare the performance of three performance ratios


PT

as three different objective functions for the model. This will result in three
trading algorithms producing different trading decisions for the same set of
CE

assets, and then we can readily assess the merits of each performance ratio in
generating trading signals. The resulting portfolio rebalancing methods are:
the Sharpe ratio RRL (SR-RRL), the Sterling ratio RRL (TR-RRL), and the
AC

Calmar ratio RRL (CR-RRL).


We show the comparison between the portfolios formulated using recurrent
reinforcement learning with the Sharpe ratio as the objective function vs. the
Calmar ratio as the objective function. In addition, we compare the recurrent
reinforcement learning based portfolios with the buy-and-hold strategy as a

18
ACCEPTED MANUSCRIPT

baseline benchmark. We use three years of weekly closing prices (January 01,
2011 - December 31, 2013) to train our θ for each asset, and two years of weekly
closing prices (January 01, 2014 - December 31, 2015) for testing. The value of
M is set to 104, the number of weeks in the two years of testing data. In order to

T
generate the signals and the weights from the recurrent reinforcement learning

IP
model, we set the number of evaluations for the tanh model to a maximum of
10000 and for the logsig model to 500. The logsig model at this stage is given

CR
a small alteration to the equally weighted portfolio weights due to the number
of assets.

US
4.1. Sharpe ratio recurrent reinforcement learning portfolios

In this model, the objective function that we need to optimize is the Sharpe
ratio. The Sharpe ratio is a performance measure that can be maximized by
AN
maximizing the mean of the return of the portfolio and minimizing the standard
deviation of the return. In the SR-RRL we are using the deferential of the Sharpe
ratio that is calculated by averaging the Sharpe ratio. The weights in the model
M

will be updated with respect to the gradient of the Sharpe ratio. The standard
deviation is used as a measure of volatility. The more the mean returns of
ED

the portfolio vary, the higher the volatility. In other words, the volatility will
increase if the mean of the returns is varying in a positive or negative direction.
v
u T
u1 X
PT

σ=t (Rt − γ)2 (9)


T t=1

The standard deviation should not be the only measurement of the risk of a given
CE

portfolio. For example, a fund accumulating a return between 4% and 6% on


average will have a lower standard deviation than a fund accumulating a return
AC

between 4% and 14% on average. In a portfolio with different types of assets


of different volatilities, the Sharpe ratio will be a setback to the performance of
the portfolio and the decision making process.
We use the recurrent reinforcement learning method with the deferential
Sharpe ratio as the objective function to obtain two different portfolios: the

19
ACCEPTED MANUSCRIPT

Sharpe Ratio Equally Weighted (SR-RRL EW) Long/Short (L/S) Portfolio, and
the Sharpe Ratio Variable Weights (SR-RRL VW) Long/Short (L/S) Portfolio.
We use Eq. (1) as the activation function to get the signals of each asset over
the training period, and we then use equal weights and apply the signals to the

T
equal weights. Let n = the number of assets, and the weight w = 1/n for each

IP
asset. This results in:
n
X
|Fit wit | = 1 (10)

CR
i

where i is the number of assets at time t.

In the combination of the above two portfolios, Ft in Eq. (10) is the signal

US
from Eq. (1) and wt is the weight from Eqs. (6) and (8). We select four port-
folios from the Markowitz efficient frontier shown in Fig. (4) and the equally
AN
weighted buy & hold portfolio to compare them with the different RRL portfo-
lios. Fig. (5) shows the SR-RRL equally weighted and variable weights portfolio
compared in terms of cumulative returns with the four portfolios from the ef-
M

ficient frontier using Markowitz mean-variance optimization namely (minimum


variance, maximum return, maximum Sharpe ratio and a Pareto optimal) and
the buy & hold portfolio. Both of the SR-RRL portfolios are outperformed by
ED

the Pareto optimal portfolio by the end of the investment horizon. In this test
we choose µ = 100 and δ = 0bp which is one basis point per stock traded. When
PT

examining the return per asset of the SR-equally weighted long/short portfo-
lio, it shows fluctuation of the asset returns. Since it is an equally weighted
portfolio, the portfolio is affected by each asset movement equally. The return
CE

per asset conclude that the DEM ETF is drawing down our portfolio as it is
the only asset that has strong drawdowns within the portfolio. Fig. (6) shows
the asset returns in the buy & hold strategy, indicating that the DEM ETF
AC

is causing sharp negative returns. By examining the return of each asset in


the SR-variable weights long/short portfolio we conclude that the portfolio acts
closely to the equally weighted portfolio. Since it is based on the same signals,
the minor changes are causing the portfolio to out-perform due to the higher

20
ACCEPTED MANUSCRIPT

weight of the SPY ETF until week 40 and IWD ETF in weeks 40-100. It is
placing less weight on CLY in weeks 60-80, which minimizes the loss caused by
the ETF.

T
IP
CR
US
Figure 4: Efficient frontier portfolios.
AN
M
ED

Figure 5: Portfolio performance RRL with Sharpe ratio (104 weeks).


PT

4.2. Sterling ratio recurrent reinforcement learning portfolios


TR-RRL is a model with the objective function as the differential of the
CE

Sterling ratio introduced in Moody & Saffell (2001). The output of the model
produces trading decisions to maximize the Sterling ration in Eq. (11). The
AC

change of the weights in the model will be based on the gradient of the Sterling
ratio. Let
γ
TR = (11)
M DD
where T R is the Sterling ratio, γ is the mean of the return, and M DD is the
maximum drawdown. The deferential of the Sterling ratio in Moody & Saffell

21
ACCEPTED MANUSCRIPT

T
IP
CR
US
Figure 6: Asset returns buy & hold strategy (104 weeks).
AN
(2001) is calculated empirically by using the exponential moving average of the
ratio. The maximum drawdown is calculated as follows in Eq. (12):
v
u T
u1 X
M

M DDT = t min[Rt , 0]2 (12)


T t=1

In order to train this model to have the same number of cycles as the others,
ED

we need to choose a number close to zero (e.g. 0.0001) to avoid division by zero
during training when evaluating the minimum of the return (Rt ) and zero. We
developed two portfolios: the Sterling Ratio Equally Weighted (TR-RRL EW)
PT

Long/Short (L/S) Portfolio, and the Sterling Ratio Variable Weight (TR-RRL
VW) Long/Short (L/S) Portfolio. In Fig. (7), the TR-RRL equally weighted
CE

and variable weights portfolios are compared with the four portfolios from the
efficient frontier and the buy & hold portfolio. The TR-RRL portfolios outper-
form all the portfolios by the end of the investment horizon, where µ = 100 and
AC

δ = 0bp.

4.3. Calmar ratio recurrent reinforcement learning portfolios

As in the previous experiment, the training set is three years of weekly


closing prices and the testing set is two years of weekly closing prices. We use the

22
ACCEPTED MANUSCRIPT

T
IP
CR
Figure 7: Portfolio performance RRL with Sterling ratio (104 weeks).

Calmar ratio (defined in Eq. (2)) instead of the Sharpe ratio to obtain the signals

US
and weights of our portfolios, and the objective function is the derivative of the
Calmar ratio. The difference between the two objective functions is that the
Calmar ratio is more sensitive to extreme losses while the Sharpe ratio considers
AN
average deviations. Our goal is to identify whether the large losses would make
differences in the dynamic optimization process. Using the expected maximum
drawdown in the Calmar ratio allows us to increase the number of function
M

evaluations to 10,000 because the expected maximum drawdown is based on


the mean and standard deviation of returns where by definition the expected
ED

maximum drawdown will not cause a division by zero error. On the other hand,
the basic Sterling ratio used by Moody will stop at some point due to a division
by zero error. We need to use a number that is close to zero in the maximum
PT

drawdown evaluation when computing the minimum of the return (Rt ) and
zero. In this experiment, we test the following two portfolios: the Calmar Ratio
CE

Equally Weighted (CR-RRL EW) Long/Short (L/S) Portfolio, and the Calmar
Ratio Variable Weights (CR-RRL VW) Long/Short (L/S) Portfolio.
In Fig. (8), we show the performance of the portfolios developed using the
AC

recurrent reinforcement learning and the differential Calmar ratio as the ob-
jective function where µ = 100 and δ = 0bp. Where the CR-RRL portfolio is
compared with the four portfolios from the efficient frontier and the buy & hold
portfolio. The CR-RRL portfolios are superior to the other portfolios over most

23
ACCEPTED MANUSCRIPT

of the investment horizon till the end of the horizon.

T
IP
CR
Figure 8: Portfolio performance Calmar ratio RRL (104 weeks).

US
In Table (2), we show the training computational time for each model where
we used a computer with multiple cores. The computational times are presented
in the hour, minute, and second format (HH:MM:SS). We observe that the
AN
differences between the methods in terms of computational time is minimal
where they differ only in a few minutes. Nevertheless, the TR-RRL method
has the least computational training time and the CR-RRL method has the
M

longest training time, and the SR-RRL method sits in between the other two.
Overall, the computational efficiency of all the strategies are not too far apart
ED

from each other since they are all based on the RRL method with the exception
of the objective function. The choosing of the objective function would affect
the calculation of the gradient and that would then affect the training of the
PT

parameters. In the case that the objective function cannot be increased in the
direction of it’s gradient, the algorithm may stop at a local maximum with no
CE

efficient training of the parameters θ.

4.4. Transaction cost sensitivity analysis


AC

It is well-established that when designing a realistic trading system, one


has to account for all transaction costs (Madhavan, 2002; Tetlock et al., 2008).
Although prior studies have conducted trading simulations, many neglect the
influence of transaction costs. The primary reason for such omission is due
to the difficulty in estimating realistically different types of transaction costs

24
ACCEPTED MANUSCRIPT

Trading strategy Computational time (HH:MM:SS)

SR-RRL 02:39:39

TR-RRL 02:37:50

T
CR-RRL 02:43:19

IP
Table 2: Model training computation time using eight core processors, 8GB RAM, Windows-
64bit, MATLAB

CR
involved, and these costs most likely differ for different asset classes and depend
on many other market characteristics. In this section, we examine the impact

US
of trading costs on the profitability of different portfolio strategies. Empirical
evidence shows that the average round-trip trading cost of large-cap stocks on
AN
NYSE is at least 20 bps (Chan & Lakonishok, 1997; Keim & Madhavan, 1998;
Mittermayer, 2004). In the cost sensitivity analysis, we applied one-way trading
costs of 10, 15, 20, and 25 bps. Even with realistic transaction costs of 10 bps
per round-trip, the portfolio strategies are superior to the hedge fund industry
M

index performance. For transaction costs of 15 bps, the Calmar ratio RRL
strategy is on average still profitable, but sustains a substantial loss potential.
ED

In general, these strategies cannot compensate transaction costs of more than


20 bps. It requires a system level design to accommodate high transaction costs
and further improve portfolio performances, which we will discuss in the next
PT

section.
In Fig. (9), we compare the Calmar ratio RRL portfolios performance with
CE

the Sharpe ratio RRL portfolios. Due to transaction costs, the Calmar ratio
RRL outperforms the Sharpe ratio RRL in terms of cumulative returns. We
can conclude from the signals generated that the Calmar ratio RRL portfolios
AC

are changing positions in some assets less frequently than the Sharpe ratio RRL
portfolios. This is clearer with the CLY ETF where the portfolio suffers from
losses due to transaction costs. The consistency in signals means that we are
holding the asset at the same position for a longer period of time which reduces

25
ACCEPTED MANUSCRIPT

T
IP
L/S
Sharpe Return (%) Maximum Num.

CR
portfolio
accumulative of
(δ = 0bps) ratio drawdown
(annualized) trades

Sharpe ratio
RRL EW
Sharpe ratio
3.1229
US6.31
(3.11)
8.24
0.5080 230
AN
2.7630 0.4363 230
RRL VW (4.04)

Sterling ratio 12.99


2.3780 0.4699 235
M

RRL EW (6.29)
Sterling ratio 13.83
2.2816 0.5262 235
RRL VW (6.69)
ED

Calmar ratio 16.39


2.3245 0.2721 221
RRL EW (7.88)
PT

Calmar ratio 19.65


2.1781 0.2675 221
RRL VW (9.39)
CE

Table 3: Long-short portfolio comparison with different learning objective functions. Note:
EW and VW represent Equally Weighted and Variable Weight strategies, respectively.
AC

26
ACCEPTED MANUSCRIPT

the transaction cost. Figs. (10) and (11) show clearly the differences of signals
generated for the CLY ETF using the Calmar ratio and the Sharpe ratio, re-
spectively. This is reasonable because the Calmar ratio based objective function
is sensitive to large losses which occur less frequently than in the Sharpe ratio.

T
As a result, due to the frequent rebalancing signals generated from Sharpe ratio

IP
objective function, transaction costs are relatively higher compared with the
Calmar ratio portfolios.

CR
L/S
Sharpe Return (%) Maximum Num.
portfolio

(δ = 10bps)

Sharpe ratio
ratio
US
accumulative
(annualized)

-2.77
drawdown
of
trades
AN
1.6926 0.9771 230
RRL EW (-1.39)
Sharpe ratio -1.55
2.0105 0.8575 230
RRL VW (-0.78)
M

Sterling ratio 3.67


1.8071 0.9968 235
RRL EW (1.82)
ED

Sterling ratio 3.46


1.5061 1.0000 235
RRL VW (1.71)
PT

Calmar ratio 7.63


2.5760 0.4914 221
RRL EW (3.74)
Calmar ratio 9.75
CE

2.3367 0.4692 221


RRL VW (4.76)

Table 4: Long-short portfolio comparison with different learning objective functions and trans-
AC

action costs (δ = 10bps). Note: EW and VW represent Equally Weighted and Variable Weight
strategies, respectively.

Table 3 shows the Sharpe ratio of each portfolio through backtesting with
both the Sharpe ratio, Sterling ratio and the Calmar ratio as objective functions.

27
ACCEPTED MANUSCRIPT

Overall, the Calmar ratio portfolios outperform the Sharpe ratio and Sterling
ratio portfolios consistently. When the transaction cost increases, the perfor-
mance of the Sharpe ratio based portfolios decreases, while the Calmar ratio
based portfolios maintain almost the same performance (see Tables 4 and 5).

T
The Calmar ratio portfolios are actually increasing in Sharpe ratio measure due

IP
to the fact that the transaction cost is affecting the standard deviation of the
returns more than the mean of the returns due to the high returns generated by

CR
the portfolio with a low standard deviation. The Sterling ratio based portfolios
perform between the Sharpe ratio and the Calmar ratio portfolios both with
and without transaction costs. When the transaction cost increases to 15 bps,

US
both the Sharpe ratio and Sterling ratio portfolios start to generate negative
annualized returns. Under no transaction cost, the Calmar ratio based portfo-
lio retains high performance. Under a high transaction cost (δ = 20bps), the
AN
Sharpe ratio based portfolio suffers a large loss, while the Calmar ratio based
portfolios performance is impacted only by a slight decrease in returns. In Sec-
tion 5, we propose a stop-loss control to the system to limit the transaction cost
M

effect (see Table 6).


ED
PT
CE
AC

Figure 9: Comparison in performance µ = 100, δ = 0bp (104 weeeks).

28
ACCEPTED MANUSCRIPT

T
IP
L/S
Sharpe Return (%) Maximum Num.
portfolio

CR
accumulative of
(δ = 15bps) ratio drawdown
(annualized) trades

Sharpe ratio
RRL EW
Sharpe ratio
0.4669

0.6165
US -7.31
(-3.72)
-6.45
0.9960

0.9878
230

230
AN
RRL VW (-3.28)

Sterling ratio -0.99


0.3943 0.9931 235
RRL EW (-0.50)
M

Sterling ratio -1.73


0.1044 0.9769 235
RRL VW (-0.87)
ED

Calmar ratio 3.25


2.6433 0.7116 221
RRL EW (1.61)
Calmar ratio 4.79
PT

2.3925 0.6619 221


RRL VW (2.37)
CE

Table 5: Long-short portfolio comparison with different learning objective functions and trans-
action costs (δ = 15bps). Note: EW and VW represent Equally Weighted and Variable Weight
strategies, respectively.
AC

29
ACCEPTED MANUSCRIPT

T
IP
L/S
Sharpe Return (%) Maximum Num.
portfolio

CR
accumulative of
(δ = 20bp) ratio drawdown
(annualized) trades

Sharpe ratio
RRL EW
Sharpe ratio
-0.2281

-0.2551
US-11.85
(-6.11)
-11.35
0.9661

0.9630
230

230
AN
RRL VW (-5.84)

Sterling ratio -5.65


-0.5334 0.9973 235
RRL EW (-2.87)
M

Sterling ratio -6.91


-0.7198 0.9869 235
RRL VW (-3.52)
ED

Calmar ratio -1.13


2.1424 0.8966 221
RRL EW (-0.57)
Calmar ratio -0.16
PT

2.1320 0.9990 221


RRL VW (-0.08)
CE

Table 6: Long-short portfolio comparison with different learning objective functions and trans-
action costs (δ = 20bps). Note: EW and VW represent Equally Weighted and Variable Weight
strategies, respectively.
AC

30
ACCEPTED MANUSCRIPT

T
IP
CR
US
Figure 10: CLY ETF Signals generated from RRL-Calmar.
AN
M
ED
PT
CE

Figure 11: CLY ETF Signals generated from RRL-Sharpe.

5. Trading system and discussion


AC

We develop an adaptive trading system based on the recurrent reinforcement


learning using three different objective functions. The recurrent reinforcement
learning system is a recursive learning system, where the system learns from
every output every time step. In this system, the trader can select an objec-

31
ACCEPTED MANUSCRIPT

tive function that would be the best for the assets of his portfolio. The system
parameters are trained based on the objective function desired. We have intro-
duced three objective functions and showed the difference based on a portfolio of
five commonly traded ETFs. In a paper by DeMiguel et al. (2009), the authors

T
showed that an equally weighted portfolio can be an efficient portfolio and they

IP
compared it to other strategies. In our trading system the default choice of the
portfolio weights is the equally weighted portfolio. In Fig. (12), we show the

CR
design of the trading system where the user will select the objective function
(the Sharpe Ratio, the Calmar Ratio, and the Sterling Ratio) that the RRL
system will maximize, and the assets along with the time frame T of the prices.

US
The user will also select the number of decision steps M where M < T . The
data of the assets will be gathered from Yahoo Finance. The RRL system will
learn and train the parameters using the historical returns of time T . After
AN
training, the system will allow the user to define the asset allocation from two
types of strategic asset allocation (Equally Weighted Portfolio (default), RRL
Defined Portfolio). The RRL system will output the long and short decisions of
M

each asset along with the strategic allocation. The system will ask the investor
if he would like to use the dynamic stop-loss exit strategy which will stop the
ED

trading and go to retraining the system again. If the investor does not want to
use the stop-loss then the output will be stored for the next use of the system
where it will continue to learn from the given outputs. The system is trained
PT
CE
AC

Figure 12: RRL based trading decision system.

32
ACCEPTED MANUSCRIPT

with a predefined transaction cost of δ = 10bps per share and µ = 100 with no
stop-loss during the training phase. In a real trading system, the investor would
be able to estimate their transaction costs based on their past trading records,
and these costs can change from period to period on the same set of assets. The

T
proposed system will then be able to adapt to these changes through retraining

IP
the system with a new cost estimation. The system recommends that the user
utilize the Calmar ratio as the objective function when δ ≥ 15bps per share,

CR
where using this objective function will help the system endure the transaction
cost effect. Also, if the investor is concerned about the drawdown of the port-
folio, the Calmar ratio is perfect due to the fact that the system will be trained

US
to minimize the expected maximum drawdown.
In Fig. (13), we compare the performance of the CR-RRL variable weights
(L/S) portfolio with Hedge Fund Research’s HFRI Equity Hedge Index (HFRIEHI)
AN
and Sunrise’s U.S. Equity Optimized Growth Program (SGUSOGP) hedge fund,
on a monthly basis over two years (2014-2015) with δ = 1bp and µ = 100.
Table 7 shows the comparison in performance between the CR-RRL variable
M
ED
PT
CE

Figure 13: CR-RRL variable weights portfolio vs. hedge funds (24 months).
AC

weights long/short portfolio, the HFRI Equity Hedge Index, and the Sunrise
U.S. equity hedge fund. This table shows that the CR-RRL portfolio is outper-
forming in terms of Sharpe ratio, annualized return and maximum drawdown.
In this comparison, we highlight the performance of the Calmar ratio RRL
system for investors’ portfolio designs.

33
ACCEPTED MANUSCRIPT

Portfolio Sharpe Return (%) Maximum


ratio Accu. (Ann.) drawdown

Calmar ratio
1.93 28.6 (13.4) 0.1474

T
RRL VW
SRUSOGP 1.71 9.29 (4.54) 0.7384

IP
HFRIEHI 1.17 0.82 (0.41) 0.8745

CR
Buy & Hold 0.62 -5.25 (-2.66) 0.9728

Table 7: Portfolio comparison with hedge funds and buy & hold strategy.

5.1. Dynamic stop-loss strategy


US
The stop-loss strategy used in the RRL trading decision system is a simple
AN
dynamic stop-loss strategy. The notion of the simple dynamic stop-loss is intro-
duced by Chevallier et al. (2012) where they applied it to a long-only portfolio.
Here we apply the concept in our trading system using the cumulative return
M

in Eq. (13):
rt−1
≤ −n (13)
σt−1
ED

where rt−1 is the cumulative return up to time t−1, σt−1 is the moving volatility
up to time t−1, and n is the number of volatility days prompting stop-loss. The
stop-loss is applied only during the testing phase at the decision making process;
PT

it is not used for training the parameters of the recurrent reinforcement learning
model. In Fig. (14), we show the stop-loss strategy effect on returns of the CR-
CE

RRL portfolio where the CR-RRL portfolio with stop-loss exits the market in
week 91 and stops trading. The system will retrain the parameters based on
the latest market movements. In Table (8), we show the Sharpe ratio with
AC

a stop-loss for CR-RRL equally weighted portfolio with different transactions


costs compared.
From Table (8), we see that the stop-loss will be able to make the portfolio
endure higher transaction costs in the case that δ ≥ 25bps as the stop-loss strat-
egy will exit the market when the volatility is high, and retrain the parameters

34
ACCEPTED MANUSCRIPT

Portfolio Sharpe Return (%) Maximum Num.


of
(δ = 0bps) ratio accu. (ann.) drawdown
trades

T
Stop-Loss 1.9915 29.65 (13.87) 0.2279 230
No Stop-Loss 2.1781 19.65 (9.39) 0.2675 221

IP
Portfolio Sharpe Return (%) Maximum Num.

CR
of
(δ = 10bps) ratio accu. (ann.) drawdown
trades

Stop-Loss 2.0852 19.27 (9.21) 0.3628 230


No Stop-Loss

Portfolio
2.3367

Sharpe
US
9.75 (4.76)

Return (%)
0.4692

Maximum
221

Num.
AN
of
(δ = 15bps) ratio accu. (ann.) drawdown
trades

Stop-Loss 2.1752 14.08 (6.81) 0.5083 230


M

No Stop-Loss 2.3925 4.79 (2.37) 0.6619 221

Portfolio Sharpe Return (%) Maximum Num.


ED

of
(δ = 20bps) ratio accu. (ann.) drawdown
trades
PT

Stop-Loss 2.2958 8.90 (4.35) 0.7844 230


No Stop-Loss 2.1320 -0.16 (-0.08) 0.9990 221
CE

Portfolio Sharpe Return (%) Maximum Num.


of
(δ = 25bps) ratio accu. (ann.) drawdown
trades
AC

Stop-Loss 2.1382 3.71 (1.84) 0.9550 230


No Stop-Loss 1.0273 -5.11 (-2.59) 0.9550 221

Table 8: Calmar ratio recurrent reinforcement learning (CR-RRL) portfolio comparison


with/without stop-loss and different transaction costs (δ).

35
ACCEPTED MANUSCRIPT

T
IP
CR
Figure 14: CR-RRL variable weightes with stop-loss (104 weeks).

of the model, and then generate new signals to reenter the market. The training

US
of the new parameters will be done using the latest returns until the exit point
to handle changing market conditions.
AN
6. Conclusion

In this paper, we use the recurrent reinforcement learning method to solve a


M

dynamic portfolio optimization problem where we develop four portfolios using


the RRL and compare them with each other and the buy & hold portfolio. We
use RRL methods to optimize the portfolio weights and rebalance the portfolio
ED

over a predefined time horizon. We compare the deferential of the Sharpe ratio
and the Calmar ratio as the objective functions in the recurrent reinforcement
learning process and examine the performance effect by the transaction costs.
PT

We compare the performance differences between the Sterling ratio proposed


by Moody & Saffell (2001) where they defined the downside risk as an exponen-
CE

tial moving average of drawdown. Due to its lack of necessary statistical proper-
ties, the Sterling ratio based RRL suffers computational breakdowns during the
optimization process. More importantly, it neutralizes the downside risks and
AC

therefore it is limited in reaching an optimal trading strategy. Through back-


testing of the constructed portfolio using ETFs, we conclude that: a) variable
weight long/short portfolios outperform the equally weighted long/short port-
folios; b) the RRL Calmar ratio based portfolios outperform the RRL Sharpe

36
ACCEPTED MANUSCRIPT

ratio based portfolios consistently; c) the E(MDD) RRL based trading system
with market condition stop-loss retraining responds to transaction cost effects
better and outperforms hedge fund benchmarks consistently. Overall, we show
that the portfolios constructed using RRL with the expected maximum draw-

T
down based Calmar ratio result in a significantly superior performance and are

IP
more transaction cost resilient than the portfolios constructed with the Sharpe
ratio.

CR
In addition, we propose an adaptive trading decision system based on the
proposed RRL portfolio rebalance strategies with both transaction costs and
market condition changes, and we show that the system consistently outper-

US
forms the benchmark and hedge fund industry average index. We specifically
demonstrate how this expected maximum drawdown based reinforcement learn-
ing approach can filter market noise and identify the significant trading signals,
AN
and how the trading decision system with transaction cost and stop-loss retrain-
ing can adapt to different market conditions.
For future studies, we plan to define a relative strength performance measure
M

using the expected maximum drawdown to minimize tracking errors with respect
to certain benchmarks. We also believe that logsig, sof tmax RRL model can
ED

be extended by adding layers to the model in order to help make a good variable
weight decision. The Calmar ratio using the expected maximum drawdown can
be applied in other reinforcement learning models and on a large set of asset
PT

classes.
CE

References

Beraldi, P., Violi, A., & De Simone, F. (2011). A decision support system for
strategic asset allocation. Decision Support Systems, 51 , 549–561.
AC

Bertoluzzo, F., & Corazza, M. (2008). Financial trading systems: Is recurrent


reinforcement learning the way? In Reflexing Interfaces: The Complex Co-
evolution of Information Technology Ecosystems (pp. 246–256). IGI Global.

37
ACCEPTED MANUSCRIPT

Bertsekas, D. P. (1995). Dynamic programming and optimal control volume 1.


Athena Scientific Belmont, MA.

Bhansali, V. (2007). Putting economics (back) into quantitative models. The


Journal of Portfolio Management, 33 , 63–76.

T
IP
Cavalcante, R. C., Brasileiro, R. C., Souza, V. L., Nobrega, J. P., & Oliveira,
A. L. (2016). Computational intelligence and financial markets: A survey and

CR
future directions. Expert Systems with Applications, 55 , 194–211.

Chan, L. K. C., & Lakonishok, J. (1997). Institutional Equity Trading Costs:


NYSE Versus Nasdaq. The Journal of Finance, 52 , 713–735.

US
Chande, T. S. (2001). Beyond technical analysis: How to develop and implement
a winning trading system volume 101. John Wiley & Sons.
AN
Chang, T.-J., Meade, N., Beasley, J. E., & Sharaiha, Y. M. (2000). Heuristics
for cardinality constrained portfolio optimisation. Computers & Operations
Research, 27 , 1271–1302.
M

Chen, B., Lin, Y., Zeng, W., Xu, H., & Zhang, D. (2017). The mean-variance
ED

cardinality constrained portfolio optimization problem using a local search-


based multi-objective evolutionary algorithm. Applied Intelligence, (pp. 1–
21).
PT

Chevallier, J., Ding, W., & Ielpo, F. (2012). Implementing a simple rule for
dynamic stop-loss strategies. The Journal of Investing, 21 , 111–114.
CE

DeMiguel, V., Garlappi, L., & Uppal, R. (2009). Optimal versus naive diversi-
fication: How inefficient is the 1/n portfolio strategy? Review of Financial
AC

Studies, 22 , 1915–1953.

Dempster, M. A., & Leemans, V. (2006). An automated fx trading system


using adaptive reinforcement learning. Expert Systems with Applications, 30 ,
543–552.

38
ACCEPTED MANUSCRIPT

Eilers, D., Dunis, C. L., von Mettenheim, H.-J., & Breitner, M. H. (2014).
Intelligent trading of seasonal effects: A decision support algorithm based on
reinforcement learning. Decision support systems, 64 , 100–108.

Feuerriegel, S., & Prendinger, H. (2016). News-based trading strategies. Deci-

T
sion Support Systems, 90 , 65–74.

IP
Gold, C. (2003a). Fx trading via recurrent reinforcement learning. In Computa-

CR
tional Intelligence for Financial Engineering, 2003. Proceedings. 2003 IEEE
International Conference on (pp. 363–370). IEEE.

Gold, C. (2003b). FX trading via recurrent reinforcement learning. In

US
IEEE/IAFE Conference on Computational Intelligence for Financial Engi-
neering, Proceedings (CIFEr) (pp. 363–370). volume 2003.
AN
Gorse, D. (2010). Application of stochastic recurrent reinforcement learning to
index trading. ESANN 2011 proceedings, 19th European Symposium on Ar-
tificial Neural Networks, Computational Intelligence and Machine Learning,
M

(pp. 123–128).

Hens, T., & Wöhrmann, P. (2007). Strategic asset allocation and market timing:
ED

A reinforcement learning approach. Computational Economics, 29 , 369–381.

Keim, D. B., & Madhavan, A. (1998). The cost of institutional equity trades.
PT

Financial Analysts Journal , 54 , 50–69.

Liagkouras, K., & Metaxiotis, K. (2016). A new efficiently encoded multi-


CE

objective algorithm for the solution of the cardinality constrained portfolio


optimization problem. Annals of Operations Research, (pp. 1–39).
AC

Lwin, K., Qu, R., & Kendall, G. (2014). A learning-guided multi-objective


evolutionary algorithm for constrained portfolio optimization. Applied Soft
Computing, 24 , 757–772.

Madhavan, A. N. (2002). Implementation of hedge fund strategies. Special


Issues (Hedge Fund Strategies: A Global Outlook), 2002 , 74–80.

39
ACCEPTED MANUSCRIPT

Magdon-Ismail, M., Atiya, A., Pratap, A., & Abu-Mostafa, Y. (2003). The
maximum drawdown of the brownian motion. In Computational Intelligence
for Financial Engineering, 2003. Proceedings. 2003 IEEE International Con-
ference on (pp. 243–247). IEEE.

T
Magdon-Ismail, M., & Atiya, A. F. (2004). Maximum drawdown. Risk Maga-

IP
zine, 17 , 99–102.

CR
Maringer, D., & Ramtohul, T. (2010). Threshold recurrent reinforcement learn-
ing model for automated trading. Lecture Notes in Computer Science (in-
cluding subseries Lecture Notes in Artificial Intelligence and Lecture Notes in

US
Bioinformatics), 6025 LNCS , 212–221.

Maringer, D., & Ramtohul, T. (2012). Regime-switching recurrent reinforce-


AN
ment learning for investment decision making. Computational Management
Science, 9 , 89–107.

Markowitz, H. (1952). Portfolio selection. The Journal of Finance, 7 , 77–91.


M

Martinez, L. C., da Hora, D. N., Palotti, J. R. d. M., Meira, W., & Pappa,
G. L. (2009). From an artificial neural network to a stock market day-trading
ED

system: A case study on the bm&f bovespa. In Neural Networks, 2009. IJCNN
2009. International Joint Conference on (pp. 2006–2013). IEEE.
PT

Mittermayer, M.-A. (2004). Forecasting intraday stock price trends with text
mining techniques. In system sciences, 2004. proceedings of the 37th annual
hawaii international conference on (pp. 10–pp). IEEE.
CE

Moody, J., & Saffell, M. (2001). Learning to trade via direct reinforcement.
IEEE transactions on neural Networks, 12 , 875–889.
AC

Moody, J., Wu, L., Liao, Y., & Saffell, M. (1998). Performance Functions and
Reinforcement Learning for Trading Systems and Portfolios. Science, 17 ,
441–470.

40
ACCEPTED MANUSCRIPT

Moral-Escudero, R., Ruiz-Torrubiano, R., & Suárez, A. (2006). Selection of


optimal investment portfolios with cardinality constraints. In Evolutionary
Computation, 2006. CEC 2006. IEEE Congress on (pp. 2382–2388). IEEE.

Pratap, A. (2004). Maximum drawdown of a Brownian motion and AlphaBoost:

T
a boosting algorithm. Ph.D. thesis California Institute of Technology.

IP
Pu, Q., Ananthanarayanan, G., Bodik, P., Kandula, S., Akella, A., Bahl, P.,

CR
Stoica, I., Mengistu, H., Lehman, J., Clune, J., Dulac-Arnold, G., Evans, R.,
Sunehag, P., Coppin, B., Shakirov, V., Zhang, J., Maringer, D., Graves, A.,
& Schmidhuber, J. (2016). Using a Genetic Algorithm to Improve Recurrent

US
Reinforcement Learning for Equity Trading. Computational Economics, 47 ,
421–434.
AN
Silva, A., Neves, R., & Horta, N. (2015). A hybrid approach to portfolio com-
position based on fundamental and technical indicators. Expert Systems with
Applications, 42 , 2036–2048.
M

Sutton, R. S., & Barto, A. G. (1998). Reinforcement learning: An introduction


volume 1. MIT press Cambridge.
ED

Sutton, R. S., Barto, A. G., & Williams, R. J. (1992). Reinforcement learning


is direct adaptive optimal control. IEEE Control Systems, 12 , 19–22.
PT

Tan, Z., Quek, C., & Cheng, P. Y. (2011). Stock trading with cycles: A fi-
nancial application of anfis and reinforcement learning. Expert Systems with
Applications, 38 , 4741–4755.
CE

Tetlock, P. C., Saar-Tsechansky, M., & Macskassy, S. (2008). More than words:
Quantifying language to measure firms’ fundamentals. The Journal of Fi-
AC

nance, 63 , 1437–1467.

Zimmermann, H., Drobetz, W., & Oertmann, P. (2003). Global asset allocation:
new methods and applications volume 197. John Wiley & Sons.

41

Common questions

Powered by AI

Advancements in artificial intelligence and machine learning have significantly addressed the curse of dimensionality in dynamic programming for financial portfolios by introducing approximate dynamic programming techniques like reinforcement learning. Recurrent reinforcement learning (RRL) reduces computational complexity by leveraging past data to optimize weight allocations in real-time, without exhaustive state-space exploration typical in traditional dynamic programming. Additionally, integrating RRL with other AI tools such as genetic algorithms provides a more efficient framework for capturing nuances in high-dimensional, stochastic environments, thus delivering more practical and adaptive financial solutions .

Recurrent reinforcement learning (RRL) differs from traditional reinforcement learning by incorporating a recurrent structure that allows it to capture temporal dependencies in financial data. This makes RRL particularly suited for financial trading systems where time series data exhibit non-linearity and regime changes. RRL models can dynamically adjust trading strategies based on observed patterns, unlike traditional models which may not account for temporal dependencies as effectively. In financial applications, RRL is often combined with regime-switching models to better handle different market conditions and improve prediction accuracy .

Different learning objective functions, such as those based on the Sharpe and Calmar ratios, significantly affect trading signals in long/short portfolio strategies. The Sharpe ratio tends to generate more frequent trading signals due to its focus on maximizing returns relative to volatility, which can lead to higher transaction costs. By contrast, the Calmar ratio prioritizes minimizing drawdowns, resulting in less frequent trades and a potentially better cost-per-trade efficiency. This focus means that using the Calmar ratio generally leads to holding positions longer with reduced transaction costs, thus maintaining portfolio stability and performance even amid volatile market conditions .

An adaptive trading system based on recurrent reinforcement learning provides the flexibility to adjust its strategies dynamically in response to changing market conditions, enhancing the robustness of trading decisions. The system can select the best-suited objective function for a portfolio by evaluating various performance metrics such as volatility, drawdowns, and transaction costs to optimize returns. This capability allows for tailored portfolio management that aligns with the investor's risk appetite and market performance goals, ensuring that trading decisions are data-driven and context-sensitive .

The main computational challenge with recurrent reinforcement learning (RRL) with regime switching is managing the increased complexity and computational load due to the non-linearity and dynamic nature of financial markets. Researchers address these challenges by applying techniques such as stochastic gradient ascent, which reduces the need to store extensive past data by focusing on immediate signals only. Additionally, advancements like threshold and smooth transition models help capture non-linearities more efficiently, thereby improving performance while attempting to keep computational demands manageable .

Considering downside risk measures such as expected maximum drawdown is crucial in financial portfolio management because they provide a clearer assessment of potential losses under adverse market conditions. These measures enhance decision-making by allowing managers to evaluate the risk of significant downturns beyond average volatility, improving the allocation of assets towards achieving smoother and more stable returns. Integrating downside risk metrics helps in optimizing portfolios by reducing exposure to high-risk investments, thereby aligning with the risk aversion goals of investors .

The Sterling and Calmar ratios handle transaction costs differently, impacting portfolio performance distinctly. The Sterling ratio tends to produce moderate performance between the Sharpe and Calmar ratios, often generating negative returns under high transaction costs due to its less robust handling of drawdowns. In contrast, the Calmar ratio is more resilient since it minimizes high-frequency trading resulting from transaction costs, thus maintaining stable performance even under increased costs. This trait makes Calmar-based strategies preferable when transaction costs are significant, as they preserve higher returns and a favorable Sharpe ratio by reducing transaction frequency .

The Calmar ratio provides advantages over the Sharpe ratio in portfolio optimization by being more sensitive to extreme losses. While the Sharpe ratio considers average deviations, the Calmar ratio focuses on the maximum drawdown, making it better suited to evaluate performance in conditions where avoiding large losses is a priority. This sensitivity helps in maintaining portfolio stability under adverse market situations and reduces transaction costs associated with frequent rebalancing seen in Sharpe ratio optimization .

Genetic programming enhances recurrent reinforcement learning (RRL) for equity trading by acting as a selector mechanism to evolve and optimize trading strategies. In recent studies, genetic programming is combined with RRL, where the genetic algorithm selects optimal parameters and indicators that are then used by the RRL system to execute trades. This combination aims to improve the stability of trading strategies compared to standard approaches such as buy and hold, although it doesn't always outperform in terms of positive Sharpe ratios .

The implementation of regime switching in recurrent reinforcement learning (RRL) models significantly enhances trading strategy outcomes by allowing the model to better capture financial market non-linearities and adapt to different market conditions. When the financial data exhibit distinct regime characteristics, regime switching RRL models outperform standard RRL models that assume a single regime. They improve accuracy in signal generation and asset allocation by dynamically adjusting strategies to various market phases, thus reducing risk and enhancing returns in volatile environments .

You might also like