0% found this document useful (0 votes)
5 views10 pages

Bandit Algorithm for Stock Picking

This document discusses improvements to multi-armed bandit algorithms and their application in stock-picking strategies. It outlines various fine-tuning methods to enhance the algorithm's performance, such as optimistic initial values and upper confidence bounds. The results indicate that while the algorithm may not consistently pick the highest-performing stock, it can still yield better returns compared to an equally weighted portfolio.

Uploaded by

nguoithamdo
Copyright
© All Rights Reserved
We take content rights seriously. If you suspect this is your content, claim it here.
Available Formats
Download as PDF, TXT or read online on Scribd
0% found this document useful (0 votes)
5 views10 pages

Bandit Algorithm for Stock Picking

This document discusses improvements to multi-armed bandit algorithms and their application in stock-picking strategies. It outlines various fine-tuning methods to enhance the algorithm's performance, such as optimistic initial values and upper confidence bounds. The results indicate that while the algorithm may not consistently pick the highest-performing stock, it can still yield better returns compared to an equally weighted portfolio.

Uploaded by

nguoithamdo
Copyright
© All Rights Reserved
We take content rights seriously. If you suspect this is your content, claim it here.
Available Formats
Download as PDF, TXT or read online on Scribd

Multi-Armed Bandits: Improvements and a

Stock-picking Application

00:00:06,600 --> 00:00:09,200


Welcome back! In this final lesson of

00:00:09,200 --> 00:00:12,700


Module 6, we are first going to describe several methods that

00:00:12,700 --> 00:00:15,300


allow us to fine-tune the bandit algorithm

00:00:15,300 --> 00:00:16,700


that we have developed so far.

00:00:17,500 --> 00:00:21,000


Second, we will develop an application of the bandit

00:00:21,700 --> 00:00:24,600


algorithm in the context of a stock-picking

00:00:24,600 --> 00:00:27,400


problem with daily frequency. That is, using

00:00:27,400 --> 00:00:30,000


the algorithms described in previous lessons, we are going

00:00:30,100 --> 00:00:33,400


to analyze the extent to which the bandit

00:00:33,400 --> 00:00:36,900


algorithm can lead to profitable trading strategies.

00:00:38,800 --> 00:00:41,300


Now, to improve the performance of

00:00:41,300 --> 00:00:44,300


the k-armed bandit algorithms, we can

MScFE | © 2022 - WorldQuant University – All rights reserved. Video Transcript | PAGE 1
00:00:44,300 --> 00:00:47,500
use several methods. Before briefly describing three

00:00:47,500 --> 00:00:50,700


specific fine-tuning methods, let us recall as

00:00:50,700 --> 00:00:53,600


we have mentioned a couple of times in Module 6, that we

00:00:53,600 --> 00:00:56,700


can devise methods that allow us to change the parameters

00:00:56,700 --> 00:01:00,000


that determine the exploration vs. exploitation

00:00:59,000 --> 00:01:01,100


balance,

00:01:02,100 --> 00:01:05,500


alpha and epsilon, as the agent learns

00:01:05,500 --> 00:01:08,300


or detects drastic changes in

00:01:08,300 --> 00:01:10,000


the rewards of actions.

00:01:10,900 --> 00:01:14,100


Aside from the adjustment of those parameters alpha and

00:01:13,100 --> 00:01:17,200


epsilon, we can also propose other fine-tuning

00:01:16,200 --> 00:01:19,500


techniques that may enhance the learning

00:01:19,500 --> 00:01:21,400


ability of the bandit algorithm.

00:01:22,700 --> 00:01:26,700


First, we can opt for "optimistic initial

MScFE | © 2022 - WorldQuant University – All rights reserved. Video Transcript | PAGE 2
00:01:25,700 --> 00:01:28,800
values". This method

00:01:28,800 --> 00:01:31,600


involves setting initial expected rewards

00:01:31,600 --> 00:01:34,700


to the actions that are very optimistic relative

00:01:34,700 --> 00:01:35,700


to our prior beliefs.

00:01:36,500 --> 00:01:39,800


By forcing this feature into the algorithm, the

00:01:39,800 --> 00:01:42,500


agent will forcefully put more emphasis on

00:01:42,500 --> 00:01:45,300


exploration instead of exploitation.

00:01:46,500 --> 00:01:49,700


Second, we have also stressed

00:01:49,700 --> 00:01:52,500


that epsilon greedy policies force the

00:01:52,500 --> 00:01:55,600


agent to explore different actions, but without

00:01:55,600 --> 00:01:58,400


any particular guidance or priority on

00:01:58,400 --> 00:02:01,700


what other actions to explore. To give priority

00:02:01,700 --> 00:02:04,800


to certain actions, we can use an upper confidence

00:02:04,800 --> 00:02:07,600


bound that adds a bias to the estimated measure

MScFE | © 2022 - WorldQuant University – All rights reserved. Video Transcript | PAGE 3
00:02:07,600 --> 00:02:10,100
of expected rewards, which increases the

00:02:10,100 --> 00:02:14,700


likelihood of choosing actions that have been relatively unexplored.

00:02:16,200 --> 00:02:20,300


A third fine-tuning method is to implement random

00:02:19,300 --> 00:02:21,300


actions

00:02:22,400 --> 00:02:25,200


based on the probabilities generated by a soft-max

00:02:25,200 --> 00:02:28,600


function. Instead of learning and

00:02:28,600 --> 00:02:31,900


updating an expected reward, we update for

00:02:31,900 --> 00:02:34,700


each action, the score that

00:02:34,700 --> 00:02:38,100


is used as an input into the soft-max function.

00:02:37,100 --> 00:02:40,600


We updated the score

00:02:40,600 --> 00:02:43,300


according to the reward obtained from the choice

00:02:43,300 --> 00:02:46,000


of a given action, in a form that is

00:02:46,800 --> 00:02:49,400


similar to the optimization algorithms that use gradient

00:02:49,400 --> 00:02:52,200


ascent to maximize an objective function.

MScFE | © 2022 - WorldQuant University – All rights reserved. Video Transcript | PAGE 4
00:02:53,600 --> 00:02:56,200
One important aspect to bear in

00:02:56,200 --> 00:02:59,900


mind is that all these fine-tuning methods can work well in the

00:02:59,900 --> 00:03:02,300


stationary case and have no relative

00:03:02,300 --> 00:03:06,200


advantage in non-stationary setups. In those

00:03:05,200 --> 00:03:08,800


cases, the reinitializing of

00:03:08,800 --> 00:03:12,400


changes in the parameters of the algorithm will

00:03:11,400 --> 00:03:14,600


be more effective in enhancing the

00:03:14,600 --> 00:03:16,000


performance of the algorithm.

00:03:17,300 --> 00:03:20,100


So, in the final application of this module,

00:03:20,900 --> 00:03:23,600


we are going to study the performance of an algorithm

00:03:23,600 --> 00:03:26,400


that prescribes a stock to

00:03:26,400 --> 00:03:28,800


pick among a pre-selected subset of stocks.

00:03:29,500 --> 00:03:32,500


We have chosen a set of 5 stocks in the technological

00:03:32,500 --> 00:03:36,100


sector that are very well known to us: Apple, Amazon, Microsoft,

MScFE | © 2022 - WorldQuant University – All rights reserved. Video Transcript | PAGE 5
00:03:35,100 --> 00:03:37,800
Nvidia and Twitter.

00:03:38,900 --> 00:03:41,800


The problems we are going to model is that of an

00:03:41,800 --> 00:03:43,200


investor who every day

00:03:43,900 --> 00:03:46,400


will purchase one of these five stocks at the

00:03:46,400 --> 00:03:49,700


closing price of each trading day and will sell

00:03:49,700 --> 00:03:53,200


it at the closing price of the next trading day.

00:03:54,300 --> 00:03:57,400


We aim to analyze the capacity of a valid algorithm to

00:03:57,400 --> 00:04:00,100


pick the right stock, the one that

00:04:00,100 --> 00:04:04,200


yields the highest returns across the five stocks selected.

00:04:03,200 --> 00:04:06,400


We are going to benchmark the returns

00:04:06,400 --> 00:04:09,200


of the strategy that arises from the bandit problem

00:04:09,200 --> 00:04:12,200


against the returns of an equally weighted portfolio of

00:04:12,200 --> 00:04:13,100


the five stocks.

00:04:14,400 --> 00:04:17,800


Because this problem is non-stationary in

MScFE | © 2022 - WorldQuant University – All rights reserved. Video Transcript | PAGE 6
00:04:17,800 --> 00:04:20,400
nature, expected returns of stocks

00:04:20,400 --> 00:04:23,200


tend to vary, so we are going to use a bandit algorithm that

00:04:23,200 --> 00:04:26,600


gives a large weight to more recent realizations of returns.

00:04:27,800 --> 00:04:30,900


Setting the value of alpha to 0.975.

00:04:32,600 --> 00:04:36,000


Moreover, we are going to allow for some exploration by

00:04:35,300 --> 00:04:38,600


fixing an epsilon parameter of 0.1 and by

00:04:38,600 --> 00:04:41,800


adding an upper confidence adjustment to

00:04:41,800 --> 00:04:42,700


the objective function.

00:04:44,600 --> 00:04:47,100


Now, here this figure depicts

00:04:47,600 --> 00:04:50,300


and summarizes the performance of

00:04:50,300 --> 00:04:53,900


the bandit algorithm to capture the best performing stocks

00:04:53,900 --> 00:04:56,700


over 1,000 episodes of a

00:04:56,700 --> 00:04:58,400


five-year interval.

00:04:59,300 --> 00:05:00,700


The upper panel here,

MScFE | © 2022 - WorldQuant University – All rights reserved. Video Transcript | PAGE 7
00:05:02,400 --> 00:05:05,800
shows that the frequency with which the algorithm picks

00:05:05,800 --> 00:05:08,800


the stock with the highest return is almost always

00:05:08,800 --> 00:05:12,400


below 50%, with large

00:05:11,400 --> 00:05:14,700


fluctuation arguably but with

00:05:14,700 --> 00:05:17,400


an average that is in the order of 20% as

00:05:17,400 --> 00:05:19,000


you can see, here. This is 0.2.

00:05:20,300 --> 00:05:23,100


This means that the algorithm in terms of picking the

00:05:23,100 --> 00:05:26,600


best performing stock, does as badly on

00:05:26,600 --> 00:05:30,000


average as a stock-picking criterion based

00:05:29,400 --> 00:05:32,700


that chooses stocks at random.

00:05:34,600 --> 00:05:37,400


However, the bottom panel depicts the

00:05:37,400 --> 00:05:39,200


cumulative return over time

00:05:40,200 --> 00:05:44,100


of the bandit-based strategy relative to the equally

00:05:43,100 --> 00:05:46,700


weighted portfolio. The bandit strategy

MScFE | © 2022 - WorldQuant University – All rights reserved. Video Transcript | PAGE 8
00:05:46,700 --> 00:05:49,500
provides a considerable increase in the returns of an investor,

00:05:49,500 --> 00:05:50,200


as you can see here,

00:05:51,400 --> 00:05:54,400


despite its low accuracy at finding the stock

00:05:54,400 --> 00:05:55,300


with the highest pay.

00:05:56,100 --> 00:05:59,400


If this is the case, it is because the algorithm may be unable to

00:05:59,400 --> 00:06:03,400


identify the stock with the highest payoff, but

00:06:03,400 --> 00:06:06,400


still selects stocks with a relatively

00:06:06,400 --> 00:06:08,500


high return above the average.

00:06:09,800 --> 00:06:12,700


Thus, the algorithm may not be right, but

00:06:12,700 --> 00:06:15,500


the wrong picks are still sufficiently good

00:06:15,500 --> 00:06:18,700


to offset the return of an equally weighted portfolio.

00:06:19,200 --> 00:06:22,100


Now, please feel free to play around with these parameters of

00:06:22,100 --> 00:06:25,500


the model to understand how sensitive

00:06:25,500 --> 00:06:28,500


this strategy can be to different setups and

MScFE | © 2022 - WorldQuant University – All rights reserved. Video Transcript | PAGE 9
00:06:28,500 --> 00:06:29,400
different choices.

00:06:30,900 --> 00:06:33,300


Now, this

00:06:33,300 --> 00:06:37,000


has been all for Module 6, hope you're enjoying

00:06:36,100 --> 00:06:39,400


it. I'll see you in the final module,

00:06:39,400 --> 00:06:42,300


Module 7, where we will describe the

00:06:42,300 --> 00:06:45,900


reinforcement learning techniques that combine the dynamic programming

00:06:45,900 --> 00:06:48,600


tools of Module 5, with the learning techniques

00:06:48,600 --> 00:06:50,600


that we have developed in Module 6.

00:06:51,200 --> 00:06:52,100


So, see you there!

MScFE | © 2022 - WorldQuant University – All rights reserved. Video Transcript | PAGE 10

You might also like