Moritz Benjamin
Moritz Benjamin
München 2018
Applications of Textual Analysis and
Machine Learning in Asset Pricing
Benjamin Moritz
Dissertation
an der Fakultät für Mathematik, Informatik und Statistik
der Ludwig–Maximilians–Universität
München
vorgelegt von
Benjamin Moritz
Diese Dissertation zeigt auf, wie neuere technische Entwicklungen helfen, fundamentale,
bisher nicht gelöste, Fragen in der Vermögenspreisbewertung zu beantworten. Das zen-
trale Ziel in der Vermögenspreisbewertung ist es, einen Preis für einen Vermögenswert zu
bestimmen, gegeben den Risiken: Der Rendite-Risiko Zusammenhang. Die Theorie der
Vermögenspreisbewertung wird unterteilt in den Gesamtmarkt im Zeitablauf und einzelne
Wertpapiere im Querschnitt.
Die Gesamtmarkttheorie befasst sich mit dem Marktportfolio, welches typischerweise mit
dem Aktienmarkt approximiert wird. Es gibt einen breiten Konsens über den Rendite-
Risiko Zusammenhang in der sehr langen Frist. In der mittleren Frist, beispielsweise zehn
Jahr, und der kurzen Frist, typischerweise eine Woche oder ein Monat, zeigen Studien, dass
der Rendite-Risiko Zusammenhang positiv, negativ, nicht-existent oder instabil über die
Zeit ist. Eine mögliche Quelle dieser Meinungsverschiedenheit ist das messen der kurzfristi-
gen Risikoerwartung. Diese Arbeit führt einen neuartigen Weg ein, wie erwartete Risiken
gemessen werden können. Methoden aus der Textanalyse, wie einfache bag-of-words Mod-
elle oder fortgeschrittenere Methoden aus der Sentimentanalyse, werden genutzt um die
Artikel der NY Times von 1851 bis heute zu analysieren. In den meisten Studien wird die
Realisation der Risiken, die einfache rollierende Standardabweichung der Aktienmarktren-
diten, stellvertretend für die Risikoerwartung genutzt. Das neu vorgestellte Maß kann als
ein besseres Proxy für die kurzfristige Risikoerwartung gesehen werden. Es kann gezeigt
werden, dass die Veränderung der Volatilität das Hauptrisiko in der kurzen Frist ist. Die
Änderung der Volatilität hängt signifikant negativ mit der Rendite zusammen.
Querschnittstheorien preisen Vermögenswerte relativ zueinander. Im Querschnitt wurden
mehr als 300 signifikante Variablen für den Rendite-Risiko Zusammenhang nachgewiesen.
Diese Zahl stieg über Jahrzehnte stetig an und brachte mittlerweile die traditionell genutzten
Analysemethoden, einfaches Sortieren und Querschnittsregressionen, vor große Probleme.
Diese Arbeit schlägt den Random Forest als neue Methode vor, um alle Variablen inklu-
sive deren Interaktionen und Nichtlinearitäten gemeinsam analysieren zu können. Der
Random Forest erweist sich dabei als überlegen gegenüber den traditionellen Methoden da
er mit den großen und komplexen Datenmengen deutlich besser umgehen kann. Analytis-
che Erweiterungen werden eingeführt, um bessere praktische Einblicke in die Methode des
Random Forest zu gelangen und um einen ausführlichen Vergleich mit den traditionellen
Methoden zu ermöglichen.
vi
Summary
This thesis shows how new technical developments can help answering fundamental, but yet
unresolved, questions in financial asset pricing. The central goal of financial asset pricing
theory is to determine the price of an asset given its risks: The risk-return relationship.
The pricing theories are classified into aggregate over time and single securities in the
cross-section.
Aggregate pricing theories are concerned with the market portfolio, which is often proxied
by the stock market. There is a broad consensus over the very long-term risk-return
relationship. For the medium-term, e.g. ten years, and short-term horizon, which is
typically a week or a month, studies show that the risk-return relationship is positive,
negative, non-existent or unstable over time. One source of this disagreement may be
the measurement of the short-term expectation of risk. This thesis introduces a novel
way to measure the expectation of risk. Textual analysis methods, such as simple bag-of-
words models or more advanced sentiment extraction methods, are used to analyze articles
from the NY Times from 1851 up to today. In most studies the realization of risk, the
simple rolling standard deviation of the stock market returns, is used as the proxy for
the expectation of risk. The new proposed measure can be seen as a better proxy for the
short-term expectation of risk and it can be shown that the change in volatility is the key
risk in the short-term. The change in volatility has a significant negative relation to the
return.
Cross-sectional theories price assets relative to each other. In the cross-section there are
more than 300 variables which account for cross-sectional risk-return relationships. The
number of variables has been growing steadily for decades and today problems are arising on
how to handle so many variables with the traditional methods: Sorting and cross-sectional
regressions. This thesis proposes to use the random forest as a new method to analyze
all those variables with their interactions and non-linearity jointly. The random forest
is superior to the traditional methods used because it can handle complex and big data
easily and it gives useful insights on which variables are important. Analytical extensions
are introduced to get a better insight into the random forest and to provide a thorough
comparison to the traditional methods used in cross-sectional asset pricing.
viii
Contents
1 Introduction 1
1.1 The current state of asset pricing . . . . . . . . . . . . . . . . . . . . . . . 1
1.2 The contribution and outline of this thesis . . . . . . . . . . . . . . . . . . 2
1.3 Contributing manuscripts . . . . . . . . . . . . . . . . . . . . . . . . . . . 3
A First Appendix 75
B Second Appendix 77
B.1 Estimating expected returns with standard methods . . . . . . . . . . . . . 77
B.1.1 Portfolio sort . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 77
B.1.2 Fama-MacBeth regressions . . . . . . . . . . . . . . . . . . . . . . . 77
B.2 Transaction costs . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 81
B.2.1 Fama-MacBeth with recent returns only . . . . . . . . . . . . . . . 85
B.3 Robustness of the discovered structure . . . . . . . . . . . . . . . . . . . . 87
B.3.1 Including firm characteristics . . . . . . . . . . . . . . . . . . . . . 87
B.3.2 Expanded set of return functions . . . . . . . . . . . . . . . . . . . 91
B.3.3 Estimation by size categories . . . . . . . . . . . . . . . . . . . . . . 91
B.4 Illustration of a conditional portfolio sort . . . . . . . . . . . . . . . . . . . 94
B.5 Greedy algorithm . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 98
Chapter 1
Introduction
best in economic research. Cerniglia et al. (2016) explain that quantitative computational
methods are critical for success in stock market investing. They give a short overview over
new computational methods and argue that it is important to understand the intuition,
assumptions, and strengths and weaknesses of those new computational approaches.
Financial asset pricing distinguishes the pricing of major asset classes over time and the
pricing of individual securities within each asset class in the cross-section. The literature
on both questions is rising steadily and in some parts no consensus has been reached.
Cochrane (2011) wrote an agenda where he summarized the conflicting results in asset
pricing and which led to many discussions until today.
The current state in time-series asset pricing is as follows: In the very long-run it is
common sense that the risk-return relation of the major asset classes is linear. In the
medium-run, e.g. ten years, and short-run, e.g. one week or one month, the relation is
ambiguous.
The current state in cross-sectional asset pricing can be described as: There are more
than 300 factors documented which describe the cross-sectional differences in expected
stock returns. These factors have been found by comparing them to existing factors relying
on statistical models from the 1970s, which have been found to be unable to handle, and
correct for, so many factors simultaneously.
Chapter 2 shows how methods from textual analysis may give more insights into the relation
of the return and risk of the stock market. This thesis contributes to the time-series asset
pricing literature twofold. First, a new measure of real-time assessment of the business
cycle based on textual analysis is introduced. Second, this measure points to the fact, that
for a short-term investor the major risk is the change in volatility and not the level in
volatility.
Chapter 3 presents a new approach in pricing individual stocks with the random forest
algorithm from the field of machine learning. The contribution to the cross-sectional as-
set pricing literature is twofold as well. A new method, the random forest algorithm, is
introduced to present a possible solution for handling the hundreds of variables simultane-
ously. With an appplication to return-based variables, it can empirically be shown, that
the short-term returns become the most important predictors for future returns.
Chapter 2 is joint work with Stefan Mittnik. The research idea and the approach were
developed jointly, as was the data analysis. The idea of using textual analysis is from
Benjamin Moritz, who also collected the data and prepared the manuscript.
Chapter 3 and Moritz and Zimmermann (2016) are identical and is joint work with Tom
Zimmermann. The development of the research idea, the analysis of the data and the
preparation of the manuscript was done in collaboration.
Moritz, B. and T. Zimmermann (2016). Tree-Based Conditional Portfolio Sorts: The Rela-
tion between Past and Future Stock Returns. [Link]
cfm?abstract_id=2740751
4 1. Introduction
Chapter 2
2.1 Introduction
In the very long-run the risk-return relation of the major asset classes is linear. As figure
2.1 shows, when volatility is the measure of risk, Ibbotson et al. (2016) show a very linear
relation between the realized volatility and returns of large-cap stocks, small-cap stocks,
short-term government bonds, long-term government bonds and long-term corporate bonds
measured over almost 90 years.
In this study we focus on the time-varying or conditional case. Because the market
portfolio is commonly proxied by the stock market1 in asset pricing studies, we focus on
the stock market solely to investigate the risk-return relationship.
In the medium-term, approximately 2 to 50 years, there can be substantial variation
in the returns and volatility of the asset classes as has been documented by for example
Cochrane (2011), Ibbotson et al. (2016), Ilmanen (2011) or recently Jordà et al. (2017).
The reasons for the medium-term variations are manifold, including structural changes
and the business cycle. Because returns exhibit a positive return expectation in the very
long-run, deviations from the long-term path always led to mean-reversion in returns of
the major asset classes in the U.S.2 Most studies investigating the time-varying stock
market risk premia consider the medium-term while most studies regarding the risk-return
relationship focus on a very short horizon.
Over a very short-run horizon, e.g. one day, one week or one month, the relation between
stock market risk and return is ambiguous. For example Campbell (1987), Glosten et al.
(1993) and Whitelaw (1994) find a negative relation, Ghysels et al. (2014) find a nonlinear,
while Bollerslev et al. (1988), Ghysels et al. (2005) and Lundblad (2007) find a positive
relation.
The reasons are diverse, ranging from volatility as a bad measure of risk to insufficient
1
Back (2017)
2
Albeit this fact is not true for every country, as shown by Dimson et al. (2009).
6 2. The Relation between Stock Market Risk and Return
Figure 2.1: Historical geometric mean and standard deviation of monthly returns for the major U.S.
asset classes from 1926 to 2015. Source is Ibbotson et al. (2016).
We rely on the findings of French et al. (1987), who find evidence that on a monthly
horizon, the expected stock market return has a positive albeit not strong relation with the
2.1 Introduction 7
predictable part of the volatility but a strongly negative relation with the unexpected part
of the volatility. French et al. (1987) state, that this leads to the conclusion, that a negative
ex post relation between excess holding period returns and the unexpected component of
volatility is consistent with a positive ex ante relation between risk premiums and volatility.
This notion leads us to further distinguish between the horizon of realized and expected
returns: If volatility rises and prices fall contemporaneously than medium-term expected
returns are higher and medium-term expected volatiliy is lower because of the medium-term
mean-reversion in prices and volatility.
Two very important points in the theoretical continuous time model of Merton (1973)
are, that the choice of horizon is open to the researcher and that the market conditional
variance is unobserved. The expected volatility can be decomposed into todays volatility
and expected change in volatility: Et [V art+1 ] = Et [V art+1 − V art ] + V art . Most studies
use the realized volatility V art as a proxy for the expected volatility Et [V art+1 ] in equation
2.1. But what matters for the risk-return relation is the change in volatility from todays
volatility: E[V art+1 − V art ]. French et al. (1987) show statistically that V art+1 − V art
has a strong negative relation with rt+1 . We support this view by introducing a daily
measure of the business cycle expectations. Our measure shows that most of the time the
expectation of the business cycle turns better when volatility is at the highest and prices
are at the lowest point in the cycle. This leads us to the conclusion that the change in
volatility is what matters for short-term investors and not the level of volatility alone.
This strong negative relation between realized returns and the change in volatility is not
only a matter of fact on a monthly horizon. It holds for the full business cycle frequency.
It is well established, that the level of prices is procyclical and the level of volatility is
countercyclical as shown by Schwert (1989a), Schwert (1989b) and Panetta et al. (2006)
for example. Therefore while the autocorrelation in monthly stock market returns is not
significantly different from zero, there is yet a business cycle frequency in the stock returns.
This fact is not surprising, because the efficient market hypothesis states that a market is
called ”efficient” if prices always ”fully reflect” available information. Therefore every little
piece of information influences stock prices and therefore returns and volatility. There is
a lot of noise in the stock market prices. Only a small part of daily price changes can
be explained by economic reasons or news, as shown by Niederhoffer (1971), Roll (1988),
Cutler et al. (1989), Fair (2002), Cornell (2013) and Campbell et al. (2018). But nearly
all major price changes which occur over months can be attributed to the business cycle,
as we confirm in this study.
To get an index which summarizes all the information of economic announcements,
political outcomes, geopolitical news, etc instantaneously and well, we use the articles of
the New York Times from 1851 to 2018. We construct a daily measure of the business cycle
expectations by counting (from 1851 to 2018) and analyzing (from 1981 to 2018) the articles
which mention the word ”recession” in the text. Counting the word ”recession” in articles
and using as a leading indicator for the business cycle has been used first, as far as we
know, by The Economist in 19923 . The Economist has named that index the R-word index,
3
No news is good news. - The Economist, July 11, 1992; pg. 39; Issue 7767.
8 2. The Relation between Stock Market Risk and Return
has published updates of that index regularly and wrote that it serves useful on dating the
begin and the end of an economic recession. The index has a quarterly frequency. We built
on that idea but count the word on a daily frequency. Using the word ”recession” only has
the advantage, that it occurs solely in economics. A search for other words, like expansion,
have meanings beyond the business cycle. However, methods from textual analysis offer
us much more possibilities in extracting more information out of the articles mentioning
”recession”. It could be the case that we are at the turning point in the midst of a recession,
but the newspapers are still writing on the reasons of a recession. Or the articles can be
more optimistic about the future while writing about the recent recession. With textual
analysis methods we can extract the tone, or sentiment, from an article. Sentiment is here
defined as the positive or negative tendency of an article. That should not be confused
with the definition of Baker and Wurgler (2007). They define investor sentiment as a belief
about future cash flows and investment risks that is not justified by the facts at hand. This
is not the case here. We want to measure exactly the belief about the facts about future
cash flows and investment risks. Textual analysis has its roots decades ago but gained a
lot of importance in finance and economics in the last years. Surveys on textual analysis
in finance are for example Das et al. (2014), Kearney and Liu (2014), Tetlock (2014) and
Loughran and McDonald (2016). There are some applications on the stock market, but
most studies focus on a daily prediction of stock market activity. For example, Antweiler
and Frank (2004) show that internet stock message boards help predict volatilities on a
horizon of 15 minutes. Tetlock (2007) shows that the number of negative words in the
“Abreast of the Market” column of the Wall Street Journal predicts stock returns at the
daily frequency. Garcia (2013) shows that the fraction of positive and negative words
in two columns of financial news from the New York Times predicts stock returns on a
daily horizon but in recessions only. Da et al. (2014) show that an index constructed out
of user’s google internet searches can forecast asset prices at a daily horizon. There are
many other applications as well. Doms and Morin (2004) build a recession word-index
and analyze how the media affects consumers sentiment regarding the economy. Lawrence
et al. (2017) determine the relevance of academic papers in economics and finance with
textual analysis. Baker et al. (2016) create indexes of policy-related economic uncertainty
based on newspaper coverage frequency. Manela and Moreira (2017) construct a text-based
measure of uncertainty. Beckers et al. (2017) extract the sentiment from newspaper and
find that this indicator is competitive in forecasting inflation against a large set of common
predictors. Their sentiment indicator is superior to simple word-count indicators.
NASDAQ that have a CRSP share code of 10 or 11 at the beginning of month t, good
shares and price data at the beginning of t, and good return data for t.
For the news data we take the data from the New York Times as in Garcia (2013),
who writes that the two main media sources with regular coverage of business news were
the Wall Street Journal and the New York Times. The New York Times provides digital
access to the whole archive from 1851 until today. It is possible to do a full-text-search on
all articles from 1851 until today. The content for every article can be reached digitally for
all articles from 1981 until today.
We analyze the data with methods from the textual analysis literature. We proceed as
follows:
2. At first we scan all articles of the New York Times from 1851 until today for the word
”recession”. It is a simple word counting approach which is commonly used in the
finance literature. For each article containing the word ”recession” we get back the
timestamp, the headline, some metadata and the URL to reach the full document.
3. In the next step we loop over each URL to scrap the full html-Website und afterwards
we parse the article to extract the full body text.
4. Once we have the full text, we parse the text into a vector of sentences. The articles
before 1981 are all digitalized but not accessible through a html format but only
pdf through an application. So we focus on the full text starting in 1981 to get the
sentiment or tone of each article.
5. The sentiment is calculated using several dictionaries, which assign positive and neg-
ative values to single words or sentences to express a positive or negative tone. There
are several dictionaries available which can be used5 .
6. To calculate the total tone/sentiment of each article, we take the average over all
sentences in each article. Taking the average gives us the central tendency or mean
emotional valence.
7. Finally we average over all articles in one day to get the total tone/sentiment for
each day.
For analyzing the risk-return relation we follow French et al. (1987) and regress the
monthly stock market returns against the predictable and unpredictable components of
the standard deviations or variances of stock market returns with weighted least squares.
p
rmt − rf t = α + β σ̂mt + t (2.2)
5
See: [Link]
10 2. The Relation between Stock Market Risk and Return
p pu
rmt − rf t = α + β σ̂mt + γσmt + t (2.3)
where:
rmt = return of the stock market
rf t = return of the risk free rate
p
σ̂mt = expected risk of the stock market
pu p p
σmt = unexpected risk of the stock market = σmt − σ̂mt
with p = 1 for standard deviation and p = 2 for volatility
2.3 Empirics
2.3.1 Regressions of returns on volatility
We follow Moreira and Muir (2017) and estimate the expected risk of the stock market
with the standard deviation of realized returns using a 22-day rolling window. In table 2.1
we present the same form of regression as in French et al. (1987). The first sample is the
same sample as in French et al. (1987) and the second sample is the full data we use in our
study. The results are equal to French et al. (1987). While the relation between returns
and expected volatility is weakly positive and in some cases even negative, the relation
between returns and unexpected volatility is strongly negative.
Figure 2.2: The figure shows the number of occurrences of the word ”recession” in the New York Times
Articles (20 day moving average). Shadings show the U.S. recessions according to NBER recession dates.
2.3 Empirics 13
Figure 2.3: The figure shows the number of occurrences of the word ”recession” in the New York
Times Articles (20 day moving average). As well the cumulated return of the U.S. stock market index
(logarithmic) and the stock market standard deviation (22 day rolling window). Shadings show the U.S.
recessions according to NBER recession dates.
14 2. The Relation between Stock Market Risk and Return
Prices are rising and volatility is falling usually earlier, at the bottom of the recession not
at the end of a recession. Therefore it is not a good indicator for a real-time assessment of
the economy. It seems that after the recession has bottomed out, the media is continuing
writing on the recession. Maybe discussing its origins, the turning point or potential
repercussions of the recession. That’s why we need a measure on the tone of these articles
to get a sense of the sentiment of the articles.
Figure 2.4: The figure shows the mean sentiment over all articles of one day. The red line is the 100 day
moving average.
In Figure 2.6 we compare the number of articles with the trend of sentiment, the stan-
dard deviation of the sentiment, the cumulated market return and the market standard
deviation from 1981 to 2018. We can observe some major points which we will investigate
further. At first, we can see that in the recessions of 1982, 1991, 2002 and 2008 sentiment
has its lowest point around the stock market low. Other US leading indicators typically
bottom out a bit later and return to normal much later. For example the ISM Manufac-
turing Index hit its lowest point in December 2008 which has been announced on Januar
2.3 Empirics 15
Figure 2.5: The figure shows the mean sentiment over all articles of one day (100 day moving average)
for the four dictionaries. Black = Syuzhet, Red = Bing, Blue = Afinn, Green = NRC
2, 2009.7 But the volatility had its highest point in that cycle in October 2008. Secondly,
sentiment trended slowly down ahead of 1987. At third, in the recessionary environment of
1998/1999 and 2011 sentiment went down and volatility shot up while prices do not trended
down. So there might be an external risk (Asian crises and European crises respectively)
which in the end did not materialize in the US and therefore the prices fluctuated wildly
because of uncertainty but they did not fall. Fourth, the volatility of the sentiment goes
down in every recession and it bottoms out at the end of each recession. There seems to
be a high consensus at the end. But it rises steadily after the end of each recession and
has a high level just before the next recession. So uncertainty over the future economic
outlook is at the highest right before a recession sets in.
Two exemplary articles may clarify the point, that extracting the sentiment from an
article gives much more insights. They are in the midst of the great recession in 2008 when
the stock market standard deviation was at its highest point in that business cycle and
the stock prices hit their temporary lowest point. Measured with the sentiment extraction
method the first article has a sentiment of -0.112 and the second article has a sentiment
score of 0.202. The following excerpts give a short view into the articles:
NY Times Article from October 11, 20088
Title: Those With a Sense of History May Find It’s Time to Invest
Text: ... He says that investors with a stomach for risk and a long time horizon
should consider following Warren E. Buffett, who in the last three weeks has
invested $8 billion in Goldman Sachs and General Electric. ...
7
More in a note on ”The Performance of the ISM Manufacturing Indexes in the 2008-2009 Finan-
cial Crisis” at [Link]
[Link]
8
See: [Link]
16 2. The Relation between Stock Market Risk and Return
Figure 2.6: The figure shows the number of occurrences of the word ”recession” in the New York Times
Articles (20 day moving average). The mean sentiment over all articles of one day (100 day moving average)
and the standard deviation of the mean sentiment (100 day rolling window). As well the cumulated return
of the U.S. stock market index (logarithmic) and the stock market standard deviation (22 day rolling
window). Shadings show the U.S. recessions according to NBER recession dates.
Figure 2.7: The figure shows the number of occurrences of the word ”recession” in the New York Times
Articles (20 day moving average). The mean sentiment over all articles of one day (100 day moving average)
and the standard deviation of the mean sentiment (100 day rolling window). As well the cumulated return
of the U.S. stock market index (logarithmic) and the stock market standard deviation (22 day rolling
window). Shadings show the U.S. recessions according to NBER recession dates.
price bottomed out at the same time but fell again in March 2009 probably because of U.S.
economic policy uncertainties. Interestingly, the stock market hit a low on March 6, 2009.
Just right after former U.S. President Barack Obama on March 3, 2009 said ”... ,what
you’re now seeing is profit and earning ratios are starting to get to the point where buying
stocks is a potentially good deal if you’ve got a long-term perspective on it.”.10 Further,
10
See: [Link]
2.3 Empirics 19
Figure 2.8: The figure shows the number of occurrences of the word ”recession” in the New York Times
Articles (20 day moving average). The mean sentiment over all articles of one day (100 day moving average)
and the standard deviation of the mean sentiment (100 day rolling window). As well the cumulated return
of the U.S. stock market index (logarithmic) and the stock market standard deviation (22 day rolling
window). Shadings show the U.S. recessions according to NBER recession dates.
the sentiment reached its pre-crisis level in January 2009, indicating no more recessionary
risks very early.
Following the four recessions we investigate four further periods which have resulted
partially in a large literature. The first is the October 1987 stock market crash. The stock
market rose until October and fell dramatically within a couple of days which can be seen in
figure 2.11. Volatility rose simultaneously. The number of articles mentioning ”recession”
20 2. The Relation between Stock Market Risk and Return
Figure 2.9: The figure shows the number of occurrences of the word ”recession” in the New York Times
Articles (20 day moving average). The mean sentiment over all articles of one day (100 day moving average)
and the standard deviation of the mean sentiment (100 day rolling window). As well the cumulated return
of the U.S. stock market index (logarithmic) and the stock market standard deviation (22 day rolling
window). Shadings show the U.S. recessions according to NBER recession dates.
was normally low before the crash. It shot up right after the crash but came down very
quickly. Interestingly the sentiment trended down the years before 1987 and reached a very
low level at the beginning in 1987 but rising again a little in the three months before the
crash. The volatility of the sentiment rose before the stock market crash, which is usually
typical preceding a recession. So the sentiment extracted some information which gave the
sign that a recession may be imminent. So according to the sentiment information it was
2.3 Empirics 21
Figure 2.10: The figure shows the number of occurrences of the word ”recession” in the New York Times
Articles (20 day moving average). The mean sentiment over all articles of one day (100 day moving average)
and the standard deviation of the mean sentiment (100 day rolling window). As well the cumulated return
of the U.S. stock market index (logarithmic) and the stock market standard deviation (22 day rolling
window). Shadings show the U.S. recessions according to NBER recession dates.
Figure 2.11: The figure shows the number of occurrences of the word ”recession” in the New York Times
Articles (20 day moving average). The mean sentiment over all articles of one day (100 day moving average)
and the standard deviation of the mean sentiment (100 day rolling window). As well the cumulated return
of the U.S. stock market index (logarithmic) and the stock market standard deviation (22 day rolling
window). Shadings show the U.S. recessions according to NBER recession dates.
near-collapse of the Hedge Fund Long Term Capital Management (LTCM) and there was
a lot of discussion if it led to a recession in the U.S. which can be shown by the number of
articles mentioning ”recession”. It shot up in 1998 to around 6 per day. Interestingly the
sentiment went down too and much earlier. The trend of sentiment reached the bottom
when volatility was at its highest point and stocks at its lowest level. So it seems to be an
excellent measure of the real-time expectations of the state of the economy.
2.3 Empirics 23
Figure 2.12: The figure shows the number of occurrences of the word ”recession” in the New York Times
Articles (20 day moving average). The mean sentiment over all articles of one day (100 day moving average)
and the standard deviation of the mean sentiment (100 day rolling window). As well the cumulated return
of the U.S. stock market index (logarithmic) and the stock market standard deviation (22 day rolling
window). Shadings show the U.S. recessions according to NBER recession dates.
The third event was the drop in stock markets in August 2011. It was because of some
fears of contagion from the european sovereign debt crisis and because the U.S. lost its
AAA rating from Standard & Poor’s for the first time in history. Again, as one can see
in figure 2.13, the sentiment went down from April 2011 and reached its low in September
2011 just when volatility was highest and the stock market prices were at the lowest point.
Afterwards volatility fell and prices and sentiment rose simultaneously. Indicating that
24 2. The Relation between Stock Market Risk and Return
what matters for risk is not the level, but the change in volatility.
Figure 2.13: The figure shows the number of occurrences of the word ”recession” in the New York Times
Articles (20 day moving average). The mean sentiment over all articles of one day (100 day moving average)
and the standard deviation of the mean sentiment (100 day rolling window). As well the cumulated return
of the U.S. stock market index (logarithmic) and the stock market standard deviation (22 day rolling
window). Shadings show the U.S. recessions according to NBER recession dates.
The fourth mentionable event was the discussion preceding the U.S. elections of 2016
about the possible economic repercussions for every candidate. It was heavily debated if
the economy may go into a recession after elections, when Donald Trump may be elected.
As can be seen in figure 2.14, the number of articles writing over ”recession” was very high
in the second half of 2016, but sentiment, the stock market prices and volatility were all
2.4 Implications for investors 25
nearly unaffected.
Figure 2.14: The figure shows the number of occurrences of the word ”recession” in the New York Times
Articles (20 day moving average). The mean sentiment over all articles of one day (100 day moving average)
and the standard deviation of the mean sentiment (100 day rolling window). As well the cumulated return
of the U.S. stock market index (logarithmic) and the stock market standard deviation (22 day rolling
window). Shadings show the U.S. recessions according to NBER recession dates.
this study is the fact, that the variance treats the upside and downside the same, but
investors generally exhibit asymmetric risk aversion, feeling losses more powerfully than
any equivalent gain. This asymmetry has been originally formulated by Kahneman and
Tversky (1979) as the prospect theory, also called loss aversion preferences. For this reason
there is a high demand for investment strategies which try to avoid losses when the stock
market falls in the short-term. We concentrate on the fact, that negative returns are only
weak positively related to the level of stock market volatility, but strong negatively related
to the change of stock market volatility and show that it is this change in volatility which
is the true risk for a short-term investor. Moreira and Muir (2017) implement a strategy
where they allocate into the stock market inverse to the level of the stock market volatility.
They find a Sharpe ratio for the strategy which is higher than that of the stock market
and call that a puzzle, because it takes relatively less risk in recessions when the level of
volatility is higher than normal. But when the level of volatility is not the right measure
of risk, but the change of volatility, then it may not be a puzzle, but an explanation for a
strategy which has its advantages and disadvantages as every strategy has.
So what can be the explanation here. If the strategy invests according to the level
of volatility then it invests less in the stock market typically in an economic recessions.
That is, it avoids the fall but also the subsequent rise in stock market prices. That is a
reasonable strategy but not a puzzle. Because in this case the strategy is to avoid stocks
when risk is high (volatility is rising) and also when risk is low (volatility is falling). So in
the phase of falling volatility, the investor of that strategy needs to be patient and out of
the stock market, when many others are investing and the sentiment / tone in the media is
rising or even positive. So it is a strategy which has its advantages and disadvantages. For
example, in the strategy of Moreira and Muir (2017), their strategy does not work in every
country. They also use leverage and a lot of the high return for the U.S. comes from the
fact, that they leverage in the stock market in times of low volatility. A leverage means,
that they invest more then 100%. Also a strategy does not always pay off. There can be
times when such a strategy does not perform. That could be ten years in a row.
In figure 2.15 we show again the great recession of 2008/2009. We added two lines.
The dashed red line shows the highest point in volatility. The dashed blue line shows
the average of the standard deviation over the shown timespan. The average standard
deviation is 22.86% and the highest point of volatility is on October 28, 2008. We compute
the annualized average of the return and of the standard deviation before and after the red
line and above and below the blue line. The mean standard deviation of the days above the
blue line is 38.77%, while it is 15.62% below. The mean return is 3.68% and 4.03%. The
mean standard deviation of the days before the red line is 19.67%, while it is 25.55% after.
The mean return is -17.22% and 21.70%. This result shows, that when losses are what
investors care about, than high volatility is not equal to risk in the short-term. Because it
has approximately the same average return as the low volatility regime. This must be the
case, when returns and volatility are strong negatively related contemporaneously. A rise
and a fall in the volatility is what investors are really worried about.
The key conclusion for the expected risk and return relation seems to be as follows: One
must distinguish between the short-term (momentum) and medium-term (mean-reversion)
2.4 Implications for investors 27
Figure 2.15: The figure shows the number of occurences of the word ”recession” in the New York
Times Articles (20 day moving average). The mean sentiment over all articles of one day (100 day moving
average) and the standard deviation of the mean sentiment (100 day rolling window). As well the cumulated
return of ths U.S. stock market index (logarithmic) and the stock market standard deviation (22 day rolling
window). Shadings show the U.S. recession according to NBER recession dates. The dashed red line shows
the highest point in standard deviation. The dashed blue line shows the average of standard deviation
over the shown timeframe.
expected return. The stock market price falls (rises) and volatility rises (falls) contempo-
raneously. Not perfectly on a monthly basis, but very strong over the business cycle. So
in the short-term the expected return rises (falls) when prices rise (fall) and volatility falls
(rises). Because the price and the volatility always mean revert, on the medium term the
28 2. The Relation between Stock Market Risk and Return
expected return is higher when price is low and volatility is high and the expected return
is lower when price is high and volatility is low. Because the exact turning points of the
mean reversion are quite difficult to estimate, the strategy of leaving the stock market
when volatility is high and investing in the stock market with leverage when volatility is
low can lead to payoffs which certain investors demand for. This can be improved if one
has forecasting skills for the future change in volatility or returns directly, beyond simple
volatility estimates like a rolling window of the standard deviation.
2.5 Conclusion
The relation between stock market risk and return is ambiguous. One reason is that many
studies focus on the stock market volatility level and relate it to future returns. But in the
original work of Merton (1973), the stock market volatility level is also the future volatility.
It is not observable and impossible to estimate exactly. We focus on the findings of French
et al. (1987) which show that realized returns and volatility are strong negatively related
contemporaneously. We show with methods from the textual analysis literature, that the
change in volatility is what short-term investors care about instead of the level of volatility.
That means: If volatility is at its cycle high, it is the best point to buy stocks. But because
timing is difficult, a strategy which goes out of the market when volatility is rising and goes
in after volatility has fallen to normal levels may be useful in some circumstances, because
one avoids the down- and upturn. But it is a strategy with its typical advantages and
disadvantages. So adjusting the portfolio inversely to the level of volatility makes sense,
but it is not a puzzle.
Chapter 3
3.1 Introduction
Consider the challenge of a portfolio manager who wants to use past information to es-
timate expected returns at the firm level. He has at his disposal an overwhelming set of
potentially correlated predictor variables, as documented by several recent survey papers.
Subrahmanyam (2010) surveys 50 earnings-based return predictive signals, McLean and
Pontiff (2016) document 82, and Harvey et al. (2016) and Green et al. (2013) both extend
the list to around 330. These variables range from classic accounting-based variables, such
as book-to-market, to return-based variables, such as the stock return over the previous
year. They may even include more exotic variables such as the creativity of a stock’s ticker.
Many of these diverse variables might interact in nontrivial ways, increasing the set even
more. Furthermore, the literature suggests stand-ins for many variables (for example, value
or quality). Which should the manager pick? Beyond these many challenges lurks the risk
of overfitting the data with any estimation method that the manager might use, rendering
the analysis worthless for new observations. How then should one go about estimating
expected returns while taking all these issues into account?
The literature in empirical asset pricing provides a few methods. We show, however,
that two of them — portfolios sorts and Fama-MacBeth regressions — can deal only with
some of the challenges outlined above. We suggest an alternative approach, motivated by
the method of conditional portfolio sorts but extending easily to large sets of predictor
variables and flexibly dealing with their interactions. In contrast to how conditional port-
folio sorts are usually applied, we use the data to estimate both the optimal conditioning
variables and associated optimal thresholds.
Our contribution to the literature is threefold: First, we import ideas from the ma-
chine learning literature and tailor them to a financial application in order to produce
a model that is suited to evaluate the independent information in the entirety of many
cross-sectional predictor variables and their potential interactions. While these methods
30 3. Tree-Based Conditional Portfolio Sorts
are data-driven, we are careful to develop valid out-of-sample validations of the model. As
the machine learning literature is often criticized for producing black box predictions, we
especially emphasize new measures to extract interpretable information about the structure
of the estimated prediction function.
Second, we apply our methodology to the prediction of future returns from past stock
returns, and we recover short-term returns (that is, the past six most recent one-month
returns) as the most important predictors. Implementable trading strategies based on
our findings have a risk-adjusted monthly return of around 2 percent per month, with an
information ratio that is about three times as high as the information ratio achievable in
a linear framework that does not account for nonlinearities and variable interactions. The
information ratio is about twice as high as in a Fama-MacBeth framework that accounts
for two-way interactions. Transaction costs cannot account for our results.
Third, we trace the improved predictions to several sources. We find that recent past
returns, not more distant past returns, contain almost all information about future returns
when information is exploited more efficiently than in a linear model. This finding addresses
the tension between the superior return of intermediate momentum in Novy-Marx (2012)
and Goyal (2011), who cannot find this effect in other countries. Reassuringly, our model
also reproduces previous facts in U.S. equity returns data. Our model finds short-term
reversal, the seasonal return effect of Heston and Sadka (2008), and momentum for longer-
dated returns. But the model also suggests a few new facts: A nonlinear relationship
between relatively recent past returns and future returns and important interactions among
past returns at different horizons.
These results pose a challenge for existing methodologies when the goal is to evaluate
many variables in a joint framework. The portfolio sort methodology, a dominant method
in analyzing cross-sectional predictor variables,1 sorts stocks into 3 to 10 portfolios each
month (or year) based on the value of a particular variable. In the next step, subsequent
returns for each portfolio are calculated and it is checked whether there is a monotone
relation between the sorting variable and these subsequent portfolio returns. In addition,
researchers often compute the equal- or value-weighted hedge return of going long (short)
in the highest quantile portfolio and going short (long) in the lowest quantile portfolio. The
relevance of the sorting variable is then assessed by comparing the hedge return to some
equilibrium model of asset prices (for example, the capital asset pricing model) and/or
by assessing the monotonicity of the returns over deciles. The portfolio sort methodology
is a powerful, nonparametric tool that works best in low dimensional cases. Problems
arise if returns are to be sorted on more than two or three predictor variables, as there
will typically be few stocks in each portfolio. But this makes controlling for information
contained in other variables challenging, or, as Fama and French (2008) put it, ”sorts are
awkward for drawing inference about which anomaly variables have unique information
about average returns.”
Multivariate Fama-MacBeth regressions (Fama and MacBeth (1973)) are able to ad-
dress this concern by showing the marginal effect of each predictor variable once all others
1
See the survey of Green et al. (2013).
3.1 Introduction 31
3
As Harvey et al. (2016) note ”it is possible that a particular factor is very important in certain economic
environments and not important in other environments. The unconditional test might conclude the factor
is marginal.”
3.1 Introduction 33
our results are not entirely driven by illiquid stocks by re-performing all computations for
large, small, and micro firms (in the terminology of Fama and French (2008)) separately.
While we find that results are stronger in small stocks and strongest in micro stocks, our
main conclusions hold throughout all size categories. We conclude that more recent past
returns are more relevant than intermediate past returns for prediction of future returns,
and more generally that past returns are related to future returns in a more complex way
than can be captured by any single past return.
Before we continue, we provide a short overview of the related literature. In his presi-
dential address, Cochrane (2011) describes the ”factor zoo” of stock market anomalies and
how it has developed over the years. Subrahmanyam (2010), Goyal (2011), Green et al.
(2013) and Harvey et al. (2016) review as many as 330 anomalies that have been found by
academic research and call for a synthesis of the existing literature. While early attempts
in this direction focused on smaller sets of characteristics were undertaken by Haugen and
Baker (1996), Daniel and Titman (1997) and Brennan et al. (1998), Cochrane (2011) argues
that different methods might be required to find the independent information for average
returns in the entirety of documented predictor variables. Our paper can be read as an
attempt to provide just such a new method.
Green et al. (2014) investigate the mutual information in 100 signals and find that
up to 24 of them have predictive power for returns when used jointly. They suggest an
alternative to the standard three-factor model by Fama and French (1992) that is based
on 10 different characteristics. The paper notes the potential relevance of interactions but
does not investigate them in detail.4 Lewellen (2015) investigates the power of 15 different
firm characteristics to predict variation in the cross-section. He finds that expected stock
returns derived from the model are strongly predictive of actual stock returns for as long
as 12 months.
Fama and French (2015) follow an alternative approach that attempts to capture vari-
ation in returns by a (small) factor model. They extend the three-factor model by proxies
for profitability and investment, which appears to capture contemporaneous variation in
cross-sectional returns well, except for small stocks. The paper uses a quadruple sorting
strategy to address interactions between size, value, profitability, and investment opportu-
nities. Kogan and Tian (2015) construct all combinations of three-and four-factor models
from a set of 27 firm characteristics. They find that the best performing models are un-
stable across time periods.
The literature on momentum and reversal is too large to review comprehensively here,
but we note a few key articles. If stock prices systematically over- or underreact, future
stock returns should be predictable from past returns data alone. de Bondt and Thaler
(1985) test overreaction by sorting stocks based on the return in the previous three years
(the portfolio formation period). They find that losers (the bottom decile of returns in the
formation period) outperforms winners by about 25 percent over three years. Jegadeesh
4
They write, ”fundamental valuation type measures and market trading type measures appear to matter
across firm size. In large-cap firms the important RPS can be broadly classified as fundamental valuation
measures or trading type measures. For mid-cap and small-cap firms the themes appear slightly different.”
34 3. Tree-Based Conditional Portfolio Sorts
(1990) and Lehmann (1990) find a ”reversal” effect for portfolios that are formed based on
short-term (one week to one month) prior returns. Jegadeesh and Titman (1993), on the
other hand, find evidence for a ”momentum effect” when portfolios are sorted on medium-
term (3 to 12 months) prior return. The momentum finding survives the analysis in Fama
and French (1996), who use the three-factor model as a model of equilibrium returns.
Long-term reversal disappears as an anomaly once normal returns are approximated by
the three-factor model. For much more on momentum, we refer to Asness et al. (2014),
who use simple analysis and survey published studies to show that momentum returns are
(among other things) not too volatile, are not only a small firm phenomenon, and are not
dwarfed by tax considerations or transaction costs.
A related literature in machine learning that tries to predict stock returns has developed
largely unnoticed by the finance literature. The machine learning literature has focused
on predicting stock returns from a few return-based and accounting-based variables jointly
but has then largely ignored the structure of the prediction equation, instead analyzing the
quality of the prediction itself.56 This article can also be viewed as an attempt to connect
the two and to provide a synthesized framework that can be used in either field.
The paper is organized as follows. Section 3.2 discusses the data, sets up a motivat-
ing framework and investigates two standard methods, portfolio sorts and simple Fama-
MacBeth regressions, that a portfolio manager could employ to predict future returns.
Section 3.3 explains tree-based conditional portfolio sorts in detail. Section 3.4 applies the
method to past return predictor variables and section 3.5 has further results on transac-
tion costs, the performance during the recent financial crisis, medium-term momentum,
and a risk factor vs characteristics interpretation. Section B.3 in the appendix illustrates
robustness of our results along many dimensions. Section 3.6 concludes.
Before we introduce tree-based conditional portfolio sorts, we analyze a few standard ap-
proaches that an investor might try: single variable selection, that is, investing based on
the single best-performing variable in historical data over a certain time window; standard
Fama-MacBeth regressions, that is, a multivariate prediction that combines historically
important signals; and Fama-MacBeth regressions that include variable interactions.
5
For example, Tsai et al. (2011) or Huerta et al. (2013).
6
The variables that this literature uses for prediction are typically not motivated by results from the
finance literature, but they are chosen based on their availability in different datasets (convenience sam-
ples). In analyzing the predictions itself, the joint hypothesis problem (Fama (1965, 1970)) is usually
ignored and the evaluation is conducted for raw return estimates.
3.2 Data, motivating framework and standard methods 35
3.2.1 Data
Since we will use the relation between past returns and future returns as a running example
throughout the article, we start by describing the data and the variable construction first.
The basis for our analysis is the universe of monthly U.S. stock returns from the Center
for Research in Security Prices (CRSP) from 1963 to 2012. Since we use firm characteristics
from Compustat and IBES in some robustness checks, we match stock price data to those
datasets first, and we focus our analysis on those firms that can be linked in all datasets.
Firm characteristics include traditional variables like size, book-to-market, dividend yield,
gross profitability, and 82 others that are described in more detail in appendix B.3.1.
The number of firms in our sample varies over time between 1182 and 6626. Size, value,
momentum factors, and the risk-free interest rate are taken from Kenneth French’s data
library.7
Figure 3.1 illustrates how return-based predictor variables are constructed. Suppose
that the investor wants to form a portfolio at the formation time, tf . Return-based pre-
dictor variables can be defined by two parameters: the gap between the time of portfolio
formation and the most recent month that is included in the return calculation, and the
length of the return computation horizon. We denote the former by g, the latter by l and
a return function by Ri,tf (g, l) which maps returns into cross-sectional decile ranks. For
example, Ri,tf (1, 11) = 10 implies that firm i is in the highest decile of returns at time tf
for the return that is computed over the previous 12 months and that leaves out the most
recent one.
Figure 3.1: Construction of past return-based characteristics: The investor forms a port-
folio at time tf . Return-based predictor variables can be defined by two parameters; the
gap between the time of portfolio formation and the most recent month that is included
in the return calculation, and the length of the return computation horizon. We denote
the former by g, the latter by l and a return function by Ri,tf (g, l) maps returns into
cross-sectional decile ranks.
Our benchmark set of predictors contains all one-month returns over the two years
before portfolio formation — that is, Ri,t (g, 1), g = 0, . . . , 24. Much of the related literature
is based on sorting firms into 1 of 10 deciles depending on the values of a sorting variable.
7
[Link]
36 3. Tree-Based Conditional Portfolio Sorts
When we consider return-based strategies below, we refer to buying the upper decile and
selling the lower decile based on Ri,t (g, l). As in Novy-Marx (2012), we will use the notation
Ri,t (g, l) to denote both the return for portfolio formation and the strategy return based
on that simple sorting strategy.8
The problem of predicting future returns based on past returns has the ingredients that
make it difficult for an investor to find the relevant signals: Should momentum be measured
over the most recent 6 or 12 months? What if the signals go in opposite directions? Should
one leave out the most recent month? Or the most recent six (Novy-Marx (2012))? Degrees
of freedom in choosing the gap and length parameters above contribute to the fact that
these questions do not have definitive answers yet.
Here, the expectation of ri,t+1 is formed at time t (we take a period to be one month in
what follows), and the function ft () that maps the information set into expected returns
can be time-varying. The information set Θit can contain data on the firm’s past earnings,
balance sheet information, past stock return movements, and many other variables. Since
we will focus on the relation between past and future returns in this paper, and in line with
the sorting-based literature, we assume that the information set consists of decile rankings
of companies over the past two years — that is, Θit = {Ri,t (0, 1), . . . , Ri,t (24, 1)}. In
other words, we consider decile rankings for each of the most recent twenty-five one-month
returns.
With that information set, adding an additive error term and choosing the common
specification of a linear form (see, for example, Haugen and Baker (1996), Daniel and
Titman (1997) or Brennan et al. (1998)) for the function ft (), equation (3.1) can be written
as
24
X
ri,t+1 = a + βgt Ri,t (g, 1) + i,t , (3.2)
g=0
8
We have checked that results are robust when future returns are computed over the next future month,
but we skip a day to make sure that the return would actually be implementable.
9
At a greater level of generality, one could write the model as
which would also include risk factors, and zit and λt and their histories are subsumed in the information
set Θit = {zi,t , . . . , λt , . . .} at time t. We disregard this aspect for now but note that our framework easily
extends to the case where all returns are interpreted as excess returns over risk factors.
3.2 Data, motivating framework and standard methods 37
portfolios, S1 to S4 ; for example, the stocks in portfolio S1 in the figure have expected return
E[ri,t+1 |R(g (1) , 1) ≤ τ (1) , R(g (2a) , 1) ≤ τ (2a) ].
𝑺𝟏 𝑺𝟐 𝑺𝟑 𝑺𝟒
Figure 3.2: Schematic representation of a conditional portfolio sort: First, observations are
sorted into two portfolios based on past return R(g (1) , 1) and threshold τ (1) . The resulting
portfolios are then sorted again on variables R(g (2a) , 1) and R(g (2b) , 1) with thresholds τ (2a)
and τ (2b) for a total of four portfolios S1 , S2 , S3 and S4 .
A simple way to test whether R(g (2a) , 1) provides additional information over R(g (1) , 1)
would be to compare the sorts on R(g (2a) , 1) within each portfolio sorted on R(g (1) , 1).12
One could also test whether R(g (2a) , 1) creates a return spread only in the portfolio of,
say, low R(g (1) , 1) firms, thereby testing for a potential interaction between characteristics
R(g (2a) , 1) and R(g (1) , 1). In supplementary appendix B.4, we illustrate a basic conditional
portfolio sort with a few standard firm characteristics.
hard thresholds that are sensitive to small changes in the data, their predictions do not
work very well out of sample. Following Kleinberg (1990, 1996), Ho (1998), and Breiman
(2001), we average over many tree-based conditional portfolio sorts to smooth out the
decision boundary, which improves predictions significantly, as explained below in more
detail.
Our approach draws on parallel concepts from the machine learning literature. The
techniques that we use to estimate tree-based conditional portfolio sorts mirror those that
are used to estimate a so-called decision tree in computer science.13 Model averaging or
ensemble methods are also developed in that literature, and they are successfully applied
to areas as diverse as biology (DNA sequencing), psychology, and motion sensing. Applica-
tions in economics are rare,14 and our paper can also be read as an attempt to investigate
whether these techniques provide value-added to academic research in finance and eco-
nomics. This is the first paper that interprets conditional portfolio sorts from a machine
learning perspective, tailors the methodology to similar approaches well-known in finance,
and applies it to a comprehensive financial dataset.
Estimation
We start by describing how variables are selected and how thresholds are estimated. The
goal is to estimate the expectation of the return of firm i in period t + 1 conditional on
information in period t, as in equation (3.1).
To illustrate the method, start out with the conditional portfolio sort in figure 3.2.
Consider the portfolio S1 in that figure, which is defined by variable R(g (1) , 1) being less
than threshold τ (1) and variable R(g (2a) , 1) being smaller than threshold τ (2a) . Other port-
folios can be defined similarly by their relations between sorting variables and associated
thresholds. Within each portfolio Sl , the predicted expected return is modeled as the
average return, µl , of all firms in the portfolio; that is,
13
For further reading on decision-trees, see Hastie et al. (2009), Zhang and Ma (2012), Murphy (2012),
or Criminisi and Shotten (2013).
14
A few examples in a macroeconomic context use decision trees to analyze currency crises (Kaminsky
(2006)), sovereign debt crises (Manasse and Roubini (2009)), banking crises (Duttagupta and Cashin
(2011)) or to develop early warning indicators for, say, excessive credit growth (Alessi and Detken (2014)).
3.3 Estimation strategy 41
L
µ̂l 1(Firm i ∈ Sl in period t),
X
r̂i,t+1 = (3.5)
l=1
giving a portfolio-specific expected return prediction for each observation. What we have
described so far is nothing more than a formal definition of the common conditional sorting
methodology that we carried out in the previous section.
Of course, the conditional sort does not need to end after two levels but can be computed
at greater depth. We consider the case in which the depth of the conditional sort, the
sorting variables, and associated thresholds are not preselected but need to be identified
from the data.
While it can be shown that finding the optimal solution to this problem requires solving
an optimization problem for which a computationally fast solution does not exist (see Hyafil
and Rivest (1976)), there exist feasible approximations. We use a standard algorithm that
proceeds step-wise and that was first suggested in Breiman et al. (1984). For the interested
reader, the supplementary appendix B.5 explains the algorithm in detail and discusses
related methods.
Figure 3.3 illustrates the results of the procedure using the data and variables described
in section 3.2.1. Rather than showing the entire iterative sort, the figure only shows the
first few nodes. The first selected split variable is R(0,1), the return over the previous
month. The associated threshold is 6; that is, all firms with a return over the previous
month in the lowest 6 deciles are sorted into one portfolio, and the remaining ones are
sorted into the other. Conditional on this split, R(0,1) is selected again in the left branch
at the next level and R(2,1), the one-month return two months ago, is selected in the right
branch. The actual iterative sort goes deeper but, for illustration, we have computed the
one-month-ahead returns in each of the four subsets. Differences are already pretty stark:
The subset S1 , which is the set of companies that were in the lower of the two R(0,1) groups,
display the highest return, indicating short-term reversal. The right branch illustrates a
momentum effect: Stocks with higher values of R(2,1) have a higher subsequent return on
average.
Model averaging
Constructing tree-based conditional portfolio sorts in the way we have described results in
a few challenges. First, as described earlier, because of the complexity of the optimization,
we have to use a greedy algorithm to estimate the model. This algorithm, however, does not
guarantee that thresholds and split variables are selected optimally at each node. Second,
the threshold rule is discrete, and any error in the estimation of the threshold could greatly
distort the correct path for any expected return that is supposed to be predicted from the
estimated model. Third, our initial results showed that a single estimated tree-based
conditional portfolio sort summarizes the estimation data well, but the model does not
extend well to new observations. In other words, because there are so many degrees of
42 3. Tree-Based Conditional Portfolio Sorts
Figure 3.3: tree-based conditional portfolio sort using the entire data set: First nodes
freedom (variables and thresholds) at each step, the tree-based conditional portfolio sort
can often overfit the estimation sample.
These problems are well known in the machine learning literature, and we adopt a com-
mon solution suggested in Breiman (2001). The idea is to estimate a tree-based conditional
portfolio sort a number of times using only a random subset of variables each time. The re-
sulting models are less prone to overfitting because they are arguably less complex. At the
same time, we also only use subsets of the data to estimate each model. We then compute
estimates for expected returns from each model and average over all models’ estimates to
get a final prediction.
The idea of combining many predictions to construct a more accurate one can be
illustrated in a simple voting setup in which people use majority voting to make a decision
or to determine the (objective) value of an object. If everyone has the same information
set, then nothing can be learned from aggregating individual votes; instead, every single
vote is a sufficient statistic for the outcome. Only if voters differ in their information can
aggregation lead to a more precise estimate. Using subsets of data and variables induces
just such different information sets.
More formally, let B be the number of tree-based conditional sorts that are computed,
and let fˆb (Θit ) be the predicted expected return for stock i at time t that is based on model
3.3 Estimation strategy 43
In all results that follow, we construct 200 tree-based conditional portfolio sorts (that
is, B = 200) and we use 8 out of 25 regressors (that is, roughly 30 percent of the number of
regressors) in each of them. We have tried other values for the share of sampled regressors
(between 20 and 40 percent) and also larger values for the number of estimated tree-based
conditional portfolio sorts but have found that results do not vary much with these choices.
We settled on the share of 30 percent of regressors because it is a standard recommendation
in the random forest literature, and we chose B = 200 because higher values did not have
any apparent benefit for the estimation but are more costly in terms of computation.
Discussion
Our ultimate goal is to provide a new method that is capable of tracing out which firm
characteristics predict the cross-section of stock returns well. Tree-based conditional port-
folio sorts are potentially interesting because they can account for both the correlation and
the interactions of candidate characteristics. Model averaging as described earlier protects
against the risk of in-sample overfitting and deals with the hard thresholds that sorting
induces.
The flexibility of our approach does not come without costs: Model averaging loses
the simple interpretation from a single tree-based conditional portfolio sort. Moreover, we
cannot summarize our model as a simple linear equation in the space of firm characteristics
and factors. One reason for the popularity of linear regression methods certainly lies in their
apparent transparency. Our approach draws on methods from computer science that are
sometimes criticized for producing black box predictions that cannot easily be interpreted.
One contribution of this paper is to introduce measures with which the relation between
model predictions and regressors can nevertheless be evaluated transparently.
Variable importance Since the relevance of a variable is determined by both its level
and its potential interactions with other variables, summarizing statistical significance via a
simple t-test is not appropriate. Instead, we rely on a relative variable importance measure
that can be interpreted similarly to t-statistics in simple regressions.
For each predictor variable and each tree-based conditional portfolio sort, we compute
the mean squared error (MSE) of the prediction when the values of that variable are
randomly permuted, and we express its MSE relative to the model’s MSE when all variables
are at their original values. This fraction is then averaged over all iterative conditional
sorts and predictor variables are ranked by this measure, where higher values imply that
random permutations of a predictor variable cause higher increases in mean squared error,
and the predictor variable is therefore considered more relevant.
44 3. Tree-Based Conditional Portfolio Sorts
Results are typically displayed relative to the predictor variable that causes the highest
increase in mean squared error when it is permuted, a convention that we follow. For
example, a value of .8 for a predictor variable means that this variable is associated with
an MSE increase equal to 80 percent of the variable with the highest MSE increase.
g,d 1 1 1 Xˆ
r̂i,t+1 = fb (Rit (g, 1); Rit (g − , 1)).
N T B i,t,b
Repeat this for all values of d and graph the results for each past return g and each
value of d. Our method can easily be extended to varying two (or more) variables at the
same time. In section 3.4, we also report partial derivatives for two-way interactions of
return variables.
3.4 Empirics
We apply our method to the prediction of future returns based on past returns. We will
provide evidence for the following results. First, tree-based conditional portfolio sort works
3.4 Empirics 45
months
𝑡 − 60 𝑡 𝑡+1 𝑡 + 12
In-sample Out-of-sample
Figure 3.4: Out-of-sample testing: tree-based conditional portfolio sorts are re-estimated
every year with data over the past sixty months. Predicted returns are then calculated for
the next twelve months. The strategy is go long (short) the highest (lowest) decile of those
predictions each month.
46 3. Tree-Based Conditional Portfolio Sorts
well in this setting in the sense that expected return predictions are ordinally accurate.
Strategy returns and information ratios based on the model’s predictions are much higher
than those from alternative models. Second, among return-functions, the most important
ones refer to the more recent past. Third, superior predictive ability can be traced to
flexibly dealing with nonlinear relations between past and future returns, and interaction
effects between past return functions. The relation between past and future returns is more
complex (and more predictable) than can be captured by any one summary return.
Figure 3.5: Annual strategy return: The strategy is based on the predictions of tree-based
conditional portfolio sorts that relate future returns to past decile sorts of returns. Past
return sorts include decile rankings R(g,l) with length l equal to 1 and gap g between 0
and 24 months (i.e. all one-month returns over the two years before portfolio formation).
The strategy goes long the highest decile of predictions and goes short the lowest decile of
predictions each month. The figure shows the annual return for each of forty-five out-of-
sample predictions.
not greatly, above those of the standard methods in section B.1, the method seems to do
so with a large reduction in variance.
Table 3.2 sheds more light on the decile portfolios that are formed based on the models’
predictions. They show the factor loadings of each decile portfolio return for one of four risk
models. The returns of all decile portfolios appear to correlate one to one with the market
return, with the extreme portfolios experiencing a slightly higher covariance. Second, there
is no apparent spread in factor loadings for the size and the value factor. The extreme
portfolios load slightly higher on the size factor (an issue that we come back to in appendix
B.3) and slightly lower on the value factor. Third, there is a monotone relationship of decile
returns with respect to loadings on the momentum factor. Quantitatively, however, these
differences are small. Fourth, even though none of these portfolios differ much in their
loadings on risk factors, there is a strong monotone relation between the portfolios and
their (risk-adjusted) average returns. This stands in stark contrast to the seemingly very
similar portfolios in terms of risk loadings. What is more, this relation is not only driven
3.4 Empirics 49
Figure 3.6: Earned profit from investing $1 in the strategy in 1968: The strategy is based
on the predictions of a tree-based conditional portfolio sort that relates future returns to
past decile sorts of returns. Past return sorts include decile rankings R(g,l) with length l
equal to 1 and gap g between 0 and 24 months (i.e. all one-month returns over the two
years before portfolio formation). The strategy goes long the highest decile of predictions
and goes short the lowest decile of predictions each month. The figure shows the earned
profit from investing $1 in the long and the short portfolio, respectively. For reference, the
figure also includes the returns to investing $1 at the riskfree rate and for investing at the
rate of the market return over the same horizon.
15
In unreported monotonicity tests based on Patton and Timmermann (2010), we confirm that raw
and risk-adjusted returns are monotonically increasing in deciles at all levels of significance (available on
request).
50 3. Tree-Based Conditional Portfolio Sorts
Figure 3.7: Average monthly decile return for strategy return and simple return strategies:
The strategy is based on the predictions of a tree-based conditional portfolio sort that
relates future returns to past decile sorts of returns. Past return sorts include decile
rankings R(g,l) with length equal to 1 and gaps between 0 and 24 months. The strategy
goes long the highest decile of predictions and goes short the lowest decile of predictions
each month. Simple return strategies are plotted for comparison. R(1,5) is the strategy
that goes long (short) the highest (lowest) decile of returns over the past six months, leaving
out the most recent one. R(6,6) is Novy-Marx (2012)’s intermediate return strategy that
goes long (short) the highest (lowest) decile of returns that are computed over the six
months that skip the most recent six months.
Table 3.2: Factor loadings of decile portfolios: tree-based conditional portfolio sort
CAPM
Intercept -1.52 -0.73 -0.48 -0.32 -0.23 -0.13 0.01 0.06 0.23 0.72 2.23
(-8.59) (-4.97) (-3.60) (-2.47) (-1.81) (-1.03) (0.08) (0.42) (1.55) (3.97) (16.04)
MKT 1.12 1.08 1.06 1.04 1.04 1.06 1.06 1.09 1.14 1.20 0.07
(27.40) (30.01) (31.94) (32.94) (30.65) (31.83) (30.62) (28.90) (29.10) (25.30) (2.14)
Three-factor model
Intercept -1.64 -0.87 -0.63 -0.48 -0.38 -0.28 -0.13 -0.07 0.10 0.61 2.25
(-13.12) (-7.71) (-6.80) (-5.58) (-4.66) (-3.45) (-1.42) (-0.71) (0.98) (4.51) (16.51)
MKT 0.99 0.97 0.97 0.96 0.97 0.98 0.98 0.99 1.02 1.04 0.05
(27.70) (29.22) (34.92) (36.04) (38.43) (38.76) (32.98) (29.14) (32.43) (27.55) (1.53)
SMB 0.87 0.74 0.69 0.67 0.65 0.66 0.69 0.71 0.80 0.95 0.08
(8.64) (8.49) (8.76) (8.51) (8.80) (8.69) (8.51) (8.84) (10.21) (12.08) (1.40)
HML 0.23 0.25 0.27 0.30 0.29 0.28 0.27 0.25 0.25 0.21 -0.03
(2.84) (3.58) (4.44) (4.73) (4.64) (4.44) (4.24) (3.50) (3.68) (2.52) (-0.39)
Four-factor model
Intercept -1.37 -0.67 -0.48 -0.36 -0.29 -0.21 -0.07 -0.02 0.13 0.69 2.05
(-12.77) (-7.10) (-6.32) (-4.96) (-4.11) (-3.10) (-0.90) (-0.25) (1.40) (5.13) (14.54)
MKT 0.94 0.94 0.94 0.93 0.95 0.97 0.96 0.98 1.02 1.03 0.09
(30.12) (31.11) (37.89) (39.00) (41.63) (44.66) (36.30) (31.58) (35.46) (27.00) (2.78)
SMB 0.86 0.73 0.69 0.67 0.65 0.65 0.69 0.71 0.80 0.95 0.09
(11.69) (11.04) (10.88) (10.14) (10.15) (9.58) (9.15) (9.38) (10.52) (13.16) (1.69)
HML 0.15 0.19 0.23 0.26 0.25 0.26 0.25 0.23 0.24 0.18 0.04
(2.47) (3.39) (4.40) (4.91) (4.96) (4.65) (4.53) (3.62) (3.86) (2.34) (0.61)
UMD -0.27 -0.20 -0.15 -0.12 -0.10 -0.07 -0.06 -0.05 -0.03 -0.08 0.20
(-7.71) (-6.29) (-4.87) (-3.82) (-2.90) (-2.26) (-1.86) (-1.49) (-0.93) (-2.34) (5.57)
This table shows time-series regressions of decile portfolio returns on factors. Returns are specified in percent per month. Each decile is formed
51
on the predicted returns of a tree-based conditional portfolio sort that relates future returns to past decile sorts of returns. Past return sorts
include decile rankings R(g,l) with length l equal to 1 and gap g between 0 and 24 months. Predictions are based on the model in section
3.3.2. Low denotes the lowest decile of predicted returns and High denotes the highest decile of predicted returns. The first panel reports
the average return, the second panel reports CAPM estimates, the third reports the three-factor model estimates and the fourth panel adds
momentum. MKT is the market return, SMB and HML are the Fama-French factors for size and value, and UMD is the momentum factor.
52 3. Tree-Based Conditional Portfolio Sorts
While the strategy returns in our tree-based conditional portfolio sort appear high, they
could still disappear after taking transaction costs into account. Strategies that are based
on past returns generally have been found to have relatively high turnover (see de Groot
et al. (2012) or Frazzini et al. (2015)), especially so when they are based on recent past
returns. As the tree-based conditional portfolio sort mainly exploits variation in the most
recent past returns, we expect turnover to be high as well.
Appendix B.2 shows that this expectation is correct: An equal-weighted hedge strategy
that goes long $1 and short $1 in the extreme portfolios has an average monthly turnover
of 318 percent. Turnover is also high using the less extreme hedge returns that go long the
ninth or eighth decile and that go short the second or third decile, respectively.16 However,
appendix B.2 also shows that our strategy has positive excess return after transaction costs
based on an extrapolation of transaction cost estimates from Frazzini et al. (2015).
Tree-based conditional portfolio sorts appear to work well in our application in the sense
that they produce high and stable excess returns out of sample that are not explained by
standard factor models. This begs the question what the method finds that researchers
have not paid attention to. We discuss the discovered structure of predictor variables next.
Recall that we re-estimate the model each year for a total of 45 different estimated models
over time. When we compute our measure of predictor variable importance for each year,
this gives us a ranking of the importance of each variable in each year. As a first summary,
we rank past returns by their median rank in these 45 models. Table 3.3 shows the median
rank as well as the upper and lower quartile of ranks for each of the top 10 past returns.
The top four return functions are related to the most recent six months of returns;
all return functions over the most recent six months enter the top 10. In addition, some
returns that show up provide information about the intermediate return between 6 and 12
months before the formation date. In particular, it is interesting and reassuring to see past
return functions considered in the preceding literature to rank highly in the list. R(0,1),
the return over the most previous month, is the return function of Jegadeesh (1990) and
many other papers, while R(11,1), the one-month return exactly 12 months ago, is the
seasonal effect documented by Heston and Sadka (2008).
There is also considerable time variation in the exact ranks as illustrated by the in-
terquartile range of ranks for each past return. All of them were in the top half for more
than 50 percent of the time, and 7 out of the 10 return functions are in the top 10 for at
least half of the years. On the other hand, each variable also had periods during which it
appears less relevant to the prediction, as expressed in the last column of the table. We
computed the rank correlation of past returns’ importance between subsequent years and
16
These numbers are similar to those reported in de Groot et al. (2012) or Frazzini et al. (2015) for
strategies based on short-term returns.
3.4 Empirics 53
found it to be around 0.7, which points to the fact that the structure is relatively stable
over time.
The fact that the pattern of more recent returns being more relevant than more distant
past returns comes out of an agnostic search procedure is intriguing. We find that it is a
quite robust fact in the data throughout various specifications. For instance, we find very
similar results for past-return-based variables when we include other firm characteristics
in the estimation (appendix B.3.1). In appendix B.3.2, we consider an expanded set of
predictor variables that uses 126 past return functions of different gaps and different lengths
such that standard past return functions like R(0,6) (the return over the most recent six
months) are also part of the set of regressors. In that exercise, all 10 predictor variables
are related to the most recent 6 months of returns and, what is more, the top 6 return
functions are returns of length one that, taken together, summarize the most recent six-
month return. The fact that a standard return like R(0,6) is not chosen but its components
are illustrates that using the return over the previous year alone (and not the one-month
returns that it is based on) leads to a loss of relevant information. One-month returns
contain important information that is neglected when summary returns such as R(0,6) or
R(1,11) are considered. For both sets of past-return functions, we repeat the estimations
by firm size in appendix B.3.3 and again find similar results.
Our first intermediate result is, thus, that tree-based conditional portfolio sorts work
because they effectively exploit variation in relatively recent one-month returns. The next
sections look at how these variables are combined.
thus, help to identify stocks with low expected returns but do not necessarily help much to
identify stocks with high expected returns. It is only when we consider one-month returns
that are in the more distant past (more than four months out) that we find a standard
momentum effect — that is, a monotonically positive and close-to-linear relation between
past and predicted returns.
56
Figure 3.8: Average partial derivatives for return characteristics: The figure shows the average prediction when a charac-
teristic is counterfactually varied from low to high values. Details are in section 3.3.2. The first row shows results when
we use only twenty-five past one-month returns as predictors. The second row shows results for the same one-month
return functions when other firm characteristics (defined in appendix B.3.1) are included in the estimations as additional
variables. Each column shows one return characteristic and predictions are averaged over the sample period.
3. Tree-Based Conditional Portfolio Sorts
3.4 Empirics 57
The literature has not paid much attention to nonlinear relations between past and
future returns. However, given that a) predictions from our tree-based conditional portfolio
sorts make high risk-adjusted excess returns, b) short-term return functions have high
values in our predictor variable importance calculations, and c) the partial effects of these
variables cannot all be linearly related to returns, it appears that nonlinearities should be
investigated further in future research.
Figure 3.9 shows contour plots for all two-way interactions of the most recent one-
month return functions. In each panel, darker areas represent lower return predictions and
brighter areas represent higher return predictions. A couple of interesting results stand out:
First, many return variables interact in nonlinear ways. For example, the upper-left panel
shows the interaction of R(0,1), the most recent one-month return, and R(1,1), the return
over the preceding month. Return predictions generally decrease in the value of R(0,1),
reflecting short-term reversal. However, within high values of R(0,1), return predictions
increase in R(1,1), while they decrease in R(1,1) within low values of R(0,1). This type of
nonlinearity holds, to a varying extent, in many panels involving R(0,1). Second, for some
return variables, we find monotonically increasing predictions within both return variables,
mostly for those that involve returns from four or more months ago. Third, some return
predictions neither decrease nor increase monotonically in the predictor variable range but
are nonlinearly related to return predictions, once one variable is fixed. For instance, from
figure 3.8 we know that R(1,1) is nonlinearly related to returns. In figure 3.9, we see that
this nonlinearity is more pronounced when R(1,1) is interacted with intermediate returns
like R(3,1) or R(4,1).
58
Figure 3.9: Average double partial derivatives. The figure shows the average prediction when two characteristics are
counterfactually varied from low to high values. Results are based on rolling optimization of the model and predictions
are averaged over the sample period. Details are in section 3.3.2.
3. Tree-Based Conditional Portfolio Sorts
3.4 Empirics 59
Finally, we find evidence that the estimate average partial derivatives are time-varying.
Figures 3.10 and 3.11 illustrate this for two different variables. Figure 3.10 shows average
partial derivatives in eight different years, evenly spaced over the sample period, for R(0,1),
the return over the previous month. Short-term reversal is detectable across all years, but
its strength varies over time. While our model estimates indicate relatively monotone (or
regular) short-term reversal across all 10 deciles for the first half of the sample, short-
term reversal is more apparent in the extreme deciles in the second half of the sample.
Similar conclusions can be drawn from figure 3.11, which shows the same calculations for
R(5,1), the one-month return six months before portfolio formation. In the first half of the
sample, momentum is apparent and robust across all deciles. In the second half, however,
differences in average partial derivatives are more pronounced between extreme deciles
than between intermediate deciles. Interestingly, recently (in 2012, the lower-right panel),
the average partial derivative of R(5,1) has reversed such that lower values of R(5,1) are
associated with higher returns in the model estimates. Recall that the estimation period
for this panel is 2006 to 2011, which coincides with an episode of a momentum crash as
documented by Daniel and Moskowitz (2015). As we have shown in the previous section,
a trading strategy based on our model estimates has not suffered the strong crash that a
standard momentum strategy has experienced in this period. The average partial derivative
at that time indicates that the model has picked up the weakness of standard momentum
and that the estimated relationship was adjusted (in that case, reversed) accordingly.
60
Figure 3.10: Average partial derivatives in different years. The figure shows the average prediction when R(0,1), the
return over the previous month, is counterfactually varied from low to high values, and results are displayed for different
years to illustrate time-variation. Results are based on rolling optimization of the model, details can be found in section
3.3.2.
3. Tree-Based Conditional Portfolio Sorts
3.4 Empirics
Figure 3.11: Average partial derivatives in different years. The figure shows the average prediction when R(5,1), the
one-month return six months before portfolio formation, is counterfactually varied from low to high values, and results
are displayed for different years to illustrate time-variation. Results are based on rolling optimization of the model, details
can be found in section 3.3.2.
61
62 3. Tree-Based Conditional Portfolio Sorts
Figure 3.12 shows the time-variation in the CAPM beta for the long- and the short-
portfolio of the strategy. The loadings do not display systematic variation over time and
the difference between the betas fluctuates around zero. As such, it is no surprise that the
strategy return is less affected by the aggregate market.
Figure 3.12: Beta exposure of the long- and short-portfolios: Loadings of the top (solid
line) and bottom (dashed line) decile portfolios on the market factor. Deciles are based on
predicted returns of the tree-based conditional portfolio sort in section 3.4.
The solid lines in figure 3.13 show the time-varying performance of typical momentum
strategies between 2000 and 2010: Both R(11,1), the strategy that is based on the return
from exactly 12 months ago, and R(5,1), the strategy that is based on the return from
exactly 6 months ago, display a sharp performance decrease in 2008-2009. The figure
illustrates that both variables indicate reversal, especially in crisis times.
The dashed lines in figure 3.13 show that this is how both return variables were incorpo-
rated into the strategy during that time period: While the strategy is typically positively
related to both variables (a value greater than zero means that the strategy is long stocks
that had positive exposure to the simple past returns), exposure became negative or close
to zero right around the financial crisis, such that the portfolio positions were reversed
relative to a standard momentum strategy based on these variables. This illustrates the
algorithm picked up on changes in the underlying relationship between past and future
returns and was therefore able to avoid the drawdown in 2009.
3.5 Discussion and robustness 65
(a) R(5,1)
(b) R(11,1)
Figure 3.13: Simple strategy returns and exposure of the tree-based strategy: Solid lines
show the performance of simple long-short strategies that are based on stock returns from
6 months (R(5,1) or 12 months (R(11,1) ago. Dashed lines show the exposure of the
tree-based strategy to the simple return strategies.
Table 3.4: Strategy factor loadings: Short-term and intermediate-term return functions
Dependent variable
Return of short-term strategy Return of intermediate-term strategy
(1) (2) (3) (4) (5) (6) (7) (8) (9) (10)
Intercept 2.15 2.12 2.17 2.04 1.05 1.78 1.74 1.78 1.44 0.16
(19.09) (18.47) (19.35) (17.37) (9.11) (15.01) (14.90) (15.52) (13.33) (1.65)
MKT 0.03 0.01 0.03 -0.04 0.05 0.04 0.10 0.08
(1.38) (0.46) (1.64) (-2.37) (1.40) (1.32) (4.25) (4.39)
SMB 0.04 0.04 0.06 -0.05 -0.04 -0.06
(0.65) (0.85) (1.95) (-0.74) (-0.93) (-2.27)
3.5 Discussion and robustness
and standard errors were clustered using Newey-West’s adjustment for serial correlation.
68 3. Tree-Based Conditional Portfolio Sorts
In column (10), we do the same and add the short-term strategy return to the factor
regression for the intermediate-strategy return. Interestingly, alpha disappears almost
entirely once the short-term strategy return is accounted for.
We interpret this as evidence that the most important variation for return predic-
tion purposes stems from short-term variation in returns rather than intermediate-term
variation once interactions and confounding returns are included in the estimation. This
reconciles the result in Novy-Marx (2012) with Goyal and Wahal (2015), who cannot find
the intermediate-term momentum effect in 37 out of 38 markets.
This motivates table 3.6, which mirrors the analysis in Haugen and Baker (1996). It
shows the average values of various firm characteristics in each decile of expected returns.
The first panel of the table shows measures of risk across the ten deciles with no clear
3.5 Discussion and robustness 69
(monotone) pattern. Average market beta is higher in the extreme deciles. The same
holds for the profitability measures in the second panel. Interestingly, gross profitability
is very similar in each decile, but expected returns are very different. This illustrates that
our sorting is not driven by the Novy-Marx (2013) measure of gross profitability. Panel
3 shows that book-to-market is balanced across deciles, as one would expect from the
balanced factor loadings in table 3.2. The last panel shows that the firms in the extreme
deciles are, on average, smaller.
Table 3.6: Firm characteristics: Portfolios based on tree-based conditional portfolio sort 70
Dec. 1 Dec. 2 Dec. 3 Dec. 4 Dec. 5 Dec. 6 Dec. 7 Dec. 8 Dec. 9 Dec. 10
Risk
Debt to Equity 2.61 2.33 2.75 2.47 3.29 2.52 2.70 3.62 2.60 3.55
Long-term debt to Equity 1.43 0.78 1.15 0.81 1.61 0.74 0.92 1.98 0.90 2.09
Debt Ratio 0.51 0.52 0.53 0.53 0.54 0.54 0.54 0.54 0.54 0.54
Beta 1.18 1.09 1.07 1.05 1.05 1.05 1.06 1.08 1.12 1.19
Profitability
Gross Profitability 0.34 0.34 0.34 0.34 0.34 0.34 0.34 0.34 0.34 0.34
Return on Assets -0.02 0.01 0.02 0.02 0.02 0.03 0.02 0.02 0.01 -0.02
Return on Equity -0.15 0.03 -0.07 0.05 -0.15 0.04 0.04 -0.25 -0.18 -0.33
Profit Margin -2.79 -1.31 -1.20 -1.04 -1.11 -1.60 -1.11 -0.75 -1.20 -2.41
Gross Margin -1.31 -0.63 -0.26 -0.31 -0.19 -0.22 -0.27 -0.16 -0.37 -1.04
Earnings per Share 0.87 1.31 1.45 1.58 1.63 1.64 1.59 1.55 1.34 0.93
Basic Earnings Power Ratio 0.04 0.06 0.07 0.07 0.08 0.08 0.08 0.07 0.07 0.03
Price level
Price Earnings Ratio 4.40 5.18 4.68 6.68 6.15 5.46 3.87 7.04 3.99 4.52
Book to Market 0.79 0.81 0.83 0.83 0.83 0.83 0.83 0.83 0.83 0.87
Price Sales Ratio 2.00 1.35 0.91 0.88 0.84 0.88 0.88 0.71 0.93 1.42
Dividend Yield 0.04 0.03 0.03 0.04 0.04 0.03 0.04 0.03 0.03 0.03
Activity
Current Ratio 3.17 2.94 2.83 2.78 2.78 2.74 2.78 2.77 2.82 2.93
Quick Ratio 2.05 1.88 1.81 1.77 1.76 1.74 1.76 1.77 1.79 1.86
Net Working capital Ratio 0.30 0.29 0.28 0.27 0.27 0.27 0.27 0.28 0.28 0.29
Cash Ratio 1.49 1.23 1.13 1.07 1.07 1.04 1.07 1.06 1.11 1.22
Assets - Turnover Ratio 1.17 1.17 1.16 1.15 1.15 1.15 1.15 1.16 1.18 1.21
Inventory-Turnover Ratio 20.60 19.24 20.12 23.77 22.12 23.97 23.85 21.34 22.33 20.25
RandD 0.09 0.08 0.07 0.07 0.07 0.07 0.07 0.07 0.07 0.09
Others
Size 681.95 976.24 1113.54 1177.19 1180.10 1177.22 1156.71 1085.86 967.02 633.43
3. Tree-Based Conditional Portfolio Sorts
Each month, all stocks are ranked by their estimated expected return based on a tree-based conditional portfolio sort that is based on all one-
month return functions over the two years before portfolio formation. The table reports the average value of each firm characteristic in each
decile over time.
3.6 Conclusion 71
More intriguingly, since the strategy is based on the extreme deciles, it is worthwhile
to compare the average values within these two deciles. Note that the values of most firm
characteristics are very similar in these two deciles. The strategy appears to be based on
riskier, less profitable, and smaller companies, on average. Yet, within the set of these
firms, there are stark differences in returns that can be systematically predicted.
Recall that alternative strategies that are based on buying the second (third) highest
decile and selling the second (third) lowest decile rather than the extreme portfolios also
make robust excess returns. Comparing deciles two and nine, or deciles three and eight,
illustrates that the two corresponding portfolios are again very balanced throughout char-
acteristics. As a sole characteristic, return-on-equity is lower in decile nine (eight) than
in decile two (three), but other measures of profitability indicate that the portfolios are
comparable along this dimension. While the non-extreme decile portfolios display similar
characteristics, their excess returns vary and (see table 3.2) can be predicted from past
returns. Since the portfolios based on the tree-based conditional portfolio sort appear not
to be discernible based on many characteristics, we would not expect the strategy return
that is based on it to help explain other anomalies.
As the strategy return itself appears to be unpriced by equilibrium models and unrelated
to standard characteristics, the return could, in principle, be added as an additional risk
factor to standard equilibrium models. However, in unreported results, we found that the
strategy return, as expected, only weakly helps to explain other asset pricing anomalies,
which is why we prefer the interpretation from a characteristics’ rather than a risk factor
perspective.
3.6 Conclusion
Some 50 years after the Capital Asset Pricing Model of Sharpe (1964b), Lintner (1965b),
and Mossin (1966), and some 20 years after the three-factor model of Fama and French
(1992), there is still a remarkable lack of consensus about which variables can be related
to expected stock returns. To date, the literature has found more than 300 variables that
spread returns in a way that is unaccounted for by the standard equilibrium models. This
has led Green et al. (2013) to conclude that ”either U.S. stock markets are pervasively
inefficient, or there exist a much larger number of rationally priced sources of risk in
equity returns than previously thought.” Surely, many of these variables contain correlated
information, and some will not hold up out of sample, but, so far, the literature has
not rigorously identified which ones are fundamentally important. Furthermore, we have
illustrated that some variables interact in nontrivial ways, making it more challenging to
single out the important ones with standard methodologies.
We introduce a framework — tree-based conditional portfolio sorts — that can deal
with a large number of variables and their potential nonlinearities and interactions. It
also puts emphasis on systematic out-of-sample testing of all results. It connects model
evaluation in finance to the machine learning literature in computer science and can serve
to bridge the two fields.
72 3. Tree-Based Conditional Portfolio Sorts
We apply our framework to find information in past returns that can be related to future
returns. A simple, linear Fama-MacBeth framework finds moderate excess returns relative
to the four-factor model. Using the same variables in the tree-based conditional portfolio
sort framework, on the other hand, yields high and stable excess returns, indicating that
the linear framework does not exploit all relevant information in the data.
Finance has criticized machine learning for producing black box predictions without any
possibility to ”get insights into the underlying structure of the data” (Breiman, 2002). We
show that, even though the structure does not come in the form of simple equations, one
can still extract interpretable information from the resulting tree-based conditional sorts.
First, we find that, among the prior two years of one-month return functions before portfolio
formation, the more recent ones are the most important for accurate return predictions.
Second, some of these return-functions are nonlinearly related to future returns, mostly
returns between two and four months before portfolio formation. For instance, both high
and low values of the return over the second-to-last-month forecast lower returns. Third,
many of the return functions display nontrivial interactions. For instance, the one-month
return over the second-to-last month, is positively related to returns for stocks with low
returns last month, but it is positively related to returns with high returns last month. At
a minimum, our results indicate that the relation between past and future returns is more
complex than can be captured by any one summary return, such as standard momentum
or intermediate momentum. Our results are robust to including a larger set of correlated
return functions and to the inclusion of other firm characteristics. Similar structures are
also discovered within different size-sorted portfolios.
Lastly, tree-based conditional portfolio sorts can accommodate the inclusion of new
predictor variables quite easily. Starting from the observation that if a predictor variable
is relevant, it should show up among the most important variables that the method finds,
one could just add the variable in question to the existing set of variables. Running the
estimation on this extended set effectively controls for correlations with other variables and
takes potential interactions and nonlinearities into account. Our hope is that a framework
around tree-based conditional portfolio sorts can significantly speed up the process of
scientific discovery in this literature.
Chapter 4
This thesis shows how technological advances can help answering fundamental questions in
finance. With new sources of data and the use of computationally intensive methods from
the field of statistics we are able to solve issues which were quite impossible to solve years
ago. It can be shown that extracting systematically the tone from articles of a newspaper
over time can give interesting insights into the risk-return relation of the stock market.
Furthermore advanced statistical techniques, which are based on decision trees, can help
to analyze the hundreds of factors which claim to explain cross-sectional expected stock
returns simultaneously and with all their non-linearities and interactions.
This thesis gives a lot of possibilities for future research. Because standard settings
has been used in both parts, there is much room for analyzing the tuning parameters of
the algorithms. This is definitely a strength of the results presented in this thesis. But
they may even be much stronger when the parameter may be tuned robustically. Beside
the tuning, both parts left a lot of extensions open for further studies. In the time-series,
chapter 2, one can use other search words like ”economic expansion” and ”bubble”. One
can use the latent Dirichlet allocation topic modelling algorithm to analyze the text on
changing topics. Or other finance-specific lexicons may be used to analyze the text. Also
aside the New York Times, other newspaper sources may be of interest. In the cross-
section, chapter 3, more statistical methods beside the random forest can be tested. It
would also be of interest to see this study applied to the stock market of other countries.
Extending the variable set beyond the the return-based variables would also be a natural
extension.
74 4. Conclusions and Perspectives
Appendix A
First Appendix
Figure A.1: The figure shows the number of occurrences of the word ”recession” in the New York Times
Articles (20 day moving average) from 1900-1935. Shadings show the U.S. recessions according to NBER
recession dates.
76 A. First Appendix
Appendix B
Second Appendix
24
X
t
ri,t+1 = βcons + βgt Rit (g, 1) + it (B.1)
g=0
and keeps either all coefficients (kitchen sink) or uses LASSO to select the relevant
variables.
His period t + 1 forecast is computed based on the rolling average of the coefficient
estimates up to period t − 1 and then applying the linear model to Rit (g, 1), that is,
24
t−1 X t−1
r̂i,t+1 = β cons + β g Rit (g, 1), (B.2)
g=0
t−1 Pt−1
where β g = m1 j=t−1−m β̂gt . We initially use a rolling window of 120 months but, as
in Lewellen (2015), have found that results are robust to varying that parameter.
Lewellen (2015) uses a set of 15 predictor variables that are well established in the
literature. In contrast, we consider an investor who faces substantial uncertainty about
which variables he should include and, therefore, has to cast a wide net. Consistent with
our running example, the investor considers all one-month returns over the two years
before portfolio formation. Each period, he computes return predictions based on past
model estimates and sorts predictions into 10 deciles. He constructs an equal-weighted
hedge portfolio that goes long the highest decile of predicted returns and that goes short
the lowest decile of predicted returns, analogous to the strategies described earlier.
Starting with the kitchen sink model, the first four columns of table B.2 show the
strategy’s factor loadings from time-series regressions on the market, size, value, and mo-
mentum factors. The strategy has a positive and significant average return of 1.51 percent
per month and loads mostly on the market and the momentum factor. The alpha relative
to the four-factor model is about 1 percent per month, with an information ratio of about
1.
When we use the LASSO in the Fama-MacBeth framework as described earlier, results
remain almost unchanged. The last four columns of table B.2 show that the average
strategy return is again around 1.5 percent per month, and the four-factor alpha is 1
percent per month. The information ratio is close to 1, as in the kitchen sink regression.
The reason that these results are very similar is that many irrelevant regressors have
coefficients close to zero in the kitchen sink case.
to the nature of the penalty term (the sum of the absolute values of individual coefficients), the optimum
will typically set many coefficients to exact zeros, which is why the method can be viewed as a variable
selection device.
80 B. Second Appendix
24
X
t
ri,t+1 = βcons + βgt Rit (g, 1) + it .
g=0
The kitchen sink Fama-MacBeth model uses all variables in each period regardless of
their significance, and the LASSO model selects a set of relevant variables each period
based on a penalty function approach. Both procedures are described in section B.1.2.
Strategies go long the highest predicted return decile and go short the lowest predicted
return decile. The sample period covers 1968 to 2012, and all results are based on rolling
out-of-sample estimates of the models. MKT is the market return, SMB and HML are
the Fama-French factors for size and value, and UMD is the momentum factor. SR is
the Sharpe ratio and IR is the information ratio. T-statistics are in parentheses, and
standard errors were clustered using Newey-West’s adjustment for serial correlation.
B.2 Transaction costs 81
Note that the approaches so far have not included variable interactions. The Fama-
MacBeth regression framework lends itself to a simple implementation of additionally in-
cluding interactions of predictor variables. Equation (B.3) shows the regression equation
that adds all two-way interactions among past return rankings:
24
X 24 X
X
ri,t+1 = a + βgt Ri,t (g, 1) + t
γgj Ri,t (g, 1)Ri,t (j, 1) + i,t . (B.3)
g=0 g=0 j>g
Table B.3 shows strategy returns that are based on predictions from equation (B.3).2
At 1.13 percent per month, the average excess return relative to the four-factor model
is slightly higher than in the levels-only version above. The information ratio, however,
experiences a much stronger increase to 1.3-1.4. Hence, the main benefit to including
two-way interactions appears to be a reduction in variance rather than an improved mean
return.
Table B.3: Strategy factor loadings: Fama-MacBeth predictions using all vari-
ables and two-way interactions
24
X 24 X
X
t
ri,t+1 = βcons + βgt Ri,t (g, 1) + t
γgj Ri,t (g, 1)Ri,t (j, 1) + i,t .
g=0 g=0 j>g
LASSO estimation is applied to select relevant variables each period, described in more
detail in section B.1.2. Strategies go long the highest predicted return decile and go short
the lowest predicted return decile. The sample period covers 1968 to 2012, and all results
are based on rolling out-of-sample estimates of the models. MKT is the market return,
SMB and HML are the Fama-French factors for size and value, and UMD is the momen-
tum factor. SR is the Sharpe ratio and IR is the information ratio. T-statistics are in
parentheses, and standard errors were clustered using Newey-West’s adjustment for serial
correlation.
B.2 Transaction costs 83
strategy investigated in the aforementioned papers. The last row of the table subtracts
the approximate trading costs from the gross annual returns that we reported in table 3.1.
After adjusting for trading costs, the hedge strategy that trades the extreme portfolios has
an excess return of 24 percent per year. Trading the ninth minus the second decile (recall
that these companies are larger and therefore probably more suited to the extrapolation
from Frazzini et al. (2015)) yields an excess return of 5 percent per year. The excess return
of trading the eighth versus the third decile is insignificant and slightly negative. In other
words, the iterative conditional portfolio sort manages to profitably spread 40 percent of
the companies, even after adjusting for transaction costs.
While our strategy implementation is standard in the stock market anomalies literature,
more sophisticated variants could be designed for trading purposes when transaction costs
are taken into account. de Groot et al. (2012) suggest reducing turnover of the short-term
reversal strategy by holding onto the position in stocks even when they are not ranked in
the extreme portfolios. We do not pursue their implementation here, but, given the return
spread in the less extreme portfolios, it is plausible that such an implementation could be
constructed here as well in order to reduce turnover and trading costs further.4
4
For instance, Novy-Marx and Velikov (2016) find that many anomalies can be exploited by following
an (s,S)-type strategy that, e.g. buys stocks when they are in the highest decile but only sells them if they
drop out of the highest quintile.
Table B.4: Turnover and trading costs 84
Note that the returns in table ?? are still lower than the strategy returns in the original
tree-based conditional portfolio sort. The Fama-MacBeth regressions only include two-way
interactions.5 While the Fama-MacBeth regression with two-way interactions goes some
way to achieve similarly sized returns, the remaining differences can be attributed to the
actual return structure being more involved than can be captured by including levels and
two-way interactions of past returns alone.
Table ?? shows the Fama-MacBeth coefficient estimates averaged over the entire sample
and corresponding t-statistics for the three regression models in table ??.6 The second
column shows coefficients in the levels-only regression. We observe the short-term reversal
effect while all other past return variables enter with a positive sign. This is in line with
the standard reversal and momentum effects in the literature.
Column 3 illustrates how these results completely flip when interaction terms are intro-
duced in the regression. All level effects are on average negatively associated with expected
returns while interaction terms are positive. This result is robust to including further (less
relevant) interaction terms in column 4. A possible interpretation of this finding is that
momentum is more likely to exist when returns are more consistent. For instance, we
find that the effect of high returns in either the last month or in the second-to-last month
indicate low returns. When both returns are high, however, the interaction effect of this
consistently high return works against the reversal effect of the two individual returns.
Return consistency effects in momentum have been documented before by, among others,
Watkins (2003) and Grinblatt and Moskowitz (2004).
How do the estimated Fama-MacBeth interactions compare to the average double par-
tial derivatives in figure 3.9?7 We find both similarities and differences. When we calculate
the same average derivatives for the Fama-MacBeth model, we find that interactions of re-
turns display the aforementioned consistency effect; that is, consistently high past returns
predict high returns. These patterns coincide with the ones in figure 3.9. We also see that
in two-way interactions that involve R(0,1), returns are less sensitive to the more distant
returns, as in the top row of figure 3.9. On the other hand, in the Fama-MacBeth results,
the interactions sometimes overturn the reversal effect, unlike in the tree-based conditional
portfolio sort. Owing to their simplicity, the Fama-MacBeth regressions do not capture
the more involved interaction patterns between R(1,1) and more distant returns that are
apparent in the second row figure 3.9.
To summarize, we have emphasized the flexibility to control for variable interactions as
one of the strengths of tree-based conditional portfolio sorts before. Now we see that the
(two-way) interactions could have been discovered in a Fama-MacBeth regression frame-
work, too. The tree-based conditional portfolio sort, however, is an efficient way to screen
5
In unreported tests, we find that two-way interactions explain only around 10 percent of the variance of
the estimated expected returns of the tree-based conditional sort. This implies that much of the predictive
power of tree-based sorts comes from higher-order interactions.
6
Note that for the prediction exercise we based predictions on rolling estimates of past coefficients as
described above, while table ?? gives an average over the entire sample period.
7
Note that since we do not include higher-order polynomials of the past decile ranks, the average partial
derivatives with respect to each variable will be linear and therefore cannot capture nonlinear effects.
B.3 Robustness of the discovered structure 87
out the irrelevant interactions when the set of candidates is potentially large. At the same
time, it also allows controlling for more involved interactions.
Variable Importance
R(0,1) 1
R(1,1) 0.5
R(2,1) 0.45
R(5,1) 0.44
R(8,1) 0.41
R(4,1) 0.4
R(3,1) 0.39
R(11,1) 0.39
R(6,1) 0.38
R(7,1) 0.37
This table shows the most important re-
turn functions for the tree-based condi-
tional portfolio sorts that use all one-
month returns over the two years be-
fore portfolio formation and 86 additional
firm characteristics. Return functions are
sorted by their median importance over
forty-five years. Variable importance is
measured as described in section 3.3.2.
B.3 Robustness of the discovered structure 89
Figure B.1 shows double partial derivatives for return characteristics when firm char-
acteristics are included and corresponds to figure 3.9. In both cases, we observe patterns
that are qualitatively very similar and only differ in details — for example, the interaction
between R(1,1) and R(3,1) is somewhat more pronounced.
Overall, we conclude that the discovered structure among return characteristics is
largely unaffected by the inclusion of additional firm characteristics.
90
Figure B.1: Average double partial derivatives: Firm characteristics included. The figure shows the average prediction
when two characteristics are counterfactually varied from low to high values. The figure shows results for return functions
when 86 additional firm characteristics are included in the tree-based conditional portfolio sort. Results are based on
rolling optimization of the model and predictions are averaged over the sample period. Details are in section 3.3.2.
B. Second Appendix
B.3 Robustness of the discovered structure 91
Variable Importance
R(0,1) 1
R(1,1) 0.88
R(2,1) 0.69
R(6,1) 0.69
R(3,1) 0.66
R(4,1) 0.66
R(5,1) 0.61
R(0,2) 0.52
R(1,2) 0.45
R(1,3) 0.43
This table shows the most important re-
turn functions for a tree-based condi-
tional portfolio sort that uses past returns
functions R(g,l) with length g = 1, . . . , 18
and gaps g = 0, . . . , 6. Return functions
are sorted by their median importance
over forty-five years. Variable importance
is measured as described in section 3.3.2.
B.3 Robustness of the discovered structure 93
momentum sorts continue to work well when firms are sorted on intermediate momentum
first but the reverse is not true: Intermediate momentum sorts do not consistently give a
significant hedge return (only in low momentum stocks) or monotone returns (only in the
middle tercile of momentum stocks). Initial sorts on value or size leave the monotonicity
of return sorts intact, but interfere with the monotonicity and t-tests of gross profitability.
When firms are sorted on gross profitability first, equal-weighted hedge returns are signif-
icant for all variables in all terciles, but the returns to medium gross profitability firms is
not monotone when sorted by value.
The overall picture that emerges is that of return sorts being relatively stable while
accounting-based sorts are less robust to initial sorts on some other return- or accounting-
based variable. The results illustrate the potential relevance of correlated return- and
accounting-based characteristics, and the necessity to consider conditional returns when
the objective is to evaluate the importance of a new candidate predictor variable. Variable
interactions can also be relevant as is evident from the fact conditional sorts often work
only in some of the tercile portfolios.11
11
It is also possible to condition on more than one variable in this setting by first doubly sorting all
stocks on two variables into, say, three categories each for a total of nine portfolios. Within each portfolio,
one could then compute the same statistics as above, and discuss the effects of conditioning on levels and
interactions of variables. While, in principle, feasible for a few variables, the approach does not lend itself
to an easy interpretation in higher dimensions.
Table B.8: Conditional portfolio sorts: Average returns, t-statistics and p-values of monotonicity tests 96
where SC(g, τ ) is a split criterion function which we adopt from the related machine learn-
ing literature. The split criterion function selects the predictor variable and the associated
threshold that minimize the sum of mean squared errors in the resulting portfolios with
respect to the expected returns, that is,
X X
SC(g, τ ) = min (ri,t+1 − µ1 )2 + min (ri,t+1 − µ2 )2 (B.5)
µ1 µ2
Rit (g,1)∈S1 (g,τ ) Rit (g,1)∈S2 (g,τ )
and the inner minimizations are solved by equation (3.4). This algorithm reduces a complex
non-linear estimation problem into subsets of simpler linear ones. The problem is solved
in a brute-force fashion where the value of the split criterion function is computed for each
firm characteristic and each threshold value. The optimization is repeated in each of the
resulting portfolios until a. the number of observations in a node gets too small for further
splits, or b. no variable provides a sufficient improvement of the mean squared error in
equation (B.5). The result is a conditional portfolio sort with many levels.12
12
The question of when to stop adding new levels to the conditional sort relates to a standard bias-
variance tradeoff. Using many levels potentially results in overfitting, which would worsen the predictive
power of equation (3.5) out of sample. Estimating only a few levels might miss important aspects of the
data leading to bias. Within this sphere the number of levels can be chosen. We stop when the number
of firms in a portfolio is smaller than 100 and make sure to validate all our estimates out of samples as
described in section 3.3.2.
B.5 Greedy algorithm 99
Before we move on, we want to point out a few links to other estimation methods in
the literature. The greedy algorithm introduced in this section bears some resemblance
to forward-selection methods in regression models. Forward-selection starts out with the
smallest possible linear model, estimates bivariate regressions of the outcome variable on
each candidate regressor separately, and keeps the one with the highest t-statistic (or
some other selected performance criterion). The procedure is then repeated for all of the
remaining variables with the best-performing variable joining the regression each round
until no further variables are significant. As tree-based conditional sorts, forward-selection
works when there are more regressors than observations. On the other hand, forward
selection is global in nature in the sense that one regression function is fitted for the entire
sample and variable selection is based on performance over the entire sample. In addition,
interaction terms would need to be added one-by-one as well, leading to a large set of
candidate variables whereas the set of candidate variables is always less than the number
of main signals in tree-based conditional portfolio sorts.
Kernel regression is based on approximating an outcome variable by a (kernel-) weighted
average of the outcome at each value of the regressor. Tree-based conditional portfolio sorts
approximate the outcome by the average value of the outcome for a regressor region defined
by split points and threshold values. Kernel regressions are very flexible but do not extend
easily beyond the bivariate case. A small practical issue is the difficulty to display results
in higher dimensions. More importantly, since kernel regression is based on using local
averages, there are few observations in each subspace over which an average is taken as
the number of regressors becomes large. This is known as the curse of dimensionality and
one can show that the convergence rate for kernel regressions deteriorates sharply with the
dimensionality of the regressors. Local linear regressions run into analogous problems in
high dimensions.
100 B. Second Appendix
Bibliography
Alessi, L. and C. Detken (2014). Identifying excessive credit growth and leverage. ECB
Working Paper.
Ang, A. (2014). Asset management: A systematic approach to factor investing. Oxford
University Press.
Antweiler, W. and M. Z. Frank (2004). Is all that talk just noise? the information content
of internet stock message boards. The Journal of Finance 59 (3), 1259–1294.
Asness, C., A. Frazzini, R. Israel, and T. Moskowitz (2014). Fact, fiction and momentum
investing. Journal of Portfolio Management 40 (5), 75–92.
Asness, C. S. (1997). The interaction of value and momentum strategies. Financial Analysts
Journal 53 (2), 29–36.
Back, K. (2017). Asset pricing and portfolio choice theory. Oxford University Press.
Baker, M. and J. Wurgler (2007). Investor sentiment in the stock market. Journal of
Economic Perspectives 21 (2), 129–152.
Baker, S. R., N. Bloom, and S. J. Davis (2016). Measuring economic policy uncertainty.
The Quarterly Journal of Economics 131 (4), 1593–1636.
Bandarchuk, P. and J. Hilscher (2012). Sources of Momentum Profits: Evidence on the
Irrelevance of Characteristics. Review of Finance 17 (2), 809–845.
Barroso, P. and P. Santa-Clara (2015). Momentum has its moments. Journal of Financial
Economics 116 (1), 111–120.
Beckers, B., K. A. Kholodilin, and D. Ulbricht (2017). Reading between the lines: Using
media to improve german inflation forecasts.
Bollerslev, T., R. F. Engle, and J. M. Wooldridge (1988). A capital asset pricing model
with time-varying covariances. Journal of political Economy 96 (1), 116–131.
Breiman, L. (2001). Random forests. Machine Learning 45 (1), 5–32.
Breiman, L. (2002). Looking inside the black box. downloaded from
[Link]/users/breiman/[Link].
Breiman, L., J. Friedman, C. J. Stone, and R. A. Olshen (1984). Classification and regres-
sion trees. CRC press.
Brennan, M., T. Chordia, and A. Subrahmanyam (1998). Alternative factor specifications,
security characteristics, and the cross-section of expected stock returns. Journal of
Financial Economics 49 (3), 345 – 373.
102 BIBLIOGRAPHY
Campbell, G., W. Quinn, J. D. Turner, and Q. Ye (2018). What moved share prices in
the nineteenth-century london stock market? The Economic History Review 71 (1),
157–189.
Campbell, J. Y. (1987). Stock returns and the term structure. Journal of Financial
Economics 18 (2), 373–399.
Cerniglia, J. A., F. J. Fabozzi, and P. N. Kolm (2016). Best practices in research for
quantitative equity strategies. Journal of Portfolio Management 42 (5), 135.
Chen, J., H. Hong, and J. C. Stein (2002). Breadth of ownership and stock returns. Journal
of Financial Economics 66 (2), 171–205.
Cochrane, J. H. (2011). Presidential address: Discount rates. The Journal of Fi-
nance 66 (4), 1047–1108.
Cornell, B. (2013). What moves stock prices: Another look. Journal of Portfolio Manage-
ment 39 (3), 32.
Criminisi, A. and J. Shotten (2013). Decision Forests for Computer Vision and Medical
Image Analysis. Springer.
Cutler, D., J. Poterba, and L. Summers (1989). What moves stock prices? Journal of
Portfolio Management 15 (2).
Da, Z., J. Engelberg, and P. Gao (2014). The sum of all fears investor sentiment and asset
prices. The Review of Financial Studies 28 (1), 1–32.
Daniel, K. and T. J. Moskowitz (2015). Momentum crashes. Unpublished manuscript.
Daniel, K. and S. Titman (1997). Evidence on the characteristics of cross sectional variation
in stock returns. The Journal of Finance 52 (1), 1–33.
Das, S. R. et al. (2014). Text and context: Language analytics in finance. Foundations
and Trends R in Finance 8 (3), 145–261.
de Bondt, W. and R. Thaler (1985). Does the stock market overreact? The Journal of
Finance 40 (3), 793–805.
de Groot, W., J. Huij, and W. Zhou (2012). Another look at trading costs and short-term
reversal profits. Journal of Banking & Finance 36 (2), 371–382.
Dimson, E., P. Marsh, and M. Staunton (2009). Triumph of the optimists: 101 years of
global investment returns. Princeton University Press.
Doeswijk, R., T. Lam, and L. Swinkels (2014). The global multi-asset market portfolio,
1959–2012. Financial Analysts Journal 70 (2), 26–41.
Doms, M. E. and N. J. Morin (2004). Consumer sentiment, the economy, and the news
media.
Duttagupta, R. and P. Cashin (2011). Anatomy of banking crises in developing and emerg-
ing market countries. Journal of International Money and Finance 30 (2), 354–376.
Einav, L. and J. Levin (2014). Economics in the age of big data. Science 346 (6210),
1243089.
Fair, R. C. (2002). Events that shook the market. The Journal of Business 75 (4), 713–731.
Fama, E. and K. French (1992). The cross-section of expected stock returns. The Journal
BIBLIOGRAPHY 103
Harvey, C. R., Y. Liu, and H. Zhu (2016). ... and the cross-section of expected returns.
Review of Financial Studies 29 (1), 5–68.
Hastie, T., R. Tibshirani, and J. Friedman (2009). The Elements of Statistical Learning.
Springer New York Inc.
Haugen, R. A. and N. L. Baker (1996, July). Commonality in the determinants of expected
stock returns. Journal of Financial Economics 41 (3), 401–439.
Heston, S. L. and R. Sadka (2008). Seasonality in the cross-section of stock returns. Journal
of Financial Economics 87 (2), 418–445.
Ho, T. K. (1998). The random subspace method for constructing decision forests. Pattern
Analysis and Machine Intelligence, IEEE 20 (8), 832–844.
Huerta, R., F. Corbacho, and C. Elkan (2013). Nonlinear support vector machines can
systematically identify stocks with high and low future returns. Algorithmic Finance 2,
45–58.
Hyafil, L. and R. L. Rivest (1976). Constructing optimal binary decision trees is np-
complete. Information Processing Letters 5 (1), 15 – 17.
Ibbotson, R., R. J. Grabowski, J. P. Harrington, and C. Nunes (2016). 2016 Stocks, Bonds,
Bills, and Inflation (SBBI) Yearbook. John Wiley & Sons.
Ilmanen, A. (2011). Expected returns: An investor’s guide to harvesting market rewards.
John Wiley & Sons.
Jegadeesh, N. (1990). Evidence of predictable behavior of security returns. The Journal
of Finance 45 (3), 881–898.
Jegadeesh, N. and S. Titman (1993). Returns to buying winners and selling losers: Impli-
cations for stock market efficiency. The Journal of Finance 48 (1), 65–91.
Jordà, Ò., K. Knoll, D. Kuvshinov, M. Schularick, and A. M. Taylor (2017). The rate
of return on everything, 1870–2015. Technical report, National Bureau of Economic
Research.
Kahneman, D. and A. Tversky (1979). Prospect theory: An analysis of decisions under
risk. Econometrica, 263–291.
Kaminsky, G. L. (2006). Currency crises: Are they all the same? Journal of International
Money and Finance 25 (3), 503–527.
Kearney, C. and S. Liu (2014). Textual sentiment in finance: A survey of methods and
models. International Review of Financial Analysis 33, 171–185.
Keim, D. B. and A. Madhavan (1997). Transactions costs and investment style: an inter-
exchange analysis of institutional equity trades. Journal of Financial Economics 46 (3),
265 – 292.
Kleinberg, E. (1990). Stochastic discrimination. Annals of Mathematics and Artificial
intelligence 1 (1), 207–239.
Kleinberg, E. (1996). An overtraining-resistant stochastic modeling method for pattern
recognition. The Annals of Statistics 24 (6), 2319–2349.
Kogan, L. and M. Tian (2015). Firm characteristics and empirical factor models: a data-
BIBLIOGRAPHY 105
Hiermit erkläre ich an Eides statt, dass die Dissertation von mir selbstständig, ohne uner-
laubte Beihilfe angefertigt ist.