Quant Research Compilation
Quant Research Compilation
Alpha Models & Factor Investing | Option Pricing & Volatility Modeling
This document presents a curated collection of the most influential and cutting-edge quantitative research
papers across two foundational domains of modern quantitative finance: alpha models (encompassing
factor investing, empirical asset pricing, and machine-learning-driven return prediction) and option
pricing (spanning from the classical Black-Scholes framework to stochastic volatility, local volatility, and
deep-learning-based approaches). Each entry includes the paper’s full citation, a concise abstract, key
contributions, and practical significance for researchers and practitioners.
The CAPM, developed independently by William Sharpe and John Lintner, provides the foundational
relationship between systematic risk and expected return. The model posits that the expected return on
any asset equals the risk-free rate plus a risk premium proportional to the asset’s beta — its covariance
with the market portfolio divided by the market variance. The formula 𝐸[𝑅𝑖 ] = 𝑅𝑓 + 𝛽𝑖 (𝐸[𝑅𝑚 ] − 𝑅𝑓 )
became the cornerstone of modern portfolio theory, earning Sharpe the Nobel Prize in Economics in 1990.
The CAPM’s elegance lies in its parsimony: only market beta should matter for pricing. However, em-
pirical tests beginning in the 1970s revealed persistent anomalies. Basu (1977) documented that price-
to-earnings ratios predicted returns beyond beta; Banz (1981) discovered the size effect where small
firms outperformed large firms; and numerous other anomalies accumulated. These findings collectively
challenged the CAPM’s sufficiency and set the stage for multi-factor models. Despite its empirical short-
comings, the CAPM remains essential because it establishes the conceptual framework — risk premium
as compensation for bearing systematic, non-diversifiable risk — that all subsequent models extend.
“The Cross-Section of Expected Stock Returns” (1992) and “Common Risk Factors in the
Returns on Stocks and Bonds” (1993) represent arguably the most consequential empirical papers in
asset pricing since the CAPM. Eugene Fama and Kenneth French demonstrated that two firm character-
istics — size (market capitalization) and value (book-to-market ratio) — contain significant explanatory
power for the cross-section of average stock returns that market beta alone cannot capture.
Generated by [Link]
The three-factor model takes the form: 𝑅𝑖𝑡 − 𝑅𝑓𝑡 = 𝛼𝑖 + 𝛽𝑖,𝑀𝐾𝑇 (𝑅𝑀𝑡 − 𝑅𝑓𝑡 ) + 𝛽𝑖,𝑆𝑀𝐵 𝑆𝑀 𝐵𝑡 +
𝛽𝑖,𝐻𝑀𝐿 𝐻𝑀 𝐿𝑡 + 𝜖𝑖𝑡 , where SMB (Small Minus Big) captures the size premium and HML (High Minus
Low book-to-market) captures the value premium. Fama and French found that combining these factors
explains up to 95% of the returns of diversified portfolios. The model’s success lies in its ability to
account for the empirical regularities that bedeviled the CAPM: small-cap stocks earn higher returns
than predicted by their betas, and value stocks systematically outperform growth stocks. The model
has been extensively tested internationally, with the size and value effects documented across dozens of
markets worldwide. [12 ][17 ][18 ]
“Returns to Buying Winners and Selling Losers: Implications for Stock Market Efficiency”
introduced the momentum factor — the phenomenon that stocks with high returns over the past 3–12
months continue to outperform, while past losers continue to underperform. The authors documented that
a strategy of buying winner stocks and selling loser stocks generated risk-adjusted returns of approximately
1.5% per month.
Momentum is particularly significant because it is among the most persistent and pervasive anomalies
in finance. It has been replicated across international equity markets, asset classes (bonds, currencies,
commodities), and time periods. The strategy’s profitability challenges both the CAPM and the efficient
market hypothesis, as past price information — which should be fully incorporated — continues to predict
future returns. Subsequent work by Hong and Stein (1999) proposed a behavioral model where
gradual information diffusion creates momentum, while Asness (1997) linked momentum to value as
complementary factors. The momentum effect remains one of the strongest empirical regularities in asset
pricing, though it is subject to periodic crashes (Daniel and Moskowitz, 2014). [53 ][55 ][59 ]
Mark Carhart’s “On Persistence in Mutual Fund Performance” extended the Fama-French
three-factor model by adding momentum (UMD: Up Minus Down) as a fourth factor. The model
specification is: 𝑅𝑖𝑡 − 𝑅𝑓𝑡 = 𝛽𝑖 (𝑅𝑀𝑡 − 𝑅𝑓𝑡 ) + 𝑠𝑖 𝑆𝑀 𝐵𝑡 + ℎ𝑖 𝐻𝑀 𝐿𝑡 + 𝑢𝑖 𝑈 𝑀 𝐷𝑡 + 𝜖𝑖𝑡 . The momentum
factor is constructed from stocks sorted on their one-year past returns, skipping the most recent month
to avoid short-term reversal effects.
Carhart’s four-factor model became the industry standard for mutual fund performance evaluation. The
model’s practical significance is enormous: it demonstrated that much of what appeared to be manager
skill (alpha) in earlier studies was actually exposure to the momentum factor. The model’s explanatory
power is substantially higher than the three-factor framework, particularly for portfolios sorted on past
returns. The four-factor model remains widely used in both academic research and industry practice for
performance attribution and risk modeling. [14 ][19 ][22 ]
Generated by [Link]
2.2 Novy-Marx (2013) — The Gross Profitability Premium
Robert Novy-Marx’s “The Other Side of Value: The Gross Profitability Premium” demon-
strated that gross profitability (revenue minus cost of goods sold, scaled by total assets) has roughly
the same power as book-to-market in predicting the cross-section of average returns. This finding was
revolutionary because it showed that the “other side” of the value-growth dimension — profitability —
carries independent predictive power.
Novy-Marx documented a spread of 0.31% per month (3.78% annually) between the most and
least profitable firms. Crucially, profitable firms generate higher returns despite having higher valuation
ratios, which is difficult to reconcile with standard risk-based explanations. The paper established that
controlling for profitability dramatically improves value strategies, especially among large, liquid stocks.
The profitability factor correlates negatively with the value factor (approximately -0.50), making it a
valuable diversifier. This research directly influenced the construction of the Fama-French five-factor
model and the Hou-Xue-Zhang q-factor model, both of which incorporate profitability as a core dimension.
[64 ][68 ][71 ]
In “A Five-Factor Asset Pricing Model,” Fama and French expanded their three-factor framework
by adding two quality factors: profitability (RMW: Robust Minus Weak) and investment
(CMA: Conservative Minus Aggressive). The profitability factor captures the tendency of highly
profitable firms to outperform, while the investment factor reflects the empirical regularity that firms
with high asset growth tend to earn lower subsequent returns.
The five-factor model’s striking finding was that the value factor (HML) becomes redundant
once profitability and investment factors are included. Fama and French argued that value may be a
“noisy proxy” for expected returns because market capitalization is sensitive to forecasts of earnings and
investment. The model’s theoretical motivation draws on the dividend discount model and Miller-
Modigliani (1961) propositions, deriving three conclusions: higher book-to-market implies higher ex-
pected return; higher expected earnings imply higher expected return; and higher expected growth in
book equity implies lower expected return. The model has been tested internationally (Fama & French,
2017), though results vary across regions, with the value factor remaining significant in most non-US
markets. [13 ][45 ][47 ][51 ]
The q-factor model represents a theoretical departure from the Fama-French tradition. Motivated
by the investment CAPM (Zhang, 2017), it prices risky assets from the perspective of their suppliers
(firms) rather than their buyers (investors). The model specifies that expected excess returns are explained
by four factors: market (MKT), size (ME), investment (I/A), and profitability (ROE).
The q-factor model’s empirical performance is striking: it fully subsumes the Fama-French six-factor
model in head-to-head factor spanning tests. The model’s construction uses 2×3×3 sorts on market
equity, investment-to-assets, and return on equity, creating 18 benchmark portfolios. The investment
CAPM’s core intuition is that high investment relative to low expected profitability implies low cost of
capital, while low investment relative to high expected profitability implies high cost of capital. This
supply-side perspective validates Graham and Dodd’s (1934) security analysis on equilibrium grounds
Generated by [Link]
within efficient markets. The model has been extended to the q^5 model (Hou, Mo, Xue & Zhang,
2020) by adding an expected growth factor that earns an average premium of 0.84% per month (t
= 10.27). [16 ][20 ][24 ]
“Mispricing Factors” (with Robert Stambaugh and Yu Yuan) took a different approach by construct-
ing factors designed to capture market mispricing rather than risk. The two mispricing factors —
management (MGMT) and performance (PERF) — are composite scores based on 11 anomaly
variables each. The management factor combines measures of net stock issuance, composite equity is-
suance, accruals, net operating assets, investment-to-assets, and changes in gross property, plant, and
equipment. The performance factor combines price momentum, earnings surprise, return on assets, asset
turnover, gross profitability, and the book-to-market ratio.
The four-factor model (market, size, MGMT, PERF) performs competitively against the Fama-French
five-factor model. The paper’s framing is behavioral: factors are interpreted as capturing systematic
mispricing rather than compensation for bearing systematic risk. This model, alongside the DHS model
(Daniel, Hirshleifer & Sun, 2018) with its financing and post-earnings-announcement-drift factors,
represents the behavioral finance challenge to risk-based factor models. [74 ]
“Choosing Factors” addressed the proliferation of factor models by proposing the maximum squared
Sharpe ratio (Sh²(f)) as a unified metric for ranking asset pricing models. The paper compared
nested models (CAPM, 3-factor, 5-factor, 6-factor) and examined non-nested variants including cash
profitability versus operating profitability, long-short spread factors versus excess return factors,
and factors constructed from small versus big stocks.
The six-factor model — adding momentum (MOM) to the five-factor framework — emerged as the
winner in nested comparisons. For non-nested choices, cash profitability outperformed operating
profitability as the variable for constructing profitability factors. The paper’s methodological con-
tribution is significant: by using the max squared Sharpe ratio rather than GRS F-tests, it enables
comparison of non-nested models that previously lacked a common evaluation metric. This framework
has been adopted in subsequent international tests (Oulu, 2020) confirming that the six-factor model
outperforms its predecessors across North America, Europe, Asia Pacific, and Japan. [74 ][78 ][79 ][75 ]
“…and the Cross-Section of Expected Returns” documented the explosive growth of claimed factors
in academic finance. The authors catalogued 316 factors published in top journals between 1967 and
2014, with the rate of discovery accelerating from a handful per year in the 1970s to over 40 per year
by the 2010s.
The paper’s central argument is the multiple testing problem: when hundreds of researchers test
hundreds of potential factors on overlapping datasets, many “discoveries” will appear significant purely
Generated by [Link]
by chance. The standard t-statistic threshold of 2.0 (corresponding to a 5% significance level) is far too
lenient given the volume of testing. After applying multiple testing corrections, the authors concluded
that by 2012, a new factor needed a t-statistic of at least 3.0 to be credible, and by 2032 the threshold
would need to be even higher. This paper fundamentally changed the standards for factor discovery in
empirical asset pricing, raising the bar for what constitutes a genuine anomaly. [46 ]
3.2 Gu, Kelly & Xiu (2020) — Empirical Asset Pricing via Machine Learning
“Empirical Asset Pricing via Machine Learning” (published in Review of Financial Studies) repre-
sents the landmark study that introduced modern machine learning to the cross-section of stock returns.
The authors compared tree-based methods (random forests, gradient boosted trees), neural
networks, and linear models on their ability to predict monthly stock returns using a zoo of 920
firm-level characteristics.
The key findings were transformative: neural networks and gradient boosted trees substantially
outperformed linear models in predicting the cross-section of returns. The tree-based methods’
success derives from their ability to capture nonlinear interactions between characteristics — for instance,
the profitability effect may be much stronger among small firms than large firms, a pattern that linear
models miss. The paper demonstrated that machine learning can distill the “factor zoo” into predictive
signals while handling the high-dimensional, noisy nature of financial data. This work has spawned an
entire subfield of ML-based asset pricing research, including subsequent papers by the same authors on
autoencoder asset pricing models (Gu, Kelly & Xiu, 2021) and the influential survey “Factor
Models, Machine Learning, and Asset Pricing” (Giglio, Kelly & Xiu, 2022). [50 ][52 ][30 ]
The most recent frontier involves large language models (LLMs) and agent-based systems for quanti-
tative investing. A comprehensive survey (2025) documents the evolution from traditional deep learning
to LLM-powered quant systems. Key developments include:
Alpha-GPT (Wang et al., 2023) introduced a human-AI interactive framework for factor mining, where
LLMs propose factor ideas, refine them based on human feedback, and generate executable code. Alpha-
GPT 2.0 (Yuan et al., 2024) automates the entire pipeline from alpha mining to modeling and analysis.
FinAgent (Zhang et al., 2024) integrates multimodal data (numerical, textual, visual) with a dual-level
reflection module, significantly outperforming state-of-the-art baselines across six financial datasets.
Multi-agent systems represent the cutting edge: TradingAgents (Xiao et al., 2024) simulates collabora-
tive trading desk dynamics with specialized agents (fundamental analysts, sentiment analysts, technical
analysts, traders) who debate and reach consensus recommendations. FINCON (Yu et al., 2025) achieves
a cumulative return of 82.87% for single-stock trading and 113.84% for portfolio management with a
Sharpe ratio of 3.269. These developments suggest a paradigm shift from statistical factor models to
AI systems capable of autonomous factor discovery, multi-source data integration, and dynamic strategy
adaptation. [25 ]
Generated by [Link]
Part II: Option Pricing and Volatility Modeling
The pricing of derivative securities — particularly options — has been transformed by the interplay of
stochastic calculus, numerical methods, and more recently, machine learning. This section traces the
field from the seminal Black-Scholes framework through increasingly sophisticated volatility models to
modern deep-learning approaches.
“The Pricing of Options and Corporate Liabilities” (Black & Scholes, 1973) and “Theory of
Rational Option Pricing” (Merton, 1973) revolutionized finance by providing the first closed-form
solution for European option pricing. The Black-Scholes formula for a call option is: 𝐶 = 𝑆𝑁 (𝑑1 ) −
2 √
𝐾𝑒−𝑟𝑇 𝑁 (𝑑2 ), where 𝑑1 = ln(𝑆/𝐾)+(𝑟+𝜎
√
𝜎 𝑇
/2)𝑇
and 𝑑2 = 𝑑1 − 𝜎 𝑇 .
The model’s key insight was that options can be perfectly replicated by dynamically trading the un-
derlying asset and risk-free bonds — the delta-hedging argument. This implies that option prices
are independent of investors’ risk preferences, allowing pricing under the risk-neutral measure. The
assumptions include constant volatility, geometric Brownian motion for the underlying, no dividends, no
transaction costs, and European exercise. Scholes and Merton were awarded the Nobel Prize in Eco-
nomics in 1997 (Black had passed away in 1995). Despite its simplifying assumptions, the Black-Scholes
model remains the benchmark against which all subsequent models are measured, and its conceptual
framework underpins the entire derivatives industry. [29 ][31 ]
The binomial model’s significance is threefold: it is intuitive and requires no advanced stochastic calculus;
it can price American options (which the original Black-Scholes formula cannot) by checking for early
exercise at each node; and it converges to Black-Scholes as the number of time steps approaches
√ √
infinity. The CRR parameterization sets 𝑢 = 𝑒𝜎 Δ𝑡
and 𝑑 = 𝑒−𝜎 Δ𝑡
, ensuring convergence. The model
remains essential for teaching and for pricing path-dependent and American-style derivatives where closed-
form solutions are unavailable. [54 ][57 ][58 ]
The jump diffusion model addresses a critical limitation of Black-Scholes: the inability to explain the
volatility smile — the empirical pattern where implied volatilities vary with strike price. Jumps generate
Generated by [Link]
fatter tails in the return distribution, producing higher implied volatilities for out-of-the-money options
and creating the smile pattern. The model is particularly important for pricing options near market
events (earnings announcements, economic releases) where discrete jumps are more likely. The Bates
(1996) model later combined Merton jumps with Heston stochastic volatility for even greater flexibility.
[35 ][44 ]
The Hull-White stochastic volatility model was among the first to relax the Black-Scholes assump-
tion of constant volatility. In their framework, the asset price and its volatility are driven by separate but
correlated Brownian motions. The model demonstrated that stochastic volatility produces skewed and
leptokurtic return distributions, more consistent with empirical observations than the lognormal
distribution assumed by Black-Scholes.
Hull-White’s contribution was showing that when volatility is uncorrelated with the asset price, options
can still be priced by integrating the Black-Scholes formula over the distribution of average variance.
When correlation is introduced (the “leverage effect”), the model generates asymmetric volatility smiles
— a pattern consistently observed in equity markets where implied volatility increases as strikes decrease.
This work established the theoretical foundation for all subsequent stochastic volatility models. [36 ]
Steven Heston’s “A Closed-Form Solution for Options with Stochastic Volatility with Ap-
plications to Bond and Currency Options” is the most widely cited stochastic volatility model in
practice. The Heston dynamics under the risk-neutral measure are:
The model’s defining feature is the mean-reverting Cox-Ingersoll-Ross (CIR) process for variance
𝑉𝑡 , with parameters: 𝜅 (mean reversion speed), 𝜃 (long-term variance), 𝜂 (volatility of variance), and 𝜌
(correlation between price and volatility shocks). The correlation 𝜌 is typically negative, capturing the
leverage effect where volatility rises as prices fall. Heston’s breakthrough was deriving a semi-closed-
form solution using Fourier inversion of the characteristic function, enabling fast calibration to market
prices. The model captures the volatility smile and is widely used for pricing equity, FX, and interest
rate derivatives. Limitations include the Feller condition (2𝜅𝜃 > 𝜂2 ) required to keep variance positive
and the model’s inability to perfectly fit short-dated smiles. [26 ][28 ][33 ]
Generated by [Link]
5.3 Bates (1996) — Stochastic Volatility with Jumps
David Bates combined the Heston stochastic volatility framework with Merton-style jumps in the
asset price, creating the SVJ (Stochastic Volatility + Jumps) model. The dynamics add a compound
Poisson process to the Heston asset price equation: 𝑑𝑆𝑡 = 𝜇𝑆𝑡 𝑑𝑡 + √𝑉𝑡 𝑆𝑡 𝑑𝑊𝑡𝑆 + (𝐽𝑡 − 1)𝑆𝑡− 𝑑𝑁𝑡 .
The Bates model addresses two key empirical features that Heston alone cannot capture: short-dated
smiles (which require jumps, as stochastic volatility operates too slowly) and the steepness of the smile
at short maturities. By combining both features, the model provides a more complete description of the
volatility surface across all maturities. The model retains the Fourier-based pricing approach, making
it computationally tractable. Extensions include the SVJJ model (Eraker, 2004) which adds jumps
to the volatility process itself, and the Barndorff-Nielsen-Shephard (BNS) model which uses a
non-Gaussian OU process for volatility. [35 ][44 ]
Bruno Dupire’s local volatility model represents a fundamentally different approach: instead of making
volatility stochastic, it makes volatility a deterministic function of the underlying asset price and
time: 𝜎𝑙𝑜𝑐 (𝑆𝑡 , 𝑡). Dupire derived the famous inversion formula that extracts the local volatility surface
directly from market option prices:
𝜕𝐶
+ 𝑟𝐾 𝜕𝐶
𝜎𝑙𝑜𝑐 (𝐾, 𝑇 ) = √ 𝜕𝑇 𝐾 2 𝜕 2 𝐶𝜕𝐾
2 𝜕𝐾 2
The local volatility model’s key property is that it perfectly fits all European option prices in
the cross-section — by construction, if the local volatility function is chosen appropriately, the model
reproduces the entire implied volatility surface. This makes it invaluable for pricing exotic options
consistently with vanilla options. The model also satisfies the Markovian projection theorem (Gyöngy,
1986): any multivariate model shares some local volatility model with the same marginal distributions for
the underlying, meaning local volatility does not lose pricing flexibility for European options. Practical
challenges include numerical instability when computing derivatives of noisy market data and the fact
that local volatility predicts unrealistic future volatility smile dynamics. [36 ][40 ][43 ]
The Stochastic Alpha Beta Rho (SABR) model, introduced by Patrick Hagan, Deep Kumar,
Andrew Lesniewski, and Diana Woodward, has become the industry standard for interest rate
derivatives pricing. The model dynamics are:
where 𝑑𝑊𝑡1 𝑑𝑊𝑡2 = 𝜌𝑑𝑡, 𝐹𝑡 is the forward rate, 𝛼𝑡 is the stochastic volatility, 𝛽 controls the backbone
(from normal at 𝛽 = 0 to lognormal at 𝛽 = 1), 𝜈 is the volatility of volatility, and 𝜌 is the correlation.
Generated by [Link]
The SABR model’s breakthrough was Hagan’s asymptotic expansion, which provides an approximate
closed-form formula for implied volatility as a function of strike. This formula enables fast calibration to
market smiles and direct computation of Greeks. The model is parsimonious (four parameters), intuitive,
and produces realistic smile dynamics — when the underlying moves, the smile shifts in the same direction.
The SABR model dominates in interest rate markets for pricing swaptions, caps, and floors. Extensions
include the shifted-SABR (for negative rates), normal SABR, and free-boundary SABR models
developed in response to post-2008 negative interest rate environments. [37 ][38 ][41 ][42 ]
Jim Gatheral’s work on the SVI (Stochastic Volatility Inspired) parameterization provided a parsi-
2
monious functional form for the implied volatility surface: 𝜎𝐵𝑆 (𝑘, 𝑇 ) = 𝑎+𝑏{𝜌(𝑘−𝑚)+√(𝑘 − 𝑚)2 + 𝜎2 },
where 𝑘 = ln(𝐾/𝐹 ) is log-moneyness. The five SVI parameters (𝑎, 𝑏, 𝜌, 𝑚, 𝜎) control the level, skew, cur-
vature, and asymptotic behavior of the smile.
Gatheral demonstrated that SVI fits market smiles with remarkable accuracy while ensuring absence
of static arbitrage when calibrated appropriately. His “Smile Dynamics” series of papers (Bergomi
& Gatheral, 2004–2015) introduced the Skew Stickiness Ratio (SSR) — the ratio of the change in
at-the-money-forward skew to the change in at-the-money-forward volatility when the spot moves. This
ratio distinguishes between “type I” models (like Heston) where the SSR tends to 1 and “type II” models
(like the Bergomi model) where the SSR stabilizes around 1.5–2.0, matching empirical equity market
behavior. The SSR framework has become essential for understanding the joint dynamics of spot and
implied volatilities and for arbitraging differences between implied and realized smile dynamics. [39 ][61 ][62 ]
“Pricing Options and Computing Implied Volatilities using Neural Networks” demonstrated
that neural networks can learn the mapping from option inputs (strike, maturity, underlying price, risk-
free rate, volatility) to option prices. The authors showed that a multilayer perceptron (MLP) trained
on synthetic Black-Scholes prices could generalize to pricing options under the Heston model with high
accuracy.
The key insight was that neural networks can act as universal function approximators (Hornik et al.,
1989), learning the complex nonlinear relationship between model parameters and option prices. This
approach is particularly valuable for models lacking closed-form solutions, where traditional methods
require computationally expensive numerical integration or Monte Carlo simulation. The neural network,
once trained, prices options in a single forward pass — orders of magnitude faster than traditional
methods. [73 ]
Neural Stochastic Differential Equations (NSDEs) represent a paradigm shift: instead of specifying
parametric models for the drift and diffusion terms, these approaches parameterize them with neural
networks. The general form is: 𝑑𝑋𝑡 = 𝑓𝜃 (𝑋𝑡 , 𝑡)𝑑𝑡 + 𝑔𝜙 (𝑋𝑡 , 𝑡)𝑑𝑊𝑡 , where 𝑓𝜃 and 𝑔𝜙 are neural networks.
Generated by [Link]
The calibration problem becomes a simulation optimization problem: neural network parameters are
adjusted to minimize the difference between model-predicted prices (computed via Monte Carlo simulation
through the neural SDE) and market prices. This approach combines the flexibility of nonparametric
models with the economic structure of SDEs. The resulting models can capture complex market
dynamics that parametric models miss, while maintaining the no-arbitrage guarantees of continuous-
time finance. Challenges include training stability and the computational cost of Monte Carlo simulation
during training. [70 ]
“Deep Learning Option Pricing with Market Implied Volatility Surfaces” introduced a vari-
ational autoencoder (VAE) framework that compresses high-dimensional volatility surfaces into a
low-dimensional latent space. Using S&P 500 options data (2018–2023), the authors construct arbitrage-
free volatility surfaces on a 41×20 grid and compress them into 10 latent dimensions via a VAE with
convolutional encoder/decoder architecture.
A multilayer perceptron then maps these latent variables plus option-specific inputs (strike, maturity) to
prices for American puts and arithmetic Asian options. The staged training (VAE pre-training
→ pricing network → end-to-end fine-tuning) achieves high accuracy with prediction errors concentrated
near long maturities and ATM strikes — precisely where bid-ask spreads are largest. The method requires
only a single neural network forward pass, supports GPU parallelization, and naturally improves
with additional data. This represents the state of the art in combining deep learning with the full
information content of the volatility surface. [72 ]
“Option Pricing with Deep Learning: A Long Short-Term Memory Approach” demonstrated
that LSTM networks outperform both the Black-Scholes and Heston models for options with maturities
greater than 3 months. The LSTM’s time-sequencing nature enables it to infer short-term volatil-
ity from the past five trading days of S&P 500 prices, reducing dependence on explicit volatility
estimates.
SHAP interpretability analysis revealed that strike price and underlying value are the most important
features, with volatility playing a surprisingly limited role within the LSTM framework — the model
learns pricing dynamics from historical sequences rather than explicit volatility inputs. The LSTM model
outperformed MLP benchmarks, particularly during periods of market stress, because it is less exposed
to volatility measurement errors. This finding has important practical implications: LSTM-based
pricers may be more robust than traditional models during rapidly changing market conditions when
volatility estimates are most unreliable. [81 ]
Generated by [Link]
Paper Authors Year # Factors Key Innovation Factor Definitions
Black- Black, Scholes, 1973 Constant vol Delta hedging, Yes [29 ]
Scholes Merton risk-neutral pricing
Jump Merton 1976 Jump process Poisson jumps in Yes [35 ]
Diffusion asset price
CRR Cox, Ross, 1979 Discrete time American option N/A [54 ]
Binomial Rubinstein pricing
Hull-White Hull, White 1987 Stochastic vol Uncorrelated vol Approx. [36 ]
SV process
Heston Heston 1993 Stochastic vol Mean-reverting CIR Yes (Fourier) [26 ]
Model variance
Local Dupire 1994 Local vol 𝜎(𝑆, 𝑡) from market Formula [40 ]
Volatility prices
Bates Bates 1996 SV + Jumps Combines Heston + Yes (Fourier) [35 ]
Model Merton jumps
SABR Hagan et al. 2002 Stochastic vol Asymptotic implied Approx. [37 ]
Model vol formula
Generated by [Link]
Table 2 – continued
Paper Authors Year Model Class Key Feature Closed Form?
CAPM — — — — — 60–70%
FF (SMB) (HML) — — — 90–95% [12 ]
3-Factor
Carhart — — (UMD) 92–96% [14 ]
4-Factor
FF (redun- (RMW) (CMA) — 93–97% [51 ]
5-Factor dant)
q-Factor (ME) — (ROE) (I/A) — 94–98% [16 ]
FF (MOM) 95–98% [78 ]
6-Factor
Mispricing — — — —(in PERF) 92–96% [74 ]
4-Factor
The evolution from single-factor to machine-learning-based models reveals several important themes.
First, the factor zoo critique (Harvey, Liu & Zhu, 2016) has fundamentally raised the evidentiary
bar for new factor discovery — researchers must now demonstrate t-statistics above 3.0 and account
for multiple testing. Second, the competition between risk-based and behavioral explanations
remains unresolved: the q-factor model provides an elegant risk-based foundation for profitability and
investment effects, while mispricing models capture similar patterns through behavioral channels. Third,
machine learning has shifted the frontier from hand-crafted factors to learned representations: Gu,
Kelly & Xiu (2020) showed that neural networks and tree-based methods substantially outperform linear
models, and the latest LLM-based approaches (Alpha-GPT, FinAgent, TradingAgents) suggest a future
where AI systems autonomously discover and exploit predictive signals.
Generated by [Link]
9.2 For Option Pricing
The option pricing landscape has undergone a similar transformation from analytical tractability to com-
putational flexibility. The Black-Scholes model remains the conceptual foundation, but practitioners
rely on Heston for equity derivatives, SABR for interest rate derivatives, and local volatility for ex-
otic option pricing. The emergence of deep learning approaches (neural SDEs, VAE-based pricers,
LSTM models) offers a new paradigm: models that learn from data rather than being specified para-
metrically. These approaches are particularly valuable for exotic options and complex derivatives where
traditional models are computationally prohibitive. The trade-off is interpretability — while a Heston
model’s parameters have clear economic meanings (mean reversion speed, long-term variance), neural
network parameters do not, posing challenges for risk management and regulatory compliance.
This compilation was prepared to serve as a comprehensive reference for quantitative researchers, portfolio
managers, derivatives traders, and graduate students in financial economics. Each paper represents a
significant contribution to the understanding of asset pricing and derivative valuation, and collectively
they trace the intellectual evolution of two of the most active frontiers in quantitative finance.
Generated by [Link]