Évaluer les primes factorielles, la diversification et les crises
Résumé
Ce notebook évalue les rendements factoriels publiés comme éléments potentiels de portefeuille. Il examine trois questions : le rendement moyen d’un facteur dépasse-t-il un seuil statistique exigeant, les facteurs se diversifient-ils mutuellement pour des raisons susceptibles de perdurer, et comment se sont-ils comportés pendant certaines périodes de crise ? Il s’appuie sur les longues séries de AQR et de Fama-French, corrige les statistiques de rendement moyen pour tenir compte de l’autocorrélation et compare les seuils de significativité conventionnels à un seuil plus élevé, destiné à tenir compte de la recherche extensive de facteurs.
L’analyse examine aussi les corrélations entre classes d’actifs, les rendements pendant les périodes de crise et les performances avant et après la publication des facteurs. Les éléments présentés sont descriptifs : les séries de recherche dans leur version actuelle peuvent être révisées ou complétées rétrospectivement, les historiques diffèrent selon les facteurs et les rendements sont bruts des coûts de financement, de rotation, de vente à découvert et des contraintes de capacité. Les périodes de crise et les découpages autour des dates de publication ont été choisis a posteriori ; ils ne peuvent donc établir une protection future ni prouver que la publication a causé l’affaiblissement ultérieur des rendements.
Idées clés
- La sélection des facteurs doit tenir compte des nombreux candidats testés dans la littérature.
- La diversification dépend des liens entre les sources de rendement, et pas seulement du fait que chacune ait généré une prime.
- L’autocorrélation doit être prise en compte pour évaluer la robustesse statistique des rendements moyens factoriels.
- Les gains historiques observés pendant certaines périodes de crise décrivent des épisodes choisis et n’établissent pas une protection future.
- Une performance plus faible après publication documente une évolution, mais n’en identifie pas la cause.
Étiquettes
Texte intégral
# Which return sources actually diversify each other
# Which return sources actually diversify each other
**Docker image**: `ml4t`
## Purpose
An allocator needs things to allocate between, and the useful question about a candidate is not
whether it earned a premium but whether it earns one when the others do not. This notebook works
through the published evidence on that, using close to a century of factor returns from AQR and
Kenneth French's data library.
Three things get separated that are easy to run together. Whether a factor's premium is real, on
a bar appropriate to a literature that has tested hundreds of candidates. Whether two factors
diversify each other, and whether there is a reason for it beyond the sample. And whether a
factor pays when a portfolio most needs it to, which is a question the available evidence can
describe and cannot test.
## Learning objectives
- Compute a serial-correlation-corrected t-statistic for a factor's mean return, and judge it
against a threshold appropriate to the number of candidates the literature has tried.
- Measure the correlation between two factors across asset classes, and separate a correlation
that has a structural explanation from one that does not.
- Read a crisis-window table without treating a set of windows chosen after the fact as a test
of anything.
- Split a factor's history at its publication date and say precisely what a weaker second half
does and does not establish.
## Book reference
Chapter 17, Section 17.4 (Defining baseline allocators).
## Prerequisites
- None. Everything here comes from AQR and Fama-French published series, not from the case
studies.
## Data Requirements
### Fama-French Data (automatic)
Data is fetched automatically from Kenneth French's Data Library on first use.
### AQR Data (one-time download required)
AQR factor data must be downloaded once before running this notebook:
```python
from ml4t.data.providers import AQRFactorProvider
AQRFactorProvider.download() # Downloads ~50MB of Excel files from AQR
```
Both data sources are freely available for academic and research use.
## Setup
```python
"""Factor Allocation Evidence - assess diversification across published factor returns."""
import importlib
import logging
import warnings
from datetime import datetime
import numpy as np
import pandas as pd
import plotly.graph_objects as go
import polars as pl
import structlog
from plotly.subplots import make_subplots
from scipy import stats
from utils import DATA_DIR
from utils.style import COLORS, ml4t_diverging, show_plotly_with_alt
structlog.configure(wrapper_class=structlog.make_filtering_bound_logger(logging.WARNING))
# The provider package imports optional Oanda code whose Python 3.14 syntax warnings are
# unrelated to these factor providers. Scope suppression to that one optional-package import.
with warnings.catch_warnings():
warnings.simplefilter("ignore", SyntaxWarning)
provider_module = importlib.import_module("ml4t.data.providers")
AQRFactorProvider = provider_module.AQRFactorProvider
FamaFrenchProvider = provider_module.FamaFrenchProvider
```
### What each setting decides
There is no date range to choose: each factor's history is whatever its publisher has released,
and they differ, which is itself something the figures have to show rather than hide.
What is a choice is the evidence bar. Two thresholds are bound below: the conventional one, which
assumes a single hypothesis was tested, and the higher one proposed for a literature that has
searched a large space of candidates. Which is appropriate is the subject of Part 2.
```python
DISCOVERY_T_THRESHOLD = 3.0
CONVENTIONAL_T_THRESHOLD = 2.0
```
## Initialize Data Providers
Two providers supply everything below: AQR Capital Management publishes the long-history and
cross-asset series as Excel workbooks, and Kenneth French's Data Library publishes the
three-, five- and six-factor US equity series.
```python
try:
aqr = AQRFactorProvider(data_path=DATA_DIR / "factors" / "aqr")
except FileNotFoundError as err:
raise RuntimeError(
"AQR factor data is required for this notebook. Run AQRFactorProvider.download() first."
) from err
# Fetch French source archives afresh so the signed run exercises download and parsing
# rather than certifying a stale local cache hit. AQR's downloaded workbooks are the
# declared canonical inputs and are reconciled separately in the evidence bundle.
ff = FamaFrenchProvider(cache_path=DATA_DIR / "factors" / "fama-french", use_cache=False)
print(f"AQR: {len(aqr.list_datasets())} datasets in {len(aqr.list_categories())} categories")
print(f"French: {len(ff.list_datasets())} datasets in {len(ff.list_categories())} categories")
```
These provider checks look mundane, but they matter because the notebook combines
several sources with different histories and naming conventions. Verifying coverage
up front avoids drawing strong conclusions from mismatched sample periods later on.
## Load Core Factor Data
We'll load the key factor datasets for our analysis:
1. **Fama-French factors**: FF3, FF5, Momentum (1926-present)
2. **AQR equity factors**: QMJ and BAB
3. **Cross-asset**: VME (1972-present), TSMOM (1985-present)
4. **Long-history**: Century of Factor Premia (1926-2024 in this source vintage)
```python
# Fama-French core factors
ff3 = ff.fetch("ff3") # Market, SMB, HML, RF
ff5 = ff.fetch("ff5") # Adds RMW, CMA
mom = ff.fetch("mom") # Momentum factor
# Join once-fetched source frames on their common monthly timestamps.
ff4 = ff3.join(mom, on="timestamp", how="inner")
ff6 = ff5.join(mom, on="timestamp", how="inner")
print(
f"Three-factor: {len(ff3):,} months, "
f"{ff3['timestamp'].min():%Y-%m} to {ff3['timestamp'].max():%Y-%m}"
)
print(f"Five-factor: {len(ff5):,} months")
print(f"Momentum: {len(mom):,} months")
```
```python
# AQR factors (if available)
qmj = aqr.fetch("qmj_factors")
bab = aqr.fetch("bab_factors")
vme = aqr.fetch("vme_factors")
tsmom = aqr.fetch("tsmom")
century = aqr.fetch("century_premia")
for label, frame in (
("Quality minus junk", qmj),
("Betting against beta", bab),
("Value and momentum everywhere", vme),
("Time-series momentum", tsmom),
("Century of factor premia", century),
):
print(f"{label:<32} {len(frame):>6,} months")
print(f"Longest history runs through {century['timestamp'].max():%Y-%m}")
```
## Return Units and Conventions
Every factor series here is a **monthly decimal return**: a hundredth means one percent for the
month. The providers reach that convention differently and it is worth knowing which is which,
because mixing the two silently rescales everything by a hundred.
- `FamaFrenchProvider` divides the raw French data, which is published in percent, on fetch.
- `AQRFactorProvider` returns its files as they come, already in decimals.
These are **current-vintage published research series**. Providers can revise,
backfill, or extend their histories. Observation timestamps identify return months,
not the dates on which the final research series became available to investors.
The notebook therefore presents historical evidence, not a live-vintage backtest.
**Terminology**:
- `Mkt-RF`: Market excess return (market return minus risk-free rate)
- `HML`, `SMB`, `MOM`: Long-short factor portfolios (already excess returns)
```python
# Convert to pandas for analysis (provider returns Polars)
ff3_pd = ff3.to_pandas().set_index("timestamp")
mom_pd = mom.to_pandas().set_index("timestamp")
ff4_pd = ff4.to_pandas().set_index("timestamp")
# Sanity check: verify decimal units (fail if provider has a bug)
for col in ["Mkt-RF", "HML", "SMB", "MOM"]:
if col in ff4_pd.columns:
median_abs = ff4_pd[col].dropna().abs().median()
assert median_abs < 0.20, (
f"{col}: median |r| = {median_abs:.2f} suggests percent units (provider bug)"
)
print(
f"Combined French panel: {len(ff4_pd):,} months, "
f"{ff4_pd.index.min():%Y-%m} to {ff4_pd.index.max():%Y-%m}"
)
```
The unit sanity check is worth keeping in a teaching notebook because factor datasets
are notorious for mixing percent and decimal conventions. A silent unit error would
completely distort every Sharpe ratio and t-statistic that follows.
## Reference Data: NBER Recessions
Source: NBER Business Cycle Dating Committee (https://www.nber.org/research/data/us-business-cycle-expansions-and-contractions).
The static list includes NBER recessions through the 2020 contraction. Future
contractions require an explicit source update.
```python
# NBER recession dates (peak to trough)
RECESSIONS = [
("1929-08-01", "1933-03-01"), # Great Depression
("1937-05-01", "1938-06-01"), # 1937 Recession
("1945-02-01", "1945-10-01"), # Post-WWII
("1948-11-01", "1949-10-01"), # 1948 Recession
("1953-07-01", "1954-05-01"), # 1953 Recession
("1957-08-01", "1958-04-01"), # 1957 Recession
("1960-04-01", "1961-02-01"), # 1960 Recession
("1969-12-01", "1970-11-01"), # 1969 Recession
("1973-11-01", "1975-03-01"), # Oil Crisis
("1980-01-01", "1980-07-01"), # 1980 Recession
("1981-07-01", "1982-11-01"), # Early 80s Recession
("1990-07-01", "1991-03-01"), # 1990 Recession
("2001-03-01", "2001-11-01"), # Dot-com Bust
("2007-12-01", "2009-06-01"), # Global Financial Crisis
("2020-02-01", "2020-04-01"), # COVID-19
]
# Major crisis periods for detailed analysis
CRISES = {
"1987 Black Monday": ("1987-10-01", "1987-10-31"),
"1998 LTCM": ("1998-08-01", "1998-10-31"),
"2000-02 Dot-com": ("2000-03-01", "2002-10-31"),
"2008-09 GFC": ("2007-10-01", "2009-03-31"),
"2020 COVID": ("2020-02-01", "2020-03-31"),
"2022 Inflation": ("2022-01-01", "2022-10-31"),
}
```
---
# Part 1: A Century of Factor Performance
## Are Factor Premia Real?
One of the strongest criticisms of factor investing is that factor premia were discovered
through data mining on the 1963-1990 period. The **Century of Factor Premia** dataset
from AQR addresses this with a history beginning in 1926, much of which predates
the academic publication of these factors.
> "The fact that value and momentum premia exist in data that predates their discovery
> provides compelling pre-publication evidence." - Ilmanen et al. (2021)
**Important distinction**: This is "pre-discovery" or "pre-publication" evidence, not
strictly "out-of-sample" in the workflow sense (which would require pre-registration
before seeing any data). The evidence is strong because it's less susceptible to
post hoc window selection.
One colour assignment for the whole notebook, so a factor keeps its identity from figure to
figure. Where more than five factors appear at once, one is emphasised and the rest are neutral
rather than each getting its own hue: a reader cannot hold eight colours in mind, and a chart
that asks them to is a chart nobody reads.
```python
factor_colors = {
"Mkt-RF": COLORS["blue"],
"HML": COLORS["copper"],
"SMB": COLORS["neutral"],
"MOM": COLORS["amber"],
"Value": COLORS["copper"],
"Momentum": COLORS["amber"],
"Carry": COLORS["positive"],
"Defensive": COLORS["slate"],
}
labels = {
"Mkt-RF": "Market (Mkt-RF)",
"HML": "Value (HML)",
"SMB": "Size (SMB)",
"MOM": "Momentum (MOM)",
"RMW": "Profitability (RMW)",
"CMA": "Investment (CMA)",
"QMJ": "Quality (QMJ)",
"BAB": "Low-Vol (BAB)",
"TSMOM": "Trend (TSMOM)",
}
```
```python
cum_returns = (1 + ff4_pd[["Mkt-RF", "HML", "SMB", "MOM"]]).cumprod()
terminal_growth = cum_returns.iloc[-1]
fig = go.Figure()
for col in ["Mkt-RF", "MOM", "HML", "SMB"]:
fig.add_trace(
go.Scatter(
x=cum_returns.index,
y=cum_returns[col],
name=labels.get(col, col),
line=dict(color=factor_colors[col], width=2),
hovertemplate="%{x|%Y-%m}<br>%{y:.2f}x<extra></extra>",
)
)
```
### Add recession context
Recession shading separates secular compounding from crisis behavior.
```python
# Add recession bands
for start, end in RECESSIONS:
start_dt = datetime.strptime(start, "%Y-%m-%d")
end_dt = datetime.strptime(end, "%Y-%m-%d")
if start_dt >= cum_returns.index.min():
fig.add_vrect(
x0=start_dt,
x1=end_dt,
fillcolor=COLORS["silver_muted"],
opacity=0.35,
layer="below",
line_width=0,
)
wealth_tick_values = [0.25, 0.5, 1, 2, 5, 10, 20, 50, 100, 200, 500, 1_000]
wealth_tick_text = [f"{value:g}x" for value in wealth_tick_values]
fig.update_layout(
title=(
f"Market and momentum dominate long-run factor wealth<br><sup>Growth of $1 in "
f"published long-short factor returns, {cum_returns.index.min():%Y-%m} to "
f"{cum_returns.index.max():%Y-%m}; log scale; shaded bands are NBER recessions</sup>"
),
xaxis_title="Month",
yaxis_title="Growth of $1 (log scale)",
yaxis_type="log",
yaxis=dict(tickmode="array", tickvals=wealth_tick_values, ticktext=wealth_tick_text),
legend=dict(yanchor="top", y=0.99, xanchor="left", x=0.01),
hovermode="x unified",
height=500,
)
show_plotly_with_alt(
fig,
"Four growth-of-one-dollar paths on a logarithmic axis from the 1920s to the present, for the market, momentum, value and size factors, with recessions shaded.",
)
```
### What the wealth paths establish
The market, value, size, and momentum series are published long-short research
portfolios. Their full-history wealth paths establish historical magnitude, not
implementable investor returns. The later inference and period-comparison sections
test whether the apparent premia are statistically distinguishable and whether SMB
weakened after its 1981 publication.
> **Important**: These are **long-short factor portfolio returns**, not tradable
> strategies. The cumulative return charts above represent theoretical wealth paths
> assuming: (1) perfect rebalancing, (2) zero transaction costs, (3) unlimited
> leverage capacity, and (4) no financing costs for short positions. Actual
> implementation of factor strategies involves significant slippage, financing
> costs, and capacity constraints. Use this data as evidence about published factor
> returns, not as a literal investable wealth trajectory.
```python
print("Growth of one dollar over the shared French history, gross of everything:")
for name, value in sorted(terminal_growth.items(), key=lambda item: -item[1]):
print(f" {name:<8} ${value:>12,.0f}")
```
Those magnitudes are what a long compounding window does to a modest monthly edge, and they are
also why long-horizon factor charts are misleading if read as achievable. Nothing has been
deducted: no financing on the short leg, no turnover, no capacity limit, no tax. The series is a
research construct measured on paper, and the gap between it and an investable version widens
with exactly the turnover that makes some of these factors work.
## Century of Factor Premia: Pre-Publication Evidence
The AQR "Century of Factor Premia" dataset provides pre-discovery evidence for
factor premia. Much of this data (1926-1963) predates the academic publication
of factors, making it less susceptible to post hoc window selection.
Note: "Pre-discovery" is more accurate than "out-of-sample" here. True out-of-sample
would require pre-registration of the strategy before seeing any data.
```python
century_pd = century.to_pandas().set_index("timestamp")
century_columns = [
"All asset classes Value",
"All asset classes Momentum",
"All asset classes Carry",
"All asset classes Defensive",
]
assert set(century_columns) <= set(century_pd.columns)
century_factors = century_pd[century_columns].dropna()
cum_century = (1 + century_factors).cumprod()
print(
f"Century panel: {len(century_factors):,} common months, "
f"{century_factors.index.min():%Y-%m} to {century_factors.index.max():%Y-%m}"
)
```
The four columns selected are the all-asset-class aggregates, which combine each factor's
stock-selection and macro sleeves into one series. Taking those rather than the per-region
variants keeps four lines on the chart instead of forty, and keeps them comparable: each is the
same factor definition applied across the same set of markets.
```python
fig = go.Figure()
for col in cum_century.columns:
short_name = col.removeprefix("All asset classes ")
fig.add_trace(
go.Scatter(
x=cum_century.index,
y=cum_century[col],
name=short_name,
line=dict(color=factor_colors[short_name], width=2),
hovertemplate="%{x|%Y-%m}<br>%{y:.2f}x<extra></extra>",
)
)
fig.update_layout(
title=(
"Four aggregate premia persist across the Century sample"
f"<br><sup>AQR published factor returns, {cum_century.index.min():%Y-%m} to "
f"{cum_century.index.max():%Y-%m}; pre-publication history is not a "
"preregistered holdout</sup>"
),
xaxis_title="Month",
yaxis_title="Growth of $1 (log scale)",
yaxis_type="log",
legend=dict(yanchor="top", y=0.99, xanchor="left", x=0.01),
height=500,
)
show_plotly_with_alt(
fig,
"Four growth-of-one-dollar paths on a logarithmic axis for the value, momentum, carry and defensive aggregates of the Century of Factor Premia data, each ending far above one, momentum highest, with value and carry roughly flat since about 2010.",
)
```
---
## Part 1 Summary: what a premium in the long history is evidence of
Value, momentum, carry and defensive all show positive premia in data reaching back to 1926,
most of which predates their publication. That makes the finding harder to explain as a window
chosen after the fact, which is the specific criticism the long history answers.
What it establishes is that the premium was there in the returns, gross. Four things stand
between a gross premium and a return an allocator receives, and none of them is measured in a
published research series: the turnover cost of holding the portfolio, the financing cost of
its short leg, the capacity at which its own trading moves the prices it trades, and the drag
from rebalancing timing and corporate actions. Chapter 16 models the first, and Chapter 18
measures the third.
---
---
# Part 2: Statistical Rigor: Do Factors Pass Significance Tests?
The conventional significance threshold assumes one hypothesis was tested. Hundreds of factors
have been published, and the ones that reached publication are the ones that cleared a bar, so
the conventional threshold is the wrong one to judge them by.
Harvey, Liu and Zhu (2016) propose raising it, on the reasoning that a literature searching a
large space needs a correspondingly higher bar for any single finding. Their conclusion is worth
quoting because it is not hedged:
> "Most claimed research findings in financial economics are likely false."
Four things about the higher threshold used below, because it is easy to apply it to the wrong
quantity:
1. It applies to the t-statistic of the **mean return**, not to the significance of a Sharpe
ratio, which is a different statistic with a different distribution.
2. It is a proposal for cross-sectional factor discovery, not a general law of inference.
3. What threshold is right depends on how many things were tried. A single pre-registered test
does not need the same bar as a search over hundreds of candidates.
4. The t-statistics here carry a Newey-West correction, because monthly factor returns are
autocorrelated and an uncorrected standard error would be too small.
### The t-statistic the threshold applies to
A Newey-West corrected t-statistic for the mean excess return. The correction matters: monthly
factor returns are serially dependent, and an uncorrected standard error understates the
uncertainty, which inflates every t-statistic in the same direction.
```python
def calculate_mean_return_tstat(returns, periods_per_year=12):
"""Calculate the Newey-West mean-return t-statistic used in the Harvey test."""
if hasattr(returns, "values"):
returns = returns.values
returns = returns[~np.isnan(returns)]
n = len(returns)
if n < 12:
return {"mean_tstat": np.nan, "mean_return": np.nan, "nw_se_monthly": np.nan}
mean_ret = np.mean(returns)
# Newey-West lag selection
max_lag = int(np.floor(4 * (n / 100) ** (2 / 9)))
max_lag = max(1, min(max_lag, n // 4))
# Compute Newey-West HAC variance
resid = returns - mean_ret
gamma_0 = np.sum(resid**2) / n
gamma_sum = 0
for j in range(1, max_lag + 1):
weight = 1 - j / (max_lag + 1) # Bartlett kernel
gamma_j = np.sum(resid[j:] * resid[:-j]) / n
gamma_sum += 2 * weight * gamma_j
nw_var = (gamma_0 + gamma_sum) / n
nw_se = np.sqrt(max(nw_var, 1e-10))
mean_tstat = mean_ret / nw_se if nw_se > 0 else 0
# Annualization: mean_return scales by periods_per_year; nw_se kept monthly.
# Inference uses mean_tstat (computed on monthly data), not annualized SE.
return {
"mean_tstat": mean_tstat,
"mean_return": mean_ret * periods_per_year,
"nw_se_monthly": nw_se,
}
```
### Sharpe Ratio Uncertainty
The Sharpe ratio t-statistic is **different** from the Harvey threshold. The
approximation below rescales both the monthly estimate and its standard error to
annual units, then inflates uncertainty for return autocorrelation. It is a
transparent Lo-style diagnostic, not the full covariance expression in Lo (2002).
```python
def compute_autocorrelation_adjustment(returns, max_lag: int) -> tuple[list[float], float]:
autocorrs = []
for k in range(1, max_lag + 1):
if len(returns) > k + 1:
rho_k = np.corrcoef(returns[:-k], returns[k:])[0, 1]
autocorrs.append(np.clip(rho_k, -0.95, 0.95))
else:
autocorrs.append(0.0)
lr_var_factor = 1.0
for k, rho_k in enumerate(autocorrs, start=1):
weight = 1 - k / (max_lag + 1)
lr_var_factor += 2 * weight * rho_k
return autocorrs, max(lr_var_factor, 0.1)
```
Reuse the autocorrelation adjustment inside the Sharpe function so the notebook can
keep the statistical assumptions explicit without burying them in one large cell.
```python
def calculate_sharpe_stats(returns, periods_per_year=12):
"""Calculate annualized Sharpe statistics with a serial-correlation adjustment."""
if hasattr(returns, "values"):
returns = returns.values
returns = returns[~np.isnan(returns)]
n = len(returns)
monthly_mean = np.mean(returns)
monthly_std = np.std(returns, ddof=1)
mean_ret = monthly_mean * periods_per_year
std_ret = monthly_std * np.sqrt(periods_per_year)
monthly_sharpe = monthly_mean / monthly_std if monthly_std > 0 else 0
sharpe = monthly_sharpe * np.sqrt(periods_per_year)
max_lag = int(np.floor(4 * (n / 100) ** (2 / 9)))
max_lag = max(1, min(max_lag, n // 4))
autocorrs, lr_var_factor = compute_autocorrelation_adjustment(returns, max_lag)
se_monthly = np.sqrt((1 + 0.5 * monthly_sharpe**2) / n)
se_sharpe = se_monthly * np.sqrt(lr_var_factor * periods_per_year)
sharpe_tstat = sharpe / se_sharpe if se_sharpe > 0 else 0
ci_lower = sharpe - 1.96 * se_sharpe
ci_upper = sharpe + 1.96 * se_sharpe
p_value = 2 * stats.norm.sf(abs(sharpe_tstat))
mean_stats = calculate_mean_return_tstat(returns, periods_per_year)
rho1 = autocorrs[0] if autocorrs else 0.0
return {
"sharpe": sharpe,
"sharpe_tstat": sharpe_tstat,
"t_stat": mean_stats["mean_tstat"],
"p_value": p_value,
"ci_lower": ci_lower,
"ci_upper": ci_upper,
"ann_return": mean_ret,
"ann_vol": std_ret,
"n_months": n,
"autocorr": rho1,
"lr_var_factor": lr_var_factor,
"sharpe_se": se_sharpe,
}
```
### Maximum Drawdown
Simple utility for computing peak-to-trough drawdown.
```python
def calculate_max_drawdown(returns):
"""Calculate maximum drawdown from return series."""
if hasattr(returns, "values"):
returns = returns.values
returns = returns[~np.isnan(returns)]
cum_returns = np.cumprod(1 + returns)
rolling_max = np.maximum.accumulate(cum_returns)
drawdowns = cum_returns / rolling_max - 1
return np.min(drawdowns)
```
### Compute Factor Statistics
Apply the statistical functions to all available factors from Fama-French and AQR.
```python
# Calculate statistics for all factors
factor_stats = {}
# French factors use each source's full available history for inference. The
# inner-joined FF4 panel remains appropriate only for shared-history comparisons.
for col in ["Mkt-RF", "HML", "SMB"]:
factor_stats[col] = calculate_sharpe_stats(ff3_pd[col].dropna())
factor_stats["MOM"] = calculate_sharpe_stats(mom_pd["MOM"].dropna())
# FF5 additional factors (convert once, reuse)
ff5_pd = ff5.to_pandas().set_index("timestamp")
for col in ["RMW", "CMA"]:
if col in ff5_pd.columns:
factor_stats[col] = calculate_sharpe_stats(ff5_pd[col].dropna())
# AQR factors
if qmj is not None:
qmj_pd = qmj.to_pandas().set_index("timestamp")
if "USA" in qmj_pd.columns:
factor_stats["QMJ"] = calculate_sharpe_stats(qmj_pd["USA"].dropna())
elif "Global" in qmj_pd.columns:
factor_stats["QMJ"] = calculate_sharpe_stats(qmj_pd["Global"].dropna())
if bab is not None:
bab_pd = bab.to_pandas().set_index("timestamp")
if "USA" in bab_pd.columns:
factor_stats["BAB"] = calculate_sharpe_stats(bab_pd["USA"].dropna())
elif "Global" in bab_pd.columns:
factor_stats["BAB"] = calculate_sharpe_stats(bab_pd["Global"].dropna())
```
```python
stats_df = pd.DataFrame(factor_stats).T
stats_df["significant_conventional"] = stats_df["t_stat"] > CONVENTIONAL_T_THRESHOLD
stats_df["significant_harvey"] = stats_df["t_stat"] > DISCOVERY_T_THRESHOLD
print(
f"Clear the conventional bar of {CONVENTIONAL_T_THRESHOLD}: "
f"{int(stats_df['significant_conventional'].sum())} of {len(stats_df)} factors"
)
print(
f"Clear the Harvey bar of {DISCOVERY_T_THRESHOLD}: "
f"{int(stats_df['significant_harvey'].sum())} of {len(stats_df)} factors"
)
factor_order = ["Mkt-RF", "MOM", "HML", "RMW", "CMA", "QMJ", "BAB", "SMB"]
factor_order = [f for f in factor_order if f in factor_stats]
```
The Harvey threshold belongs on the mean-return t-statistic itself. Bars above
three clear the conservative discovery screen; annualized Sharpe remains in the
hover text as a separate economic-magnitude diagnostic.
```python
fig = go.Figure()
fig.add_trace(
go.Bar(
x=[labels.get(factor, factor) for factor in factor_order],
y=[factor_stats[factor]["t_stat"] for factor in factor_order],
marker_color=[
COLORS["blue"]
if factor_stats[factor]["t_stat"] > DISCOVERY_T_THRESHOLD
else COLORS["neutral"]
for factor in factor_order
],
customdata=np.array([[factor_stats[factor]["sharpe"]] for factor in factor_order]),
hovertemplate="<b>%{x}</b><br>Mean t (NW): %{y:.2f}<br>Annualized Sharpe: "
"%{customdata[0]:.2f}<extra></extra>",
)
)
fig.add_hline(y=DISCOVERY_T_THRESHOLD, line_dash="dash", line_color=COLORS["amber"], line_width=2)
fig.add_hline(
y=CONVENTIONAL_T_THRESHOLD, line_dash="dot", line_color=COLORS["neutral"], line_width=2
)
fig.add_hline(y=0, line_color=COLORS["neutral"], line_width=1)
fig.update_layout(
title=(
"Two thresholds, and the factors that fall between them"
f"<br><sup>Newey-West t-statistics on mean monthly returns. Dotted line: the conventional "
f"threshold of {CONVENTIONAL_T_THRESHOLD:.0f}. Dashed line: Harvey's discovery threshold "
f"of {DISCOVERY_T_THRESHOLD:.0f}. Bars between them clear the first and not the second. "
"Histories differ by factor.</sup>"
),
xaxis_title="Published factor portfolio",
yaxis_title="Mean-return t-statistic (Newey-West)",
showlegend=False,
height=450,
)
show_plotly_with_alt(
fig,
"Bars of the Newey-West t-statistic on each factor's mean monthly return, with a dotted line at the conventional threshold of two and a dashed line at the discovery threshold of three; bars clearing the higher line are coloured.",
)
```
### Test the SMB decay claim directly
A full-sample SMB statistic cannot establish post-publication decay. Splitting at
the 1981 publication year makes the comparison visible while remaining explicitly
ex-post and descriptive.
```python
smb_returns = ff3_pd["SMB"].dropna()
smb_periods = {
"Pre-1981": smb_returns[smb_returns.index < "1981-01-01"],
"1981-present": smb_returns[smb_returns.index >= "1981-01-01"],
}
smb_period_stats = {name: calculate_sharpe_stats(values) for name, values in smb_periods.items()}
fig = go.Figure(
go.Bar(
x=list(smb_period_stats),
y=[stats["ann_return"] * 100 for stats in smb_period_stats.values()],
marker_color=[COLORS["blue"], COLORS["neutral"]],
text=[f"t = {stats['t_stat']:.2f}" for stats in smb_period_stats.values()],
textposition="outside",
hovertemplate="<b>%{x}</b><br>Annualized mean: %{y:.2f}%<br>%{text}<extra></extra>",
)
)
fig.add_hline(y=0, line_color=COLORS["neutral"], line_width=1)
fig.update_layout(
title=(
"The size premium's two halves differ across its publication year"
"<br><sup>Annualized mean monthly return; labels show Newey-West t-statistics; "
"ex-post period split</sup>"
),
xaxis_title="Sample period",
yaxis_title="Annualized mean return (%)",
height=420,
)
show_plotly_with_alt(
fig,
"Two bars of annualized mean return for the size factor, before and from 1981, each labelled with its Newey-West t-statistic.",
)
```
```python
for period, period_stats in smb_period_stats.items():
print(
f"{period:<14} annualized mean {period_stats['ann_return']:>7.1%} "
f"t-statistic {period_stats['t_stat']:>5.2f}"
)
```
The split date is the factor's publication year, chosen with the benefit of knowing what happened
after it. That makes the comparison a description of two halves and not a test: three
explanations fit it equally well. The premium may have been arbitraged away once it was known.
The market may have changed for unrelated reasons over four decades. Or the first half's estimate
may simply have been high, in which case the second half is the truth and the decay is an
artifact of where the line was drawn.
Distinguishing them needs something this split does not have, which is a reason chosen in advance
to expect the break at that date and not another.
---
# Part 3: Value and Momentum Everywhere
Asness, Moskowitz & Pedersen (2013) document value and momentum premia not just in US
equities, but across **8 different asset classes globally** - a result that mitigates
the single-market data-mining critique applied to factor work on US equities alone.
> "We find consistent value and momentum return premia across eight diverse markets
> and asset classes." - Asness, Moskowitz & Pedersen (2013)
```python
vme_pd = vme.to_pandas().set_index("timestamp")
print(f"Value-and-momentum panel: {len(vme_pd):,} months, {len(vme_pd.columns)} series")
```
Map AQR VME column suffixes (`VALLS_VME_XX90`, `MOMLS_VME_XX90`) to
display labels. `EQ`, `FX`, `FI`, `COM` cover cross-asset buckets;
`US90`, `UK90`, `ROE90`, `JP90` cover regional equity blocks.
```python
asset_class_map = {
"US90": "US Equities",
"UK90": "UK Equities",
"ROE90": "Europe ex-UK",
"JP90": "Japan",
"EQ": "All Equities",
"FX": "Currencies",
"FI": "Fixed Income",
"COM": "Commodities",
}
vme_stats = []
```
Aggregate value/momentum row (the "Everywhere" line) - uses the bundled
`VAL` / `MOM` columns rather than any single asset class.
```python
if "VAL" in vme_pd.columns and "MOM" in vme_pd.columns:
val_data = vme_pd["VAL"].dropna()
mom_data = vme_pd["MOM"].dropna()
if len(val_data) > 12 and len(mom_data) > 12:
common_idx = val_data.index.intersection(mom_data.index)
val_stats = calculate_sharpe_stats(val_data.loc[common_idx])
mom_stats = calculate_sharpe_stats(mom_data.loc[common_idx])
corr = val_data.loc[common_idx].corr(mom_data.loc[common_idx])
vme_stats.append(
{
"Asset Class": "EVERYWHERE",
"Value SR": val_stats["sharpe"],
"Value t": val_stats["t_stat"],
"Momentum SR": mom_stats["sharpe"],
"Momentum t": mom_stats["t_stat"],
"Val-Mom Corr": corr,
"Common Months": len(common_idx),
}
)
```
```python
# Then add by asset class.
for suffix, display_name in asset_class_map.items():
val_col = f"VALLS_VME_{suffix}"
mom_col = f"MOMLS_VME_{suffix}"
assert val_col in vme_pd.columns and mom_col in vme_pd.columns
val_data = vme_pd[val_col].dropna()
mom_data = vme_pd[mom_col].dropna()
common_idx = val_data.index.intersection(mom_data.index)
assert len(common_idx) > 12
val_stats = calculate_sharpe_stats(val_data.loc[common_idx])
mom_stats = calculate_sharpe_stats(mom_data.loc[common_idx])
vme_stats.append(
{
"Asset Class": display_name,
"Value SR": val_stats["sharpe"],
"Value t": val_stats["t_stat"],
"Momentum SR": mom_stats["sharpe"],
"Momentum t": mom_stats["t_stat"],
"Val-Mom Corr": val_data.loc[common_idx].corr(mom_data.loc[common_idx]),
"Common Months": len(common_idx),
}
)
```
```python
assert len(vme_stats) == 9
vme_df = pd.DataFrame(vme_stats)
class_corr = vme_df.loc[vme_df["Asset Class"] != "EVERYWHERE", "Val-Mom Corr"]
median_vme_corr = float(class_corr.median())
min_vme_corr = float(class_corr.min())
max_vme_corr = float(class_corr.max())
```
The cross-asset comparison matters more than any single region. The figure below
separates economic magnitude (Sharpe) from the diversification statistic (correlation).
Two-panel frame: Sharpe ratios by asset class
on the left, value-momentum correlation on the right.
```python
fig = make_subplots(
rows=1,
cols=2,
subplot_titles=("Annualized Sharpe", "Value-Momentum Correlation"),
)
```
Left panel - Sharpe-ratio bars for Value and Momentum side-by-side.
```python
_ = fig.add_trace(
go.Bar(
x=vme_df["Asset Class"],
y=vme_df["Value SR"],
name="Value",
marker_color=factor_colors["Value"],
),
row=1,
col=1,
)
_ = fig.add_trace(
go.Bar(
x=vme_df["Asset Class"],
y=vme_df["Momentum SR"],
name="Momentum",
marker_color=factor_colors["Momentum"],
),
row=1,
col=1,
)
```
Right panel - value-momentum correlation, with blue for negative
(diversifying) observations and red for positive observations.
```python
_ = fig.add_trace(
go.Bar(
x=vme_df["Asset Class"],
y=vme_df["Val-Mom Corr"],
name="Correlation",
marker_color=[
COLORS["negative"] if value > 0 else COLORS["blue"] for value in vme_df["Val-Mom Corr"]
],
showlegend=False,
),
row=1,
col=2,
)
```
```python
fig.add_hline(y=0, line_dash="dash", line_color=COLORS["neutral"], row=1, col=2)
fig.update_layout(
title=(
"Value and momentum are negatively correlated in every asset class"
"<br><sup>Monthly published long-short returns, before any implementation cost</sup>"
),
barmode="group",
height=480,
legend=dict(yanchor="top", y=0.99, xanchor="right", x=0.45),
)
fig.update_xaxes(title_text="Asset class", tickangle=-30)
fig.update_yaxes(title_text="Annualized Sharpe ratio", row=1, col=1)
_ = fig.update_yaxes(title_text="Correlation", range=[-1, 1], row=1, col=2)
```
```python
show_plotly_with_alt(
fig,
"Two panels by asset class: annualized Sharpe ratios for value and momentum side by side on the left, and their correlation on the right, every bar on the right below zero.",
)
```
```python
everywhere_corr = float(vme_df.loc[vme_df["Asset Class"] == "EVERYWHERE", "Val-Mom Corr"].iloc[0])
print("Value-momentum correlation across the sleeves:")
print(f" lowest {min_vme_corr:+.2f}")
print(f" median {median_vme_corr:+.2f}")
print(f" highest {max_vme_corr:+.2f}")
print(f" pooled across everything: {everywhere_corr:+.2f}")
```
### The Diversification Benefit
A crucial finding is that **value and momentum are negatively correlated** in
every VME sleeve shown. This creates diversification potential, but the published
long-short returns do not include a reader's implementation costs or constraints.
The mechanism behind the negative correlation is that the two strategies buy, by construction,
nearly opposite things. Value buys what has fallen and is now cheap on a fundamental ratio;
momentum buys what has risen recently. An asset that has just dropped enters the value book and
leaves the momentum one on the same information, so when one side is being rewarded the other
tends not to be.
That is a structural argument rather than an empirical one, which is what makes it worth more
than the correlation estimate itself. A negative correlation measured on one sample can be
sampling noise; a negative correlation with a reason to be negative is likelier to persist.
---
# Part 4: Factor Correlations and Diversification
Correlations need every series present on the same months, so the panel below is restricted to
the window in which all of them have history. That is set by the shortest series, and it discards
decades of the longest ones - a real cost, and the alternative is a correlation matrix whose
entries are computed on different samples and therefore not comparable with one another.
```python
combined_pl = ff6.select(["timestamp", "Mkt-RF", "SMB", "HML", "RMW", "CMA", "MOM"]).with_columns(
pl.col("timestamp").cast(pl.Datetime("us"))
)
# Add the AQR factors to the common-period panel.
if qmj is not None and "USA" in qmj.columns:
qmj_usa = qmj.select([pl.col("timestamp"), pl.col("USA").alias("QMJ")]).with_columns(
pl.col("timestamp").cast(pl.Datetime("us"))
)
combined_pl = combined_pl.join(qmj_usa, on="timestamp", how="full", coalesce=True)
if bab is not None and "USA" in bab.columns:
bab_usa = bab.select([pl.col("timestamp"), pl.col("USA").alias("BAB")]).with_columns(
pl.col("timestamp").cast(pl.Datetime("us"))
)
combined_pl = combined_pl.join(bab_usa, on="timestamp", how="full", coalesce=True)
# Filter to common period (1972+) and drop nulls
combined_pl = combined_pl.filter(pl.col("timestamp") >= pl.date(1972, 1, 1)).drop_nulls()
print(
f"Common eight-factor panel: {len(combined_pl):,} months, "
f"{combined_pl['timestamp'].min():%Y-%m} to {combined_pl['timestamp'].max():%Y-%m}"
)
# Convert to pandas only for correlation and visualization
combined_data = combined_pl.to_pandas().set_index("timestamp")
```
```python
# Compute the common-period factor correlation matrix.
corr_matrix = combined_data.corr()
# Create labels for display
corr_labels = {col: labels.get(col, col) for col in corr_matrix.columns}
corr_display = corr_matrix.copy()
corr_display.index = [corr_labels.get(c, c) for c in corr_display.index]
corr_display.columns = [corr_labels.get(c, c) for c in corr_display.columns]
```
Mask the duplicate upper triangle and use a fixed diverging scale centered at zero.
```python
corr_plot = corr_display.mask(np.triu(np.ones(corr_display.shape, dtype=bool), k=1))
corr_text = corr_plot.apply(
lambda column: column.map(lambda value: "" if pd.isna(value) else f"{value:.2f}")
)
fig = go.Figure(
go.Heatmap(
z=corr_plot.to_numpy(),
x=corr_plot.columns,
y=corr_plot.index,
text=corr_text.to_numpy(),
texttemplate="%{text}",
colorscale=ml4t_diverging(),
zmin=-1,
zmax=1,
zmid=0,
colorbar=dict(title="Correlation"),
hovertemplate="%{y} vs %{x}<br>Correlation: %{z:.2f}<extra></extra>",
)
)
```
Add a message-first title and explicit dimensions.
```python
_ = fig.update_layout(
title=(
"Pairwise correlations among the factors, over their shared history"
"<br><sup>Pairwise correlations over the window in which all factors have history</sup>"
),
height=560,
margin=dict(l=135, r=90, b=120),
xaxis_title="Factor",
yaxis_title="Factor",
)
show_plotly_with_alt(
fig,
"Lower-triangle correlation heatmap of eight factors over their shared window, each cell labelled, on a diverging scale centred at zero.",
)
```
```python
print(
"Value against momentum, US equity factors over the shared window: "
f"{corr_matrix.loc['HML', 'MOM']:+.2f}"
)
print(f"The same pair pooled across all VME asset classes: {everywhere_corr:+.2f}")
```
### Key Correlation Insights
1. **Value against momentum**: the two lines printed above the heatmap put the US equity
estimate beside the cross-asset one. Both are negative and they are not the same number,
which is what the cross-asset evidence adds: a relationship that holds in one market can be
a property of that market, and one that holds in eight is harder to explain that way.
2. **Quality (QMJ) and low-volatility (BAB) against the market**: three properties are easy to
run together and are not the same. *Long-short* says only that the portfolio holds both
sides. *Dollar-neutral* says the two legs are equal in size. *Beta-neutral* says their market
exposures cancel, which equal dollars do not deliver when the two legs have different betas.
BAB is built for the third: it levers its low-beta leg up to a beta of one, de-levers the
high-beta leg down to one, and holds them against each other, so its *estimated* net beta is
zero by construction. Realized exposure can still differ, because those betas are estimated
on past returns and move. The correlations in the lower triangle say how much each factor
co-moved with the market over this window. A correlation is not a beta - it is scaled by the
two volatilities - so these cells support a statement about co-movement, and the regression
that gives exposure is what `01_portfolio_metrics` runs on a return series.
---
# Part 5: Crisis Performance: Who Provides "Crisis Alpha"?
A factor that pays when the rest of the book is losing is worth more to an allocator than the
same premium earned in calm markets, because it is the one that lets the whole portfolio be
held through the episode. **Crisis alpha** is the name for a positive return during an equity
drawdown.
These windows are selected ex post. They describe historical co-movement; they do
not establish that a factor will insure a future crisis.
```python
# Calculate crisis period returns
crisis_factors = ["Mkt-RF", "HML", "SMB", "MOM"]
if "QMJ" in combined_data.columns:
crisis_factors.append("QMJ")
if "BAB" in combined_data.columns:
crisis_factors.append("BAB")
crisis_returns = {}
for crisis_name, (start, end) in CRISES.items():
start_dt = pd.to_datetime(start)
end_dt = pd.to_datetime(end)
mask = (combined_data.index >= start_dt) & (combined_data.index <= end_dt)
crisis_data = combined_data[mask]
if len(crisis_data) > 0:
crisis_returns[crisis_name] = {}
for col in crisis_factors:
if col in crisis_data.columns:
data = crisis_data[col].dropna()
if len(data) > 0:
cum_ret = (1 + data).prod() - 1
crisis_returns[crisis_name][col] = cum_ret
crisis_df = pd.DataFrame(crisis_returns).T
```
```python
tsmom_pd = tsmom.to_pandas().set_index("timestamp")
assert "TSMOM" in tsmom_pd.columns
for crisis_name, (start, end) in CRISES.items():
tsmom_crisis = tsmom_pd.loc[pd.to_datetime(start) : pd.to_datetime(end), "TSMOM"].dropna()
if len(tsmom_crisis) > 0:
crisis_df.loc[crisis_name, "TSMOM"] = (1 + tsmom_crisis).prod() - 1
```
A signed-return heatmap makes the crisis pattern legible without assigning seven
categorical colors. Blank cells denote factor histories that had not started.
```python
available_factors = [
f for f in ["Mkt-RF", "MOM", "HML", "QMJ", "BAB", "TSMOM", "SMB"] if f in crisis_df.columns
]
crisis_plot = crisis_df[available_factors].rename(columns=labels) * 100
tsmom_positive = int((crisis_df["TSMOM"].dropna() > 0).sum())
tsmom_observed = int(crisis_df["TSMOM"].notna().sum())
fig = go.Figure(
go.Heatmap(
z=crisis_plot.to_numpy(),
x=crisis_plot.columns,
y=crisis_plot.index,
text=np.round(crisis_plot.to_numpy(), 1),
texttemplate="%{text:.1f}%",
colorscale=ml4t_diverging(),
zmid=0,
colorbar=dict(title="Return (%)"),
hovertemplate="%{y}<br>%{x}: %{z:.1f}%<extra></extra>",
)
)
```
```python
fig.update_layout(
title=(
"Trend following was positive in most of these chosen windows"
"<br><sup>Cumulative monthly returns over windows chosen after the fact; blanks predate "
"a factor's history</sup>"
),
xaxis_title="Published factor portfolio",
yaxis_title="Selected crisis window",
height=520,
margin=dict(l=105, r=75, b=100),
)
show_plotly_with_alt(
fig,
"Heatmap of cumulative factor return, one row per named crisis window and one column per factor, red for losses and blue for gains, with blank cells where a factor's history had not begun.",
)
```
**Interpretation**: Crisis performance is where correlations and premia become
economically meaningful. A factor that compounds nicely on average but fails exactly
when the rest of the portfolio is under stress is much less valuable in allocation.
### Crisis Performance Observations
1. **Market (Mkt-RF)**: By definition, large negative returns during crises.
2. **Momentum (MOM)**: Mixed results. Provided crisis alpha in some events but
suffered the famous "momentum crash" in 2009 when the trend reversed violently.
3. **Value (HML)**: Generally negative during crises as cheap stocks get cheaper.
The 2020 COVID crash was particularly painful for value.
4. **Quality (QMJ)**: Its observed crisis returns vary by episode; the heatmap
shows when the historical "flight to quality" description does and does not fit.
5. **Trend (TSMOM)**: The aggregate published series is positive in most, but not
all, observed windows. Its 1985 start leaves earlier crises unmeasured.
---
# Part 6: Risk and Return in One View
### Summary Helper
The final chart uses one numeric row layout across French and AQR sources.
```python
def append_summary_row(
summary_stats: list[dict[str, object]],
factor: str,
source: str,
returns: pd.Series,
stats_dict: dict[str, float],
) -> None:
"""Append one numeric summary row for a factor series."""
summary_stats.append(
{
"Factor": labels.get(factor, factor),
"Source": source,
"Period": f"{returns.index.min():%Y}-{returns.index.max():%Y}",
"Ann. Return": stats_dict["ann_return"],
"Ann. Volatility": stats_dict["ann_vol"],
"Sharpe Ratio": stats_dict["sharpe"],
"t-statistic": stats_dict["t_stat"],
"Max Drawdown": calculate_max_drawdown(returns),
"Skewness": returns.skew(),
"Months": len(returns),
}
)
```
```python
# Build comprehensive summary table
summary_stats = []
# Fama-French factors
for factor in ["Mkt-RF", "HML", "SMB", "MOM"]:
if factor in factor_stats:
source_data = mom_pd if factor == "MOM" else ff3_pd
returns = source_data[factor].dropna()
s = factor_stats[factor]
append_summary_row(summary_stats, factor, "French", returns, s)
# FF5 additional factors (ff5_pd already converted above)
for factor in ["RMW", "CMA"]:
if factor in factor_stats:
returns = ff5_pd[factor].dropna()
s = factor_stats[factor]
append_summary_row(summary_stats, factor, "French", returns, s)
# AQR factors
for factor, data in [("QMJ", qmj), ("BAB", bab)]:
if data is not None and factor in factor_stats:
data_pd = data.to_pandas().set_index("timestamp")
col = "USA" if "USA" in data_pd.columns else data_pd.columns[0]
returns = data_pd[col].dropna()
s = factor_stats[factor]
append_summary_row(summary_stats, factor, "AQR", returns, s)
summary_df = pd.DataFrame(summary_stats).set_index("Factor")
label_positions = {
"Market (Mkt-RF)": "middle left",
"Momentum (MOM)": "top center",
"Value (HML)": "top center",
"Size (SMB)": "bottom center",
"Profitability (RMW)": "top right",
"Investment (CMA)": "bottom right",
"Quality (QMJ)": "top right",
"Low-Vol (BAB)": "top center",
}
```
A labeled risk-return map replaces the redundant summary dump. Comparisons remain
historical and descriptive because factor histories begin on different dates.
```python
fig = go.Figure(
go.Scatter(
x=summary_df["Ann. Volatility"] * 100,
y=summary_df["Ann. Return"] * 100,
mode="markers+text",
text=summary_df.index,
textposition=[label_positions[label] for label in summary_df.index],
marker=dict(
size=12,
color=[
COLORS["blue"] if value > DISCOVERY_T_THRESHOLD else COLORS["neutral"]
for value in summary_df["t-statistic"]
],
),
customdata=summary_df[["Sharpe Ratio", "t-statistic", "Max Drawdown", "Period"]],
hovertemplate=(
"<b>%{text}</b><br>Annualized return: %{y:.1f}%<br>Annualized volatility: "
"%{x:.1f}%<br>Sharpe: %{customdata[0]:.2f}<br>Mean t (NW): "
"%{customdata[1]:.2f}<br>Max drawdown: %{customdata[2]:.1%}<br>Period: "
"%{customdata[3]}<extra></extra>"
),
)
)
fig.update_layout(
title=(
"Annualized return against volatility for eight published factors"
"<br><sup>Published monthly factor returns; blue markers clear the discovery "
"threshold; "
"sample histories differ</sup>"
),
xaxis_title="Annualized volatility (%)",
yaxis_title="Annualized mean return (%)",
height=500,
margin=dict(l=75, r=75, b=80),
showlegend=False,
)
show_plotly_with_alt(
fig,
"Scatter of eight factors, annualized volatility against annualized mean return, each point labelled, coloured by whether its t-statistic clears the discovery threshold.",
)
```
```python
print(f"Median value-momentum correlation across asset classes: {median_vme_corr:+.2f}")
print(
f"Factor means clearing the discovery threshold: "
f"{int(stats_df['significant_harvey'].sum())} of {len(stats_df)}"
)
print(f"Crisis windows in which trend following was positive: {tsmom_positive} of {tsmom_observed}")
```
### What this run produced
Three counts, all printed above the chart that produced them. The first is how many of the
measured factors clear the conventional t-statistic of 2 on their autocorrelation-adjusted mean
monthly return; the second is how many clear Harvey's discovery threshold of 3; the difference
between them is what raising the bar for multiple testing costs on this set. The third is the
number of hand-picked crisis windows in which trend following returned a positive number.
None of the three is a test of anything prospective. Every series is a published research
portfolio selected into the literature for having cleared a bar, and the crisis windows were
chosen because they were crises.
---
## Key takeaways
1. **Two factors that are negatively correlated for a structural reason are worth more than two
with high individual Sharpe ratios.** Value buys what has fallen and momentum buys what has
risen, so the same piece of news moves an asset into one book and out of the other. That is a
mechanism, not a correlation estimate, and it is the reason to expect the relationship to
hold outside the sample it was measured on.
2. **The bar for believing a published factor is higher than the bar for publishing one.**
Hundreds of candidates have been tested and the surviving ones were selected for clearing a
threshold, so judging them against the threshold they were selected on is circular. Raising
the bar to Harvey's discovery threshold is what tests that, and it disqualifies fewer of
these than of the literature at large - which is expected, because this set is the
most-replicated factors in the field rather than a random draw from what has been published.
3. **A long history is not the same as strong evidence.** Every series here is current-vintage
published research, revisable and backfillable by its provider, and none of it is net of
financing, turnover or the capacity limits a real allocation would meet.
4. **Crisis windows chosen after the fact cannot test insurance.** Trend following was positive
in most of the windows shown, and the windows were picked because they were crises. That is
worth knowing and it is not a prospective claim about the next one.
5. **A factor weakening after publication is a decay to document, not a cause to assert.** The
size premium is weaker after its publication date on this split. Whether publication caused
it, whether the market changed, or whether the original estimate was lucky are three
explanations this evidence does not separate.
### Known limitations
- Every series is a published long-short research portfolio, gross of implementation. Financing,
shorting costs, turnover and capacity all come out of these numbers before an allocator sees
them, and none is measured here.
- Histories differ by factor, so a comparison across the whole set is either restricted to the
shortest window or is comparing statistics computed on different samples. Both appear, and
which one a figure uses is stated on it.
- The crisis windows and the pre- and post-publication split dates are chosen with knowledge of
what happened. They describe; they do not test.
**Next:** [`06_hierarchical_risk_parity`](06_hierarchical_risk_parity.ipynb) builds an allocator
that reads the correlation structure above as a tree rather than inverting it. Section 17.4
develops baseline allocators and factor diversification.
## References
- Asness, Moskowitz & Pedersen (2013), "Value and Momentum Everywhere"
- Harvey, Liu & Zhu (2016), "...and the Cross-Section of Expected Returns"
- Ilmanen et al. (2021), "How Do Factor Premia Vary Over Time?"
- Moskowitz, Ooi & Pedersen (2012), "Time Series Momentum"
- Lo (2002), "The Statistics of Sharpe Ratios"
Full bibliography and Data Library / AQR licensing details live in the
chapter prose.







Reproduit dans son intégralité avec attribution, conformément à la licence de la source. Licence: MIT
Ce résumé a été rédigé par l’agent de recherche de Stratmill à partir de la source originale ; il n’en est pas une copie.