Evaluating Portfolio Returns, Drawdowns, Benchmarks, and Stress Periods
Summary
This notebook presents a framework for assessing an ETF allocation backtest using daily portfolio returns and a matched SPY benchmark. It covers risk and return measures such as Sharpe, Sortino, Calmar, and maximum drawdown; rolling metrics; drawdown recovery; benchmark-relative alpha, beta, tracking error, and information ratio; and performance in selected stress periods. Returns come from a registered, cost-aware backtest, with execution assumptions inherited from that artifact rather than deducted again.
The strategy is selected as the highest validation-Sharpe allocator in the case-study registry, so the analysis describes a selected backtest rather than an unbiased estimate of future performance. Interpretation depends on choices such as the risk-free rate and annualization frequency; a zero cash rate can flatter excess-return ratios. The notebook also cautions that one strategy and one history provide no confidence intervals, hand-picked stress windows are retrospective, and empirical VaR and CVaR cannot project losses beyond the observed sample.
Key ideas
- Sharpe alone omits drawdown depth and duration and penalizes upside volatility, so it should be read with complementary measures.
- Rolling performance reveals periods of weakness that a full-sample average can conceal.
- Alpha, beta, tracking error, information ratio, and capture ratios describe different aspects of benchmark-relative returns.
- Selected stress windows show how a strategy behaved in those episodes but do not establish general stress resilience.
- Validation-based selection, risk-free-rate assumptions, and limited history constrain what portfolio statistics can establish.
Tags
Full text
# Portfolio Performance Analysis
# Portfolio Performance Analysis
**Docker image**: `ml4t`
`ml4t-diagnostic` computes the risk-return metrics, drawdown episodes and
benchmark-relative statistics a portfolio report is built from, and replaces the
unmaintained pyfolio. The strategy
under analysis is the highest-validation-Sharpe ETF allocation backtest as recorded in
`case_studies/etfs/run_log/registry.db` - resolved at runtime via
`resolve_best_backtest_runs(...)` so the metrics always reflect the current
best-Sharpe allocator on validation, not a baked-in hash. Daily portfolio
returns are loaded from the case study's run-log, so every metric below traces
to a registered, cost-aware engine backtest on real ETF data. The validation
ranking chooses the artifact; the displayed return series is a strategy diagnostic,
not a fresh unbiased model-selection estimate.
**Learning Objectives**:
- Compute summary statistics (Sharpe, Sortino, Calmar, max drawdown)
- Analyze rolling performance metrics over configurable windows
- Create drawdown analysis with recovery period tracking
- Compare strategy performance against benchmarks (alpha, beta, IR)
- Perform event analysis during market stress periods
**Book Reference**: Chapter 17, Section 17.3 (Portfolio evaluation metrics)
**Prerequisites**: The ETF case study (`case_studies/etfs/`) must have been run
to produce `run_log/backtest/<hash>/daily_returns.parquet`.
```python
"""Compute risk-return metrics, drawdowns, and stress-period analysis for ETF portfolios."""
import hashlib
import json
import numpy as np
import pandas as pd
import plotly.express as px
import plotly.graph_objects as go
import polars as pl
# ml4t-diagnostic imports
from ml4t.diagnostic.evaluation import PortfolioAnalysis
from ml4t.diagnostic.evaluation.portfolio_analysis import (
alpha_beta,
annual_return,
annual_volatility,
calmar_ratio,
conditional_var,
information_ratio,
max_drawdown,
omega_ratio,
stability_of_timeseries,
value_at_risk,
)
from ml4t.diagnostic.metrics import sharpe_ratio, sortino_ratio
from ml4t.diagnostic.visualization import (
combine_figures_to_html,
create_portfolio_dashboard,
)
from case_studies.utils.registry.queries import resolve_best_backtest_runs
from data import load_etfs
from utils.paths import get_case_study_dir, get_output_dir
from utils.reproducibility import set_global_seeds
from utils.style import COLORS, show_plotly_with_alt
```
Five settings decide what is analysed. `BACKTEST_HASH` left unset means the notebook picks the
strategy from the case study's registry rather than naming one: it takes the allocation-stage
backtest with the highest validation Sharpe on the 21-session forward-return target, so the
analysis follows whatever that case study currently produces instead of a hash that goes stale
the next time it is rebuilt. Setting it to a twelve-character hash pins one instead.
`BENCHMARK_SYMBOL` is what alpha, beta and the capture ratios are measured against, and it has
to be something the strategy could plausibly have been held instead of - SPY, a broad US equity
fund, for a strategy allocating across US-listed ETFs. `ETF_LABEL` names the target the
case study's models predicted, which is what makes the registry lookup unambiguous when a case
study fits more than one.
```python
BACKTEST_HASH = None
BENCHMARK_SYMBOL = "SPY"
ETF_CASE_STUDY = "etfs"
ETF_LABEL = "fwd_ret_21d"
SEED = 42
```
```python
set_global_seeds(SEED)
OUTPUT_DIR = get_output_dir(17, "portfolio_metrics")
OUTPUT_DIR.mkdir(parents=True, exist_ok=True)
```
## 1. Strategy and Benchmark Returns
We load the daily portfolio return series from a registered ETF case-study
allocation backtest and use SPY as the benchmark. Rather than hard-code a
`prediction_hash` or `backtest_hash`, the notebook queries
`registry.db` via `resolve_best_backtest_runs()` and selects the allocation-stage
run with the highest validation Sharpe. The strategy is therefore whatever
the current case-study pipeline ranks first among ETF allocators on real OOS
data; see `case_studies/etfs/run_log/backtest/<hash>/spec.json` for its full
specification. The backtest artifact already includes its registered commission,
slippage, and next-bar execution assumptions; this notebook does not subtract costs again.
```python
# Record the immutable registry before selecting its registered artifact.
registry_path = get_case_study_dir(ETF_CASE_STUDY) / "run_log" / "registry.db"
registry_sha256 = hashlib.sha256(registry_path.read_bytes()).hexdigest()
print(f"ETF registry SHA-256: {registry_sha256}")
# Resolve the backtest hash from the registry unless one was injected by Papermill.
if BACKTEST_HASH is None:
best_runs = resolve_best_backtest_runs(
case_study=ETF_CASE_STUDY,
label=ETF_LABEL,
split="validation",
stage="allocation",
top_n=1,
)
if best_runs.is_empty():
raise RuntimeError(
f"No allocation-stage backtests found for {ETF_CASE_STUDY}/{ETF_LABEL} on "
"validation. Run the ETF case study (Ch16/17 stages) before this notebook."
)
best_row = best_runs.row(0, named=True)
BACKTEST_HASH = best_row["backtest_hash"]
backtest_spec = json.loads(best_row["spec_json"])
execution = backtest_spec["backtest_config"]["execution"]
commission = backtest_spec["backtest_config"]["commission"]
slippage = backtest_spec["backtest_config"]["slippage"]
allocator = backtest_spec["strategy"]["allocation"]
if execution["execution_mode"] != "next_bar":
raise RuntimeError("Selected backtest is not a causal next-bar execution artifact")
print(
f"Resolved highest-validation-Sharpe allocation backtest: {BACKTEST_HASH} "
f"(prediction={best_row['prediction_hash']}, Sharpe={best_row['sharpe']:.3f})"
)
print(
f"Registered assumptions: allocator={allocator['method']}, top_k={allocator['top_k']}, "
f"execution={execution['execution_mode']}/{execution['execution_price']}, "
f"commission={commission['model']}, slippage={slippage['model']}"
)
else:
print(f"Using user-pinned BACKTEST_HASH={BACKTEST_HASH}")
```
```python
# Load strategy daily returns from the registered backtest
backtest_dir = get_case_study_dir(ETF_CASE_STUDY) / "run_log" / "backtest" / BACKTEST_HASH
returns_path = backtest_dir / "daily_returns.parquet"
if not returns_path.exists():
raise FileNotFoundError(
f"Backtest daily returns not found at {returns_path}. Re-run the ETF case "
"study (see case_studies/etfs/) before this notebook."
)
strategy_df = (
pl.read_parquet(returns_path)
.rename({"daily_return": "strategy"})
.with_columns(pl.col("timestamp").cast(pl.Date))
.sort("timestamp")
)
# Trim the leading pre-trade window where daily_return == 0 (before the first rebalance)
first_active = strategy_df.filter(pl.col("strategy") != 0)["timestamp"].min()
if first_active is None:
raise RuntimeError(f"Registered backtest {BACKTEST_HASH} contains no active strategy returns")
strategy_df = strategy_df.filter(pl.col("timestamp") >= first_active)
strategy_returns = strategy_df.to_pandas().set_index("timestamp")["strategy"].astype(float)
strategy_returns.index = pd.to_datetime(strategy_returns.index)
```
```python
# Load SPY benchmark over the same window
spy_lo, spy_hi = strategy_returns.index.min(), strategy_returns.index.max()
spy_data = (
load_etfs()
.filter(
(pl.col("symbol") == BENCHMARK_SYMBOL)
& (pl.col("timestamp") >= pl.lit(str(spy_lo.date())).str.to_datetime())
& (pl.col("timestamp") <= pl.lit(str(spy_hi.date())).str.to_datetime())
)
.sort("timestamp")
)
spy_prices = spy_data.select(["timestamp", "close"]).to_pandas().set_index("timestamp")["close"]
spy_returns = spy_prices.pct_change().dropna()
# Align both series on the intersection of dates
common = strategy_returns.index.intersection(spy_returns.index)
strategy_returns = strategy_returns.loc[common].rename("Strategy")
spy_returns = spy_returns.loc[common]
print(
f"Strategy: {len(strategy_returns):,} days ({strategy_returns.index[0].date()} → {strategy_returns.index[-1].date()})"
)
print(f"Benchmark: {BENCHMARK_SYMBOL}, {len(spy_returns):,} days")
```
## 2. Setting up the analysis
Every metric below is computed from the same three inputs - the strategy's daily returns, the
benchmark's, and the dates they fall on - so they are bound once here rather than passed
separately to each function. Two of the arguments are choices rather than data.
The **risk-free rate** is what gets subtracted before a ratio is taken. It is set to zero here,
which makes every Sharpe and Sortino below an excess-return-over-cash figure only to the extent
that cash paid nothing; over a window covering 2022-23 it did not, so these ratios are
marginally flattering. Applying the same rate to the benchmark does **not** leave the comparison
unchanged. Sharpe is $(\mu - r_f)/\sigma$, so raising the rate by $\Delta$ costs each series
$\Delta/\sigma$ - and that is *larger* for the series with the smaller volatility. The steadier
of the two loses more of its ratio, which is what can reorder them: a low-volatility series
ranked first at $r_f = 0$ can fall behind a more volatile one once a realistic cash rate is
subtracted from both. Sortino is not invariant either, and it is not a plain division:
`periodic_sortino_ratio` measures downside deviation on returns *after* subtracting the rate, so
raising the rate shrinks the numerator and grows the denominator at the same time. It still only
ever moves the ratio down - that holds for a strategy earning less than cash too, where the
numerator is negative and both effects might have been expected to fight - but by an amount that
is not $\Delta$ over anything, so read the two ratios' sensitivities as different in kind rather
than assuming Sharpe's arithmetic carries over. What a common rate does leave alone is the
information ratio, which is computed on
active returns - the difference between the two series, from which any rate applied to both
cancels before the ratio is taken.
**Periods per year** is what an annualization multiplies by. Daily equity returns are quoted
against 252 trading sessions. Switching to 365 does not move every annualized number by the same
amount, because two different scalings are at work: volatility and the ratios built on it scale
with the square root of the count, so they rise by the square root of the ratio of the two
counts, roughly a fifth, while a compounded
annual return raises one plus the periodic return to the count itself and therefore moves with
the underlying growth rate rather than by a fixed factor.
```python
# Create analysis object
RISK_FREE_RATE = 0.0 # annual; set it to a cash rate to measure excess return over cash
PERIODS_PER_YEAR = 252 # trading sessions in a year, the grid these returns are quoted on
analysis = PortfolioAnalysis(
returns=strategy_returns.values,
benchmark=spy_returns.values,
dates=strategy_returns.index,
risk_free=RISK_FREE_RATE,
periods_per_year=PERIODS_PER_YEAR,
)
print(f"Daily returns analysed: {len(strategy_returns):,} sessions")
print(f"Measured against: {BENCHMARK_SYMBOL}")
print(f"Annual risk-free rate subtracted before every ratio: {RISK_FREE_RATE:.2%}")
print(f"Periods a year used to annualize: {PERIODS_PER_YEAR}")
```
## 3. Summary Statistics
The headline table is the smallest set that does not mislead: a return, the volatility it was
earned at, two ratios of one against the other, the worst loss along the way, and three
benchmark-relative figures. Reporting the return alone says nothing about what was risked to
get it; reporting the Sharpe ratio alone hides how far the path fell before recovering.
```python
# Compute all metrics.
metrics = analysis.compute_summary_stats()
headline_metrics = pd.DataFrame(
{
"Total Return": [f"{metrics.total_return:.2%}"],
"Annual Return": [f"{metrics.annual_return:.2%}"],
"Annual Volatility": [f"{metrics.annual_volatility:.2%}"],
"Sharpe": [f"{metrics.sharpe_ratio:.3f}"],
"Sortino": [f"{metrics.sortino_ratio:.3f}"],
"Max Drawdown": [f"{metrics.max_drawdown:.2%}"],
"Alpha": [f"{metrics.alpha:.2%}" if metrics.alpha is not None else "N/A"],
"Beta": [f"{metrics.beta:.3f}" if metrics.beta is not None else "N/A"],
"Information Ratio": [
f"{metrics.information_ratio:.3f}" if metrics.information_ratio is not None else "N/A"
],
},
index=[BACKTEST_HASH],
)
headline_metrics
```
### What each metric is sensitive to
The nine above answer different questions, and each is blind to something the next one sees.
What separates them is the denominator: what each one divides return by decides which kind of
bad outcome it can register at all.
| Metric | What it divides return by | What it cannot see |
|--------|---------------------------|--------------------|
| **Sharpe ratio** | total volatility | which side the volatility came from - an upside surprise penalises it as much as a loss |
| **Sortino ratio** | volatility of losses only | how long a loss lasted, or how deep the path went |
| **Calmar ratio** | the largest peak-to-trough loss | everything except that one episode |
| **Omega ratio** | the probability-weighted mass of losses | nothing about ordering: the same returns shuffled give the same number |
| **Tail ratio** | the size of the worst losses against the largest gains | the middle of the distribution, which is most of it |
| **Alpha** | nothing - it is a return, net of what the benchmark's moves explain | whether the deviations that produced it were large or small |
| **Information ratio** | the volatility of the deviations from the benchmark | the direction of the market the deviations were taken in |
| **Up capture** | the benchmark's return in its rising periods | anything about falling ones |
| **Down capture** | the benchmark's return in its falling periods | anything about rising ones |
There is no threshold that makes one of these good. A Sharpe of 1 is unremarkable for a
strategy trading at daily frequency and hard to reach for one holding positions for a quarter,
and the same number computed over three years and over thirty means different things because
the estimate's own standard error shrinks with the square root of the sample. What the table is
for is reading them together: a strategy that scores well on all of them is doing something
different from one that scores well on Sharpe alone.
## 4. Rolling Metrics
A single Sharpe ratio over the whole backtest is an average over every market the strategy
traded through. Recomputing it inside a moving window shows whether that average describes a
steady process or two different regimes with a crossover in the middle - which is the question
an investor with a finite horizon is actually asking.
Three windows are used, and the choice is a trade-off between resolution and noise. Twenty-one
sessions is about a calendar month: it responds within weeks and its Sharpe estimate swings
wildly, because a ratio estimated from twenty-one observations has a standard error close to
the ratio itself. Sixty-three sessions is a quarter. Two hundred and fifty-two is a year, which
smooths through most single episodes and therefore reacts to a regime change several months
after it happens. Reading the three together is what separates a change in the strategy from a
short run of luck.
```python
# Compute rolling metrics
ROLLING_WINDOWS = [21, 63, 252] # about a month, a quarter and a year of sessions
rolling = analysis.compute_rolling_metrics(
windows=ROLLING_WINDOWS,
metrics=["sharpe", "volatility", "returns"],
)
print(f"Sharpe, volatility and return recomputed over windows of {rolling.windows} sessions.")
```
```python
# Plot rolling Sharpe ratio
fig = go.Figure()
for window in ROLLING_WINDOWS:
if window in rolling.sharpe:
sharpe_series = rolling.sharpe[window]
fig.add_trace(
go.Scatter(
x=strategy_returns.index,
y=sharpe_series.to_numpy(),
name=f"{window}d Rolling",
opacity=0.8,
)
)
fig.add_hline(y=0, line_dash="dash", line_color=COLORS["neutral"], opacity=0.3)
fig.add_hline(
y=1, line_dash="dot", line_color=COLORS["amber"], opacity=0.3, annotation_text="Sharpe = 1"
)
fig.add_hline(
y=2, line_dash="dot", line_color=COLORS["positive"], opacity=0.3, annotation_text="Sharpe = 2"
)
fig.update_layout(
title="Rolling Sharpe ratio over 21-, 63- and 252-session windows",
xaxis_title="Date",
yaxis_title="Sharpe Ratio",
height=450,
legend=dict(yanchor="top", y=0.99, xanchor="right", x=0.99),
)
show_plotly_with_alt(
fig,
"Three lines of rolling Sharpe ratio against date, over 21, 63 and 252 sessions. The 21-session line swings across the whole vertical range while the 252-session line stays within a narrow band.",
)
```
```python
# Plot rolling volatility (annualized)
fig = go.Figure()
for window in ROLLING_WINDOWS:
if window in rolling.volatility:
vol_series = rolling.volatility[window]
fig.add_trace(
go.Scatter(
x=strategy_returns.index,
y=vol_series.to_numpy() * 100,
name=f"{window}d Rolling",
opacity=0.8,
)
)
fig.update_layout(
title="Annualized rolling volatility over 21-, 63- and 252-session windows",
xaxis_title="Date",
yaxis_title="Volatility (%)",
height=400,
)
show_plotly_with_alt(
fig,
"Three lines of annualized rolling volatility against date, over 21, 63 and 252 sessions, with the shortest window spiking far above the other two during market stress.",
)
```
## 5. Drawdown Analysis
A **drawdown** is the loss from a peak in cumulative value to the lowest point before that peak
is regained, measured as a fraction of the peak. It is the loss an investor who bought at the
worst moment actually lived through, which is why it drives the decision to abandon a strategy
in a way that volatility does not: volatility is symmetric and a drawdown is not.
Three numbers describe one: how deep it went, how long it took to reach the bottom, and how
long it took to climb back to the old peak. The last of the three is the one usually left out,
and it is the one that decides whether a strategy is holdable - the same loss recovered in two
months and recovered in four years is the same number and not the same experience. Both
durations below are counted in trading sessions, not calendar days, because that is the grid
the return series is quoted on.
```python
# Compute drawdown analysis
drawdown = analysis.compute_drawdown_analysis(top_n=5)
deepest = pd.DataFrame(
[
{
"Depth": f"{dd.depth * 100:.2f}%",
"Peak": dd.peak_date.date(),
"Valley": dd.valley_date.date(),
"Recovered": dd.recovery_date.date() if dd.recovery_date else "not yet",
"Peak to valley (sessions)": dd.duration_days,
"Valley to peak (sessions)": dd.recovery_days
if dd.recovery_days is not None
else "still under",
}
for dd in drawdown.top_drawdowns
],
index=pd.RangeIndex(1, len(drawdown.top_drawdowns) + 1, name="Rank"),
)
deepest
```
`compute_drawdown_analysis` above returns the episodes; the chart below needs the value on every
date, which is the same quantity evaluated continuously. `DrawdownResult` already carries it as
`underwater_curve`, and it is written out here as well because it is three lines - compound the
returns, carry the running maximum, take the shortfall from it - and seeing the definition once
is what makes every drawdown figure in this chapter readable. The benchmark needs the same
curve, and no analysis object was built for it, so the function earns its place twice over.
Computing it both ways is also the cheapest check available that the definition above is the
one the library uses, so the two are asserted equal rather than assumed to agree.
```python
def compute_drawdown_series(returns: pd.Series) -> pd.Series:
"""Compute drawdown time series."""
cumulative = (1 + returns).cumprod()
running_max = cumulative.expanding().max()
drawdown = (cumulative - running_max) / running_max
return drawdown
dd_series = compute_drawdown_series(strategy_returns)
dd_benchmark = compute_drawdown_series(spy_returns)
np.testing.assert_allclose(
dd_series.to_numpy(), drawdown.underwater_curve.to_numpy(), rtol=0, atol=1e-12
)
print("Hand-derived underwater curve matches DrawdownResult.underwater_curve.")
```
```python
# Plot underwater curve
fig = go.Figure()
fig.add_trace(
go.Scatter(
x=dd_series.index,
y=dd_series * 100,
name="Strategy",
fill="tozeroy",
line=dict(color=COLORS["negative"], width=1),
)
)
fig.add_trace(
go.Scatter(
x=dd_benchmark.index,
y=dd_benchmark * 100,
name="Benchmark (SPY)",
line=dict(color=COLORS["neutral"], width=1, dash="dash"),
)
)
fig.update_layout(
title="Underwater curves for the strategy and the SPY benchmark",
xaxis_title="Date",
yaxis_title="Drawdown (%)",
height=400,
)
show_plotly_with_alt(
fig,
"Two underwater curves against date, the strategy filled to zero and the benchmark dashed, each showing the percentage below its own running peak.",
)
```
```python
# Drawdown distribution
fig = px.histogram(
dd_series * 100,
nbins=50,
title="Distribution of the strategy's daily drawdown",
labels={"value": "Drawdown (%)", "count": "Frequency"},
)
fig.add_vline(
x=metrics.max_drawdown * 100,
line_dash="dash",
line_color=COLORS["negative"],
annotation_text=f"Max DD: {metrics.max_drawdown * 100:.1f}%",
)
fig.update_layout(height=350, showlegend=False)
show_plotly_with_alt(
fig,
"Histogram of the strategy's daily drawdown, concentrated near zero with a thin tail reaching the maximum drawdown marked by a vertical line.",
)
```
## 6. Monthly and Annual Returns
Daily returns compounded to month and year ends answer a question the aggregate figures do
not: how much of the total came from how few periods. A strategy whose annual return is
carried by two months is a different proposition from one earning the same amount steadily,
and the two are indistinguishable in a Sharpe ratio computed over the whole sample.
```python
# Compute monthly returns
monthly = analysis.compute_monthly_returns()
monthly_df = monthly.to_pandas() # Convert Polars to pandas
print("Monthly Return Statistics:")
print(f" Mean: {monthly_df['monthly_return'].mean() * 100:.2f}%")
print(f" Std: {monthly_df['monthly_return'].std() * 100:.2f}%")
print(f" Best: {monthly_df['monthly_return'].max() * 100:.2f}%")
print(f" Worst: {monthly_df['monthly_return'].min() * 100:.2f}%")
```
The heatmap below is the same monthly series as a year-by-month grid. Read down a column to see
whether a calendar month is systematically good or bad, and across a row to see how much of a
year's return came from how few of its months.
```python
monthly_df["year"] = monthly_df["year"].astype(int)
monthly_df["month"] = monthly_df["month"].astype(int)
heatmap_data = (
monthly_df.pivot(index="year", columns="month", values="monthly_return") * 100
) # Convert to percentage
# Month names
month_names = ["Jan", "Feb", "Mar", "Apr", "May", "Jun", "Jul", "Aug", "Sep", "Oct", "Nov", "Dec"]
fig = go.Figure(
data=go.Heatmap(
z=heatmap_data.values,
x=month_names,
y=heatmap_data.index,
colorscale=[
[0.0, COLORS["negative"]],
[0.5, COLORS["silver"]],
[1.0, COLORS["positive"]],
],
zmid=0,
text=np.round(heatmap_data.values, 1),
texttemplate="%{text:.1f}%",
textfont={"size": 10},
colorbar=dict(title="Return %"),
)
)
fig.update_layout(
title="Monthly return by year and calendar month",
xaxis_title="Month",
yaxis_title="Year",
height=400,
)
show_plotly_with_alt(
fig,
"Heatmap of monthly returns, years down the vertical axis and calendar months across, red for losses and green for gains, with each cell labelled.",
)
```
```python
# Annual returns comparison
annual = analysis.compute_annual_returns()
annual_df = annual.to_pandas() # Convert Polars to pandas
# Also compute benchmark annual
spy_annual = spy_returns.groupby(spy_returns.index.year).apply(lambda x: (1 + x).prod() - 1)
# Create comparison
fig = go.Figure()
fig.add_trace(
go.Bar(
x=annual_df["year"],
y=annual_df["annual_return"] * 100,
name="Strategy",
marker_color=COLORS["blue"],
)
)
fig.add_trace(
go.Bar(
x=spy_annual.index,
y=spy_annual.values * 100,
name="Benchmark (SPY)",
marker_color=COLORS["silver_muted"],
)
)
fig.update_layout(
title="Annual return, strategy against the SPY benchmark",
xaxis_title="Year",
yaxis_title="Return (%)",
barmode="group",
height=400,
)
show_plotly_with_alt(
fig,
"Grouped bars of annual return for the strategy and for SPY, one pair per calendar year.",
)
```
## 7. Benchmark-Relative Analysis
Regressing the strategy's daily returns on the benchmark's splits them in two. **Beta** is the
slope: how much the strategy moved, on average, for each unit the market moved. **Alpha** is
the intercept, annualized: the part of the return the market's moves do not explain.
Beta is what decides whether alpha is interesting. A strategy with a beta of one and no alpha
has reproduced the index; one with a beta of one half and no alpha has reproduced half of it
and kept half the capital idle. The two cutoffs below - a fifth below one and a fifth above -
are conventional labels for "materially less exposed than the market" and "materially more",
and there is nothing special about them beyond marking a fifth of the market's own movement
in each direction.
```python
alpha, beta = alpha_beta(
strategy_returns.values, spy_returns.values, periods_per_year=PERIODS_PER_YEAR
)
print(f"Alpha: {alpha * 100:.2f}% (annualized)")
print(f"Beta: {beta:.3f}")
# Interpretation
if beta < 0.8:
beta_interp = "Defensive (low market exposure)"
elif beta > 1.2:
beta_interp = "Aggressive (high market exposure)"
else:
beta_interp = "Neutral (market-like exposure)"
print(f" {beta_interp}")
```
**Tracking error** is the volatility of the difference between the two return series, and the
**information ratio** divides the average of that difference by it. Together they ask whether
the deviations from the benchmark were worth taking: a strategy can exceed its benchmark by a
wide margin through deviations so volatile that the excess is indistinguishable from luck.
Read the information ratio the way a t-statistic is read, because over $T$ years that is what
it is: multiplied by $\sqrt{T}$ it gives roughly the t-statistic of the average active return.
An information ratio of one half sustained over four years is therefore about one standard
error from zero, which is why the number needs a horizon attached before it means anything.
```python
active_returns = strategy_returns.values - spy_returns.values
tracking_error = active_returns.std(ddof=1) * np.sqrt(PERIODS_PER_YEAR)
ir = information_ratio(
strategy_returns.values, spy_returns.values, periods_per_year=PERIODS_PER_YEAR
)
years = len(strategy_returns) / PERIODS_PER_YEAR
print(f"Tracking Error: {tracking_error * 100:.2f}% (annualized)")
print(f"Information Ratio: {ir:.3f} over {years:.1f} years")
print(f"Implied t-statistic on the average active return: {ir * np.sqrt(years):.2f}")
```
**Capture ratios** split the comparison by the direction the benchmark moved. Up capture is
what the strategy returned across the periods the benchmark rose, as a fraction of what the
benchmark returned over those same periods; down capture is the same across the periods it
fell. A strategy capturing four fifths of the upside and half the downside is doing something
a single alpha number cannot express.
They are conventionally quoted on monthly periods rather than daily ones, so both series are
compounded to month ends before the ratio is taken. The frequency is part of the definition:
the same strategy scores differently at daily and monthly frequency, because a month in which
the benchmark ends up while falling for three weeks counts as an up period at one and as a mix
at the other.
The ratio is taken between the two *average* returns over those periods. `ml4t.diagnostic`
exposes an `up_down_capture` that instead divides the two compounded wealth factors, and for
down periods that inverts the reading: a strategy losing half as much as the benchmark in every
month the benchmark falls has captured half the downside, and the compounded form reports more
than the whole of it, because both products are below one and the larger numerator makes the
ratio exceed one. A down capture above one is supposed to mean the strategy fell harder than
the market. Take the ratio of the means.
```python
strat_monthly = strategy_returns.resample("ME").apply(lambda x: (1 + x).prod() - 1)
bench_monthly = spy_returns.resample("ME").apply(lambda x: (1 + x).prod() - 1)
up_months, down_months = bench_monthly > 0, bench_monthly < 0
up_capture = strat_monthly[up_months].mean() / bench_monthly[up_months].mean()
down_capture = strat_monthly[down_months].mean() / bench_monthly[down_months].mean()
print("Capture ratios, on monthly periods:")
print(f" Up Capture: {up_capture * 100:.1f}%")
print(f" Down Capture: {down_capture * 100:.1f}%")
# The gap between the two is the asymmetry: positive means the strategy kept more of the
# benchmark's up months than of its down months.
capture_spread = up_capture - down_capture
print(f" Spread (up - down): {capture_spread * 100:.1f}pp")
```
```python
# Rolling Beta
rolling_beta = pd.Series(index=strategy_returns.index, dtype=float)
window = PERIODS_PER_YEAR # one year of sessions
for i in range(window, len(strategy_returns)):
strat_window = strategy_returns.iloc[i - window : i].values
bench_window = spy_returns.iloc[i - window : i].values
_, beta_i = alpha_beta(strat_window, bench_window)
rolling_beta.iloc[i] = beta_i
fig = go.Figure()
fig.add_trace(
go.Scatter(
x=rolling_beta.index,
y=rolling_beta,
name="Rolling 1Y Beta",
line=dict(color=COLORS["copper"], width=2),
)
)
fig.add_hline(
y=1, line_dash="dash", line_color=COLORS["neutral"], annotation_text="Beta = 1 (Market)"
)
fig.add_hline(y=0, line_dash="dot", line_color=COLORS["neutral"], opacity=0.3)
fig.update_layout(
title="Rolling one-year beta against the SPY benchmark",
xaxis_title="Date",
yaxis_title="Beta",
height=400,
)
show_plotly_with_alt(
fig,
"Rolling one-year beta against SPY plotted against date, with reference lines at one and zero.",
)
```
## 8. Risk Metrics (VaR, CVaR)
**Value at Risk** at a confidence level is the loss that a stated fraction of periods exceeded.
Read at one day in twenty, it is the daily loss that one trading day in twenty was worse than;
read at one day in a hundred, it is the loss the worst day in a hundred exceeded. Both levels
below are empirical quantiles of the return series itself: the returns are sorted and the
quantile is taken, with no distribution assumed.
```python
# Value at Risk
var_95 = value_at_risk(strategy_returns.values, confidence=0.95)
var_99 = value_at_risk(strategy_returns.values, confidence=0.99)
print("Value at Risk, from the sample's own return distribution:")
print(f" 95% VaR: {var_95 * 100:.2f}% (daily)")
print(f" 99% VaR: {var_99 * 100:.2f}% (daily)")
breaches = int((strategy_returns.values < var_95).sum())
print(f" Days worse than the VaR: {breaches:,} of {len(strategy_returns):,}")
```
VaR says where the tail begins and nothing about what is inside it: a strategy whose worst day
loses a few percent and one whose worst day loses ten times that can share the same VaR.
**Conditional VaR**, also called expected shortfall, is the average of the losses that did
exceed the threshold, so it is the number that separates them. It is the reason a risk report
quotes both.
Both are read off this sample's own history. Neither is a forecast, and the one-day-in-a-hundred
quantile of a few thousand observations rests on a few dozen of them.
```python
cvar_95 = conditional_var(strategy_returns.values, confidence=0.95)
cvar_99 = conditional_var(strategy_returns.values, confidence=0.99)
print("Conditional VaR, the average loss on the days that breached VaR:")
print(f" 95% CVaR: {cvar_95 * 100:.2f}% (daily)")
print(f" 99% CVaR: {cvar_99 * 100:.2f}% (daily)")
```
```python
# Visualize return distribution with VaR
fig = go.Figure()
# Histogram of returns
fig.add_trace(go.Histogram(x=strategy_returns * 100, nbinsx=100, name="Daily Returns", opacity=0.7))
# VaR lines - stagger annotation y-positions to prevent overlap at the top
fig.add_vline(
x=var_95 * 100,
line_dash="dash",
line_color=COLORS["amber"],
annotation_text="95% VaR",
annotation_position="top right",
)
fig.add_vline(
x=var_99 * 100,
line_dash="dash",
line_color=COLORS["negative"],
annotation_text="99% VaR",
annotation_position="bottom right",
)
fig.add_vline(
x=cvar_95 * 100,
line_dash="dot",
line_color=COLORS["copper"],
annotation_text="95% CVaR",
annotation_position="top left",
)
fig.update_layout(
title="Daily return distribution with value at risk and conditional VaR",
xaxis_title="Daily Return (%)",
yaxis_title="Frequency",
height=400,
showlegend=False,
)
show_plotly_with_alt(
fig,
"Histogram of daily strategy returns with vertical lines marking the value at risk at two "
"confidence levels and the conditional value at risk further into the left tail.",
)
```
## 9. Event Analysis
Aggregate statistics average over every market the strategy traded through, so they say
nothing about the episodes an investor remembers. Cumulating the strategy and the benchmark
over five named windows - two crashes, two recoveries and one credit event - shows whether the
defensive profile the beta and capture ratios suggest actually held when it was tested.
The five windows are chosen after the fact, which is what makes this a description rather than
a test. Each is bounded by the dates the episode is conventionally dated to; the strategy's
behaviour outside them is what every other section measures.
```python
# Define market stress periods
STRESS_PERIODS = {
"COVID Crash (2020)": ("2020-02-19", "2020-03-23"),
"COVID Recovery": ("2020-03-23", "2020-08-31"),
"2022 Bear Market": ("2022-01-03", "2022-10-12"),
"2023 Banking Crisis": ("2023-03-01", "2023-03-31"),
"2023 Rally": ("2023-10-27", "2023-12-29"),
}
def analyze_period(returns: pd.Series, benchmark: pd.Series, start: str, end: str):
"""Analyze returns over a specific period."""
period_ret = returns.loc[start:end]
period_bench = benchmark.loc[start:end]
if len(period_ret) == 0:
return None
cum_ret = (1 + period_ret).prod() - 1
cum_bench = (1 + period_bench).prod() - 1
excess = cum_ret - cum_bench
return {
"strategy_return": cum_ret,
"benchmark_return": cum_bench,
"excess_return": excess,
"days": len(period_ret),
}
```
```python
period_results = []
for name, (start, end) in STRESS_PERIODS.items():
result = analyze_period(strategy_returns, spy_returns, start, end)
if result:
result["period"] = name
period_results.append(result)
print(f"\n{name}:")
print(f" Strategy: {result['strategy_return'] * 100:+.2f}%")
print(f" Benchmark: {result['benchmark_return'] * 100:+.2f}%")
print(f" Excess: {result['excess_return'] * 100:+.2f}%")
```
```python
# Visualize stress period performance
if period_results:
period_df = pd.DataFrame(period_results)
fig = go.Figure()
fig.add_trace(
go.Bar(
x=period_df["period"],
y=period_df["strategy_return"] * 100,
name="Strategy",
marker_color=COLORS["blue"],
)
)
fig.add_trace(
go.Bar(
x=period_df["period"],
y=period_df["benchmark_return"] * 100,
name="Benchmark",
marker_color=COLORS["silver_muted"],
)
)
fig.update_layout(
title="Cumulative return over five stress and recovery windows",
yaxis_title="Return (%)",
barmode="group",
height=450,
)
show_plotly_with_alt(
fig,
"Grouped bars of cumulative return for the strategy and the benchmark over five named stress and recovery windows.",
)
```
## 10. Stability Analysis
**Stability** here is the R-squared of a straight line fitted to the cumulative value curve
against time. A value near one means the curve looks like a line, which is what a strategy
compounding at a steady rate produces; a lower value means the same total return arrived in
bursts. It is a description of the path, not of the return: a strategy that doubles in one
month and flatlines for four years scores badly and still doubled.
A high R-squared does not mean the path was comfortable. A straight line fitted to eight years
of compounding absorbs a drawdown lasting several months without much loss of fit, so read
this number against the drawdown table in section 5 rather than instead of it.
```python
# Stability of returns (R² of cumulative returns vs time)
stability = stability_of_timeseries(strategy_returns.values)
print(f"Stability (R²): {stability:.3f}")
```
```python
# Cumulative returns with trend line
cumulative = (1 + strategy_returns).cumprod()
# Fit trend line
x = np.arange(len(cumulative))
coeffs = np.polyfit(x, cumulative, 1)
trend = np.polyval(coeffs, x)
fig = go.Figure()
fig.add_trace(
go.Scatter(
x=cumulative.index,
y=cumulative,
name="Cumulative Return",
line=dict(color=COLORS["blue"], width=2),
)
)
fig.add_trace(
go.Scatter(
x=cumulative.index,
y=trend,
name=f"Trend (R² = {stability:.3f})",
line=dict(color=COLORS["negative"], dash="dash"),
)
)
fig.update_layout(
title="Cumulative return with a fitted linear trend",
xaxis_title="Date",
yaxis_title="Cumulative Return",
height=400,
)
show_plotly_with_alt(
fig,
"Cumulative return against date with a fitted straight trend line overlaid, the two diverging where the path deviates from steady compounding.",
)
```
## 11. Full Performance Report
The sections above computed each family of statistics next to the reasoning that motivates it.
Collecting them into one table grouped by what they measure - return, risk, the ratio of one
to the other, and the comparison against the benchmark - is the form a report takes when it
goes to someone who did not follow the derivation.
```python
# Build the summary table
summary_data = {"Category": [], "Metric": [], "Value": []}
def add_metric(category: str, name: str, value: str) -> None:
summary_data["Category"].append(category)
summary_data["Metric"].append(name)
summary_data["Value"].append(value)
```
```python
returns_metrics = [
("Total Return", f"{metrics.total_return * 100:.2f}%"),
("Annual Return (CAGR)", f"{metrics.annual_return * 100:.2f}%"),
("Best Month", f"{monthly_df['monthly_return'].max() * 100:.2f}%"),
("Worst Month", f"{monthly_df['monthly_return'].min() * 100:.2f}%"),
]
risk_metrics = [
("Annual Volatility", f"{metrics.annual_volatility * 100:.2f}%"),
("Max Drawdown", f"{metrics.max_drawdown * 100:.2f}%"),
("95% VaR (Daily)", f"{var_95 * 100:.2f}%"),
("95% CVaR (Daily)", f"{cvar_95 * 100:.2f}%"),
]
risk_adjusted_metrics = [
("Sharpe Ratio", f"{metrics.sharpe_ratio:.3f}"),
("Sortino Ratio", f"{metrics.sortino_ratio:.3f}"),
("Calmar Ratio", f"{metrics.calmar_ratio:.3f}"),
("Omega Ratio", f"{metrics.omega_ratio:.3f}"),
]
benchmark_metrics = []
if metrics.alpha is not None:
benchmark_metrics = [
("Alpha", f"{metrics.alpha * 100:.2f}%"),
("Beta", f"{metrics.beta:.3f}"),
("Tracking Error", f"{tracking_error * 100:.2f}%"),
("Information Ratio", f"{metrics.information_ratio:.3f}"),
("Up Capture (monthly)", f"{up_capture * 100:.1f}%"),
("Down Capture (monthly)", f"{down_capture * 100:.1f}%"),
]
other_metrics = [
("Stability (R²)", f"{stability:.3f}"),
("Win Rate", f"{metrics.win_rate * 100:.1f}%"),
("Profit Factor", f"{metrics.profit_factor:.2f}"),
]
```
```python
for name, value in returns_metrics:
add_metric("Returns", name, value)
for name, value in risk_metrics:
add_metric("Risk", name, value)
for name, value in risk_adjusted_metrics:
add_metric("Risk-Adjusted", name, value)
for name, value in benchmark_metrics:
add_metric("Benchmark", name, value)
for name, value in other_metrics:
add_metric("Other", name, value)
```
```python
summary_df = pd.DataFrame(summary_data)
summary_df
```
The four groups answer four questions and none of them answers another's. Read alpha together
with beta, tracking error and the information ratio: a positive alpha earned through deviations
whose own volatility swamps it has not established that the active positions were worth taking.
The two capture ratios then say where the residual profile came from - holding on in rising
markets, or losing less in falling ones.
## 12. The same metrics, one at a time
The class above computes everything from one binding. Each metric is also available as a plain
function over an array of returns, which is what to reach for when checking a single number
somewhere else without setting up an analysis object. The values below are the same ones the
table carries, computed the other way.
```python
# Example: Computing metrics without PortfolioAnalysis class
returns_arr = strategy_returns.values
benchmark_arr = spy_returns.values
print("Standalone Metric Functions:")
print(f" sharpe_ratio(): {sharpe_ratio(returns_arr):.3f}")
print(f" sortino_ratio(): {sortino_ratio(returns_arr):.3f}")
print(f" calmar_ratio(): {calmar_ratio(returns_arr):.3f}")
print(f" omega_ratio(): {omega_ratio(returns_arr):.3f}")
print(f" max_drawdown(): {max_drawdown(returns_arr) * 100:.2f}%")
print(f" annual_return(): {annual_return(returns_arr) * 100:.2f}%")
print(f" annual_volatility(): {annual_volatility(returns_arr) * 100:.2f}%")
```
## 13. Portfolio Dashboard (Pyfolio Replacement)
The library provides `create_portfolio_dashboard()` which generates a complete
tear sheet in a single call. This is the production replacement for pyfolio's
`create_full_tear_sheet()`.
```python
tear_sheet = create_portfolio_dashboard(
analysis,
theme="default",
include_benchmark=True,
height_per_row=350,
)
print("Tear sheet generated.")
print(f" Figures included: {list(tear_sheet.figures.keys())}")
```
### Reading what the dashboard computed
The dashboard carries its own metrics object. The figures below come from it, so printing its
numbers first is what says the two are describing the same series - they are the values already
computed in section 3, reached by a different route.
```python
ts = tear_sheet.metrics
print(
f"Sharpe: {ts.sharpe_ratio:.3f} | Sortino: {ts.sortino_ratio:.3f} | Calmar: {ts.calmar_ratio:.3f}"
)
print(f"Annual Return: {ts.annual_return * 100:.2f}% | Max DD: {ts.max_drawdown * 100:.2f}%")
print(f"Alpha: {ts.alpha * 100:.2f}% | Beta: {ts.beta:.3f} | IR: {ts.information_ratio:.3f}")
```
`tear_sheet.show()` renders every figure in sequence. Pulling three out by name instead is what
to do when a report needs a selection, and it is also where each figure's title and its alt
text for a screen reader are set, neither of which the dashboard can know from the returns
alone.
```python
dashboard_titles = {
"Cumulative Returns": "Cumulative return, strategy against the SPY benchmark",
"Drawdown": "Strategy underwater curve",
"Monthly Returns Heatmap": "Monthly return by year and calendar month",
}
dashboard_alt = {
"Cumulative Returns": (
"Two cumulative return paths against date, the strategy and the SPY benchmark, "
"both compounding upward and separating over the sample."
),
"Drawdown": (
"The strategy's underwater curve against date, plotted as the percentage below the "
"running peak, with its deepest fall during the 2020 crash."
),
"Monthly Returns Heatmap": (
"Heatmap of monthly return, years down the vertical axis and calendar months "
"across, coloured red for losses and green for gains."
),
}
for name, title in dashboard_titles.items():
fig = tear_sheet.figures[name]
fig.update_layout(
title=title,
paper_bgcolor=COLORS["bg_light"],
plot_bgcolor=COLORS["bg_light"],
)
show_plotly_with_alt(fig, dashboard_alt[name])
```
### Saving the dashboard as a file
`save_html` writes every figure and its metrics into one file that opens in a browser with no
Python installed. `include_plotlyjs="cdn"` keeps the file small by loading the plotting library
from the network when it opens, which is the right trade unless the file has to work offline.
```python
# Export to HTML (self-contained, shareable file)
html_path = OUTPUT_DIR / "portfolio_dashboard.html"
tear_sheet.save_html(html_path, include_plotlyjs="cdn")
print(f"Dashboard saved to: {html_path.name}")
```
## 14. Assembling a custom report
The dashboard above is fixed: it decides which figures appear and in what order. A report going
to someone who is not reading the notebook usually needs a different selection and a sentence
beside each figure saying what to look at. `combine_figures_to_html` takes the figures already
built here plus that narrative and writes one self-contained file.
```python
# Collect figures we created earlier for a custom report
custom_figures = []
custom_sections = []
```
The first custom section focuses on cumulative growth against the benchmark.
```python
# Add cumulative returns figure
fig_cum = go.Figure()
fig_cum.add_trace(
go.Scatter(
x=strategy_returns.index,
y=(1 + strategy_returns).cumprod(),
name="Strategy",
line=dict(color=COLORS["blue"], width=2),
)
)
fig_cum.add_trace(
go.Scatter(
x=spy_returns.index,
y=(1 + spy_returns).cumprod(),
name="Benchmark (SPY)",
line=dict(color=COLORS["neutral"], width=1, dash="dash"),
)
)
fig_cum.update_layout(
title="Cumulative return, strategy against the SPY benchmark",
xaxis_title="Date",
yaxis_title="Cumulative Return",
height=400,
)
custom_figures.append(fig_cum)
custom_sections.append(
{
"title": "Strategy Performance",
"text": f"The strategy achieved a total return of {metrics.total_return * 100:.1f}% "
f"vs benchmark return of {((1 + spy_returns).prod() - 1) * 100:.1f}%. "
f"Sharpe ratio: {metrics.sharpe_ratio:.2f}.",
"figure_index": 0,
}
)
```
The second section summarizes drawdown severity and recovery behavior.
```python
# Add drawdown figure
fig_dd = go.Figure()
fig_dd.add_trace(
go.Scatter(
x=dd_series.index,
y=dd_series * 100,
fill="tozeroy",
line=dict(color=COLORS["negative"], width=1),
name="Drawdown",
)
)
fig_dd.update_layout(
title="Strategy underwater curve",
xaxis_title="Date",
yaxis_title="Drawdown (%)",
height=350,
)
custom_figures.append(fig_dd)
custom_sections.append(
{
"title": "Risk Analysis",
"text": f"Maximum drawdown reached {metrics.max_drawdown * 100:.1f}%. "
f"The strategy has a Calmar ratio of {metrics.calmar_ratio:.2f}.",
"figure_index": 1,
}
)
```
Export the assembled figures and narrative as a standalone HTML report.
```python
# Generate the custom report
report_path = OUTPUT_DIR / "custom_strategy_report.html"
combine_figures_to_html(
figures=custom_figures,
title="Strategy Performance Report",
sections=custom_sections,
output_file=report_path,
theme="default",
include_toc=True,
)
print(f"Custom report saved to: {report_path.name}")
```
### What this run produced
The strategy analysed is whichever allocation backtest the ETF case study currently ranks first
on validation Sharpe; its hash, allocator, cost model and execution convention are printed in
section 1, and every number in this notebook is computed from that one artifact. The summary table in
section 11 is the full set. Three things in it are worth reading together rather than
separately: the annualized return against the annual volatility it was earned at, the maximum
drawdown against the Calmar ratio that divides the return by it, and the alpha against the beta
and tracking error that say how much market exposure and how much active deviation produced it.
The rolling panels in section 4 are what the aggregate figures hide. Where the 252-session
Sharpe and the 21-session Sharpe disagree for an extended stretch, the aggregate is an average
over two different regimes rather than a description of one process.
## Key Takeaways
1. **One ratio cannot describe a return path.** Sharpe divides return by total volatility,
so it registers neither how far the path fell nor how long it stayed down, and it treats
an upside surprise as a cost. Report it alongside the maximum drawdown and the Sortino
ratio, which are blind to different things.
2. **Rolling metrics expose regime dependence.** Aggregate Sharpe can be positive
while rolling windows show extended negative periods - a critical warning for
investors with finite horizons.
3. **Benchmark comparison separates alpha from beta.** Alpha, beta, tracking error,
and information ratio must be read together. Up/down capture reveals asymmetric
exposure that aggregate alpha misses.
4. **Stress-period analysis tests robustness.** The maximum drawdown and event panels
turn an aggregate performance score into the timing and recovery questions that
matter to investors.
### Known limitations
- **One strategy, one benchmark, one sample.** Every number here describes a single backtest
over a single history. None of them carries a confidence interval, and the Sharpe ratio in
particular is estimated with a standard error close to 1/sqrt(years) - so two strategies
differing by a few tenths over a few years have not been separated by this evidence.
- **The strategy was chosen by validation Sharpe.** Ranking a set of candidates and then
reporting the statistics of the one that ranked first overstates them, because the ranking itself used the same
kind of noise the statistic measures. This notebook is a diagnostic of a chosen artifact, not
an unbiased estimate of what that artifact would earn next.
- **Stress periods are named by hand.** The five windows in section 9 were chosen because they
are known episodes, which means the selection is retrospective. They show how the strategy
behaved in them; they do not establish how it behaves in stress generally.
- **VaR and CVaR are empirical quantiles.** Both are read off the sample's own distribution, so
the one-day-in-a-hundred figures rest on a few dozen observations and neither extrapolates
past the worst day in the history.
- **The risk-free rate is zero.** Ratios that subtract it are therefore excess-over-nothing
rather than excess-over-cash, which flatters them over any period when cash paid a return.
**Next**: These metrics are used throughout Ch17-19 to evaluate allocation methods,
transaction cost impact, and risk controls.
**Book**: Section 17.3 discusses the full evaluation framework.












Shown in full with attribution under the source's licence. Licence: MIT
This summary was written by Stratmill's research agent from the original; it is not a copy of the source.