Ajuste de hiperparâmetros do Optuna em três bibliotecas de boosting
Resumo
Este notebook usa Optuna para ajustar XGBoost, LightGBM e CatBoost em um conjunto de dados de características de empresas dividido por período. A busca de cada biblioteca trata a função de perda — erro quadrático médio ou erro absoluto médio — como hiperparâmetro categórico, junto com a complexidade do modelo, a taxa de aprendizado, a amostragem e os parâmetros de regularização. O objetivo é o coeficiente de informação transversal médio nos dados de validação. A amostragem TPE, um podador mediano, a parada antecipada, tentativas com sementes definidas e a persistência no SQLite estruturam a busca; em seguida, o fluxo reconstrói os modelos selecionados e compara o IC de validação e de teste.
O documento relata que MAE teve o melhor desempenho entre as três bibliotecas neste conjunto de dados de retornos com caudas pesadas e apresenta a escolha da função de perda como especialmente influente nesses estudos. Também descreve a execução em GPU como estável entre reexecuções, mas não idêntica bit a bit; para exigir reprodutibilidade exata, sugere CPU e configurações determinísticas. Essas conclusões são específicas do conjunto de dados, da divisão, do espaço de busca e da configuração da execução; elas não demonstram que uma função de perda ou biblioteca terá melhor desempenho em outros alvos. O IC de teste é incluído para comparação após o ajuste, portanto sua interpretação deve levar em conta a seleção de modelos baseada na validação.
Ideias principais
- O protocolo de ajuste compara três bibliotecas de boosting em uma tarefa compartilhada de previsão financeira.
- O tipo de perda é pesquisado junto com os parâmetros do modelo, em vez de ser definido previamente.
- O coeficiente de informação transversal de validação orienta a otimização, enquanto a parada antecipada limita o crescimento das árvores.
- O notebook relata MAE como a melhor escolha de função de perda nesses estudos específicos com retornos de caudas pesadas.
- A execução em GPU pode variar ligeiramente entre rodadas, e as conclusões dependem do conjunto de dados e da configuração da busca.
Tags
Texto completo
# 05_cross_library_hpo.py
```py
# ---
# jupyter:
# jupytext:
# cell_metadata_filter: tags,-all
# text_representation:
# extension: .py
# format_name: percent
# format_version: '1.3'
# jupytext_version: 1.19.3
# kernelspec:
# display_name: Python 3 (ipykernel)
# language: python
# name: python3
# ---
# %% [markdown]
# # Cross-Library HPO: XGBoost vs LightGBM vs CatBoost
#
# **Docker image**: `ml4t-gpu`
#
# **Chapter 12, Section 12.4**: Advanced Hyperparameter Tuning with Optuna
#
# > **GPU recommended**: This notebook trains models with PyTorch/CUDA. It will run on CPU
# > but training may be very slow. For GPU acceleration:
# > ```bash
# > docker compose run --rm ml4t-gpu python 12_gradient_boosting/05_cross_library_hpo.py
# > ```
#
#
# ## Purpose
# Optuna-based HPO for XGBoost, LightGBM, and CatBoost on the Chen-Pelger-Zhu
# (2020) firm characteristics dataset. **Loss type is a hyperparameter**
# within each library's study, treating the choice of MSE vs MAE as a tunable
# decision rather than a fixed prior.
#
# ## Learning Objectives
# - Tune XGBoost, LightGBM, and CatBoost with Optuna under a shared protocol
# - Treat loss function selection as a first-class hyperparameter
# - Compare best-valid and best-test IC after model reconstruction
# - Diagnose which hyperparameters drive gains in noisy financial targets
#
# **Prerequisites**: Requires the firm characteristics dataset and supporting
# loaders from `data/__init__.py`.
#
# ## Key Concepts
# - TPE with configurable warm-up trials for each library
# - MedianPruner for early stopping of unpromising trials
# - SQLite persistence for reproducibility
# - Loss function as a categorical hyperparameter
# - Hyperparameter importance analysis (including loss_type)
#
# ## Cross-References
# - **Section 12.4**: GBM-specific tuning strategy, pruning
# - **Chapter 20**: These results feed into the cross-model synthesis
# - **Related**: `04_optuna_tuning` (ETF case study), `07_hpo_comparison` (grid vs Optuna)
# %% [markdown]
# ## 1. Setup
# %%
"""Cross-Library HPO: tune XGBoost, LightGBM, and CatBoost with Optuna on financial data."""
# Import torch before ml4t.diagnostic. ml4t.diagnostic transitively dlopens
# the older system libcudart.so.12; loading torch first makes its bundled
# CUDA runtime win symbol resolution.
import warnings
import catboost as cb
import lightgbm as lgb
import matplotlib.pyplot as plt
import numpy as np
import optuna
import polars as pl
import torch # noqa: F401
import xgboost as xgb
from optuna.pruners import MedianPruner
from optuna.samplers import TPESampler
from data import load_firm_characteristics
from utils.modeling import cross_sectional_ic_mean
from utils.paths import get_output_dir
from utils.reproducibility import set_global_seeds
from utils.style import COLORS, show_with_alt
# LightGBM records synthetic feature names when fitted on an array with an eval_set,
# and sklearn then warns at every predict on an array that has none to compare. One
# message, not the category: the fit and the predictions are unaffected.
warnings.filterwarnings(
"ignore",
message="X does not have valid feature names",
category=UserWarning,
module="sklearn.utils.validation",
)
optuna.logging.set_verbosity(optuna.logging.WARNING)
# %% tags=["parameters"]
N_TRIALS = 50
N_WARMUP_TRIALS = 5
EARLY_STOPPING_ROUNDS = 50
# 0 = all libraries, 1-3 = first N only
N_LIBRARIES = 0
SEED = 42
# %%
set_global_seeds(SEED)
# %%
OUTPUT_DIR = get_output_dir(12, "us_firm_characteristics")
CATBOOST_LOG_DIR = OUTPUT_DIR / "catboost_info"
OUTPUT_DIR.mkdir(parents=True, exist_ok=True)
# %% [markdown]
# ### Device Detection
#
# XGBoost and CatBoost support GPU acceleration. We detect availability
# at startup rather than hardcoding `device="cuda"`.
# %%
GPU_AVAILABLE = torch.cuda.is_available()
XGB_DEVICE = "cuda" if GPU_AVAILABLE else "cpu"
CB_TASK_TYPE = "GPU" if GPU_AVAILABLE else "CPU"
# CatBoost scores an MAE eval metric every fifth iteration on GPU, which it cannot
# compute on the device; declaring that period keeps the behaviour without the
# per-fit announcement on stderr.
CB_METRIC_PERIOD = 5 if GPU_AVAILABLE else 1
print(f"GPU available: {GPU_AVAILABLE} (XGB: {XGB_DEVICE}, CatBoost: {CB_TASK_TYPE})")
# %% [markdown]
# > **Reproducibility.** This search runs on a GPU for speed. GPU training is not
# > bitwise reproducible: fixed seeds pin each model's random choices but not the
# > order in which parallel hardware sums floating-point values, so the selected
# > configurations and IC values here are stable across reruns rather
# > than identical to the last decimal. For runs that must reproduce exactly, train
# > on CPU with deterministic settings; `02_gbm_comparison` measures the
# > GPU-versus-CPU prediction gap and gives the deterministic CPU recipe in full.
# %% [markdown]
# ## 2. Load Academic Dataset
# %%
train_df = load_firm_characteristics(split="train")
valid_df = load_firm_characteristics(split="valid")
test_df = load_firm_characteristics(split="test")
print(
f"Train: {len(train_df):,} (1967-1989), Valid: {len(valid_df):,} (1990-1999), Test: {len(test_df):,} (2000-2016)"
)
# %%
feature_cols = [c for c in train_df.columns if c not in ["timestamp", "ret", "split"]]
X_train = train_df.select(feature_cols).to_numpy()
y_train = train_df["ret"].to_numpy()
X_valid = valid_df.select(feature_cols).to_numpy()
y_valid = valid_df["ret"].to_numpy()
X_test = test_df.select(feature_cols).to_numpy()
y_test = test_df["ret"].to_numpy()
# Per-date timestamps for cross-sectional IC. Firm IDs are anonymized in this
# dataset, so synthesize per-row entity IDs to make the cross-section join 1:1.
dates_valid = valid_df["timestamp"].to_numpy()
entities_valid = np.arange(len(y_valid))
dates_test = test_df["timestamp"].to_numpy()
entities_test = np.arange(len(y_test))
print(f"Features: {len(feature_cols)}, Shapes: {X_train.shape} / {X_valid.shape} / {X_test.shape}")
# %% [markdown]
# ## 3. Study Configuration
#
# SQLite storage enables persistent studies. The database-backed approach
# also supports parallel workers in production settings (PostgreSQL/MySQL
# for true parallelism; SQLite for single-worker reproducibility).
#
# Persistence has one sharp edge worth flagging: with `load_if_exists=True`,
# re-running the notebook against an existing database *appends* new trials to
# the old study rather than starting over, so the trial counts and best configs
# would creep upward on every execution. To keep the results reproducible we
# delete any prior database at the top of the run, so a fresh execution always
# optimizes exactly `N_TRIALS` seeded trials per library.
# %%
STORAGE_PATH = f"sqlite:///{OUTPUT_DIR / 'hpo_studies.db'}"
# Reset persisted studies so each full run reproduces from scratch (see above).
_db_file = OUTPUT_DIR / "hpo_studies.db"
if _db_file.exists():
_db_file.unlink()
def create_study(study_name: str, direction: str = "maximize") -> optuna.Study:
"""Create or load an Optuna study with TPE sampler and pruning."""
return optuna.create_study(
study_name=study_name,
storage=STORAGE_PATH,
load_if_exists=True,
direction=direction,
sampler=TPESampler(seed=SEED, n_startup_trials=N_WARMUP_TRIALS),
pruner=MedianPruner(n_startup_trials=5, n_warmup_steps=10),
)
all_results = []
# %% [markdown]
# ## 4. Objective Functions with Loss as Hyperparameter
#
# Each library has one study where `loss_type` (MSE/MAE) is a categorical
# parameter. This treats the loss function as a tunable decision rather than
# a fixed choice.
# %%
def make_xgboost_objective():
"""Factory for XGBoost objective with loss as hyperparameter."""
def objective(trial: optuna.Trial) -> float:
loss_type = trial.suggest_categorical("loss_type", ["mse", "mae"])
xgb_obj = "reg:squarederror" if loss_type == "mse" else "reg:absoluteerror"
params = {
"objective": xgb_obj,
"n_estimators": trial.suggest_int("n_estimators", 100, 1000),
"max_depth": trial.suggest_int("max_depth", 2, 10),
"learning_rate": trial.suggest_float("learning_rate", 0.005, 0.3, log=True),
"subsample": trial.suggest_float("subsample", 0.5, 1.0),
"colsample_bytree": trial.suggest_float("colsample_bytree", 0.5, 1.0),
"min_child_weight": trial.suggest_int("min_child_weight", 1, 300),
"reg_alpha": trial.suggest_float("reg_alpha", 1e-5, 100.0, log=True),
"reg_lambda": trial.suggest_float("reg_lambda", 1e-5, 100.0, log=True),
"tree_method": "hist",
"device": XGB_DEVICE,
"early_stopping_rounds": EARLY_STOPPING_ROUNDS,
"random_state": SEED,
"verbosity": 0,
}
model = xgb.XGBRegressor(**params)
model.fit(X_train, y_train, eval_set=[(X_valid, y_valid)], verbose=False)
return cross_sectional_ic_mean(y_valid, model.predict(X_valid), dates_valid, entities_valid)
return objective
# %% [markdown]
# ### LightGBM Objective Factory
# %%
def make_lightgbm_objective():
"""Factory for LightGBM objective with loss as hyperparameter."""
def objective(trial: optuna.Trial) -> float:
loss_type = trial.suggest_categorical("loss_type", ["mse", "mae"])
lgb_obj = "regression" if loss_type == "mse" else "mae"
params = {
"objective": lgb_obj,
"n_estimators": trial.suggest_int("n_estimators", 100, 1000),
"max_depth": trial.suggest_int("max_depth", 2, 10),
"learning_rate": trial.suggest_float("learning_rate", 0.005, 0.3, log=True),
"num_leaves": trial.suggest_int("num_leaves", 8, 128),
"subsample": trial.suggest_float("subsample", 0.5, 1.0),
"colsample_bytree": trial.suggest_float("colsample_bytree", 0.5, 1.0),
"min_child_samples": trial.suggest_int("min_child_samples", 5, 300),
"reg_alpha": trial.suggest_float("reg_alpha", 1e-5, 100.0, log=True),
"reg_lambda": trial.suggest_float("reg_lambda", 1e-5, 100.0, log=True),
"random_state": SEED,
"verbose": -1,
}
model = lgb.LGBMRegressor(**params)
model.fit(
X_train,
y_train,
eval_set=[(X_valid, y_valid)],
callbacks=[lgb.early_stopping(EARLY_STOPPING_ROUNDS, verbose=False)],
)
return cross_sectional_ic_mean(y_valid, model.predict(X_valid), dates_valid, entities_valid)
return objective
# %% [markdown]
# ### CatBoost Objective Factory
# %%
def make_catboost_objective():
"""Factory for CatBoost objective with loss as hyperparameter."""
def objective(trial: optuna.Trial) -> float:
loss_type = trial.suggest_categorical("loss_type", ["mse", "mae"])
cb_loss = "RMSE" if loss_type == "mse" else "MAE"
params = {
"loss_function": cb_loss,
"iterations": trial.suggest_int("iterations", 100, 1000),
"depth": trial.suggest_int("depth", 2, 10),
"learning_rate": trial.suggest_float("learning_rate", 0.005, 0.3, log=True),
"l2_leaf_reg": trial.suggest_float("l2_leaf_reg", 1e-5, 100.0, log=True),
"random_strength": trial.suggest_float("random_strength", 0.0, 10.0),
"bagging_temperature": trial.suggest_float("bagging_temperature", 0.0, 1.0),
"task_type": CB_TASK_TYPE,
"metric_period": CB_METRIC_PERIOD,
"random_seed": SEED,
"verbose": False,
"early_stopping_rounds": EARLY_STOPPING_ROUNDS,
"train_dir": str(CATBOOST_LOG_DIR),
}
model = cb.CatBoostRegressor(**params)
model.fit(X_train, y_train, eval_set=(X_valid, y_valid), verbose=False)
return cross_sectional_ic_mean(y_valid, model.predict(X_valid), dates_valid, entities_valid)
return objective
# %% [markdown]
# ## 5. Model Building Helper
#
# Shared function to reconstruct and evaluate models from Optuna trial
# parameters. This eliminates the ~80 lines of duplicated model-building
# code that would otherwise appear in the optimization and per-loss sections.
# %%
def build_model(lib_code, params):
"""Build model instance from Optuna params and selected library."""
loss_type = params.get("loss_type", "mse")
eval_params = {k: v for k, v in params.items() if k not in ["loss_type"]}
if lib_code == "xgb":
xgb_obj = "reg:squarederror" if loss_type == "mse" else "reg:absoluteerror"
return xgb.XGBRegressor(
**eval_params,
objective=xgb_obj,
tree_method="hist",
device=XGB_DEVICE,
early_stopping_rounds=EARLY_STOPPING_ROUNDS,
random_state=SEED,
verbosity=0,
)
if lib_code == "lgb":
lgb_obj = "regression" if loss_type == "mse" else "mae"
return lgb.LGBMRegressor(
**eval_params,
objective=lgb_obj,
random_state=SEED,
verbose=-1,
)
cb_loss = "RMSE" if loss_type == "mse" else "MAE"
return cb.CatBoostRegressor(
**eval_params,
loss_function=cb_loss,
task_type=CB_TASK_TYPE,
metric_period=CB_METRIC_PERIOD,
early_stopping_rounds=EARLY_STOPPING_ROUNDS,
random_seed=SEED,
verbose=False,
train_dir=str(CATBOOST_LOG_DIR),
)
# %% [markdown]
# ### Evaluate Best-Config Model
# %%
def build_and_evaluate(
lib_code,
params,
X_tr,
y_tr,
X_va,
y_va,
X_te,
y_te,
dates_va,
entities_va,
dates_te,
entities_te,
):
"""Fit model from Optuna params and return (model, valid_ic, test_ic)."""
model = build_model(lib_code, params)
if lib_code == "xgb":
model.fit(X_tr, y_tr, eval_set=[(X_va, y_va)], verbose=False)
elif lib_code == "lgb":
model.fit(
X_tr,
y_tr,
eval_set=[(X_va, y_va)],
callbacks=[lgb.early_stopping(EARLY_STOPPING_ROUNDS, verbose=False)],
)
else:
model.fit(X_tr, y_tr, eval_set=(X_va, y_va), verbose=False)
valid_ic = cross_sectional_ic_mean(y_va, model.predict(X_va), dates_va, entities_va)
test_ic = cross_sectional_ic_mean(y_te, model.predict(X_te), dates_te, entities_te)
return model, valid_ic, test_ic
# %% [markdown]
# ## 6. Run Optimization
# %%
LIBRARY_CONFIGS = [
("XGBoost", "academic_xgboost", make_xgboost_objective(), "xgb"),
("LightGBM", "academic_lightgbm", make_lightgbm_objective(), "lgb"),
("CatBoost", "academic_catboost", make_catboost_objective(), "cb"),
]
if N_LIBRARIES > 0:
LIBRARY_CONFIGS = LIBRARY_CONFIGS[:N_LIBRARIES]
for lib_name, study_name, objective, lib_code in LIBRARY_CONFIGS:
print(f"\nOptimizing {lib_name} ({N_TRIALS} trials)...")
study = create_study(study_name)
study.optimize(objective, n_trials=N_TRIALS, show_progress_bar=True)
_, valid_ic, test_ic = build_and_evaluate(
lib_code,
study.best_params,
X_train,
y_train,
X_valid,
y_valid,
X_test,
y_test,
dates_valid,
entities_valid,
dates_test,
entities_test,
)
loss_type = study.best_params["loss_type"]
print(f" Best: Valid IC={valid_ic:.4f}, Test IC={test_ic:.4f}, Loss={loss_type.upper()}")
all_results.append(
{
"model": lib_name,
"library": lib_code.upper(),
"loss_type": loss_type.upper(),
"valid_ic": round(valid_ic, 4),
"test_ic": round(test_ic, 4),
"n_trials": len(study.trials),
}
)
# %% [markdown]
# ## 7. Per-Loss Analysis
#
# Extract each library's highest-scoring MSE and MAE configuration to
# assess whether loss function choice matters.
# %%
detailed_results = []
for lib_name, study_name, _, lib_code in LIBRARY_CONFIGS:
study = optuna.load_study(study_name=study_name, storage=STORAGE_PATH)
loss_results = {"mse": None, "mae": None}
for trial in study.trials:
if trial.state == optuna.trial.TrialState.COMPLETE:
lt = trial.params["loss_type"]
if loss_results[lt] is None or trial.value > loss_results[lt]["valid_ic"]:
loss_results[lt] = {"valid_ic": trial.value, "params": trial.params}
for lt, info in loss_results.items():
if info is None:
continue
_, valid_ic, test_ic = build_and_evaluate(
lib_code,
info["params"],
X_train,
y_train,
X_valid,
y_valid,
X_test,
y_test,
dates_valid,
entities_valid,
dates_test,
entities_test,
)
detailed_results.append(
{
"model": f"{lib_name}_{lt.upper()}",
"library": lib_code.upper(),
"loss_type": lt.upper(),
"valid_ic": round(valid_ic, 4),
"test_ic": round(test_ic, 4),
}
)
# %% [markdown]
# ## 8. Results Summary
# %%
results_df = pl.DataFrame(all_results).sort("test_ic", descending=True)
results_df.select(["model", "loss_type", "valid_ic", "test_ic", "n_trials"])
# %%
results_df.write_csv(OUTPUT_DIR / "hpo_gbm_results.csv")
# %%
detailed_df = pl.DataFrame(detailed_results).sort("test_ic", descending=True)
detailed_df.select(["model", "loss_type", "valid_ic", "test_ic"])
# %%
detailed_df.write_csv(OUTPUT_DIR / "hpo_gbm_detailed.csv")
# %% [markdown]
# ## 9. MSE vs MAE Comparison
# %%
loss_comparison = []
for lib in detailed_df["library"].unique(maintain_order=True).to_list():
lib_mse = detailed_df.filter((pl.col("library") == lib) & (pl.col("loss_type") == "MSE"))
lib_mae = detailed_df.filter((pl.col("library") == lib) & (pl.col("loss_type") == "MAE"))
if len(lib_mse) > 0 and len(lib_mae) > 0:
mse_ic = lib_mse["test_ic"][0]
mae_ic = lib_mae["test_ic"][0]
loss_comparison.append(
{
"library": lib,
"mse_test_ic": mse_ic,
"mae_test_ic": mae_ic,
"difference": round(mae_ic - mse_ic, 4),
"loss_choice": "MAE" if mae_ic > mse_ic else "MSE",
}
)
loss_comparison_df = pl.DataFrame(
loss_comparison,
schema={
"library": pl.Utf8,
"mse_test_ic": pl.Float64,
"mae_test_ic": pl.Float64,
"difference": pl.Float64,
"loss_choice": pl.Utf8,
},
)
loss_comparison_df
# %% [markdown]
# A library appears below only when its study completed a trial under each loss,
# which a short search need not do: the sampler can spend every trial on one of the
# two. Each bar is the highest-scoring configuration the search found *conditional on*
# that
# loss, scored on the held-out 2000-2016 window. The two therefore differ in depth,
# learning rate and regularization as well as in the loss, and the sampler did not
# spend equal effort on each - the loss distribution printed further down says how
# unequal. So the gap is the difference between how far the search got under each
# loss, which is the practical question, and not the isolated effect of the loss
# function on an otherwise matched model.
# %%
if loss_comparison_df.is_empty():
print(
"No library completed a trial under both losses in this run, so there is no\n"
"MSE-against-MAE pair to draw. The comparison needs one configuration of each."
)
else:
libs = loss_comparison_df["library"].to_list()
mse_ics = loss_comparison_df["mse_test_ic"].to_list()
mae_ics = loss_comparison_df["mae_test_ic"].to_list()
x = np.arange(len(libs))
width = 0.38
fig, ax = plt.subplots(figsize=(8, 5))
bars_mse = ax.bar(x - width / 2, mse_ics, width, label="MSE loss", color=COLORS["slate"])
bars_mae = ax.bar(x + width / 2, mae_ics, width, label="MAE loss", color=COLORS["amber"])
for bars in (bars_mse, bars_mae):
for bar in bars:
ax.text(
bar.get_x() + bar.get_width() / 2,
bar.get_height() + 0.0007,
f"{bar.get_height():.4f}",
ha="center",
fontsize=9,
)
ax.set_xticks(x)
ax.set_xticklabels(libs)
ax.set_xlabel("Library")
ax.set_ylabel("Held-out test IC (2000-2016)")
ax.legend()
ax.set_title("Held-out IC of each library's best MSE and MAE configuration")
show_with_alt(
fig,
"Grouped bars of held-out test information coefficient, one pair per library, "
"comparing its best MSE configuration against its best MAE configuration.",
)
# The count and the average gap are results, so they belong here rather than in
# the title, where a reader would have to check them against the chart.
n_mae = int((loss_comparison_df["loss_choice"] == "MAE").sum())
mean_gain = float(loss_comparison_df["difference"].mean())
print(
f"MAE scores higher on the held-out window for {n_mae} of "
f"{loss_comparison_df.height} libraries, by {mean_gain:+.4f} IC on average."
)
# %% [markdown]
# **What to read off it.** The gap between the two bars of one library is the
# cost of fixing the loss by prior instead of letting the search choose it. MAE
# is the more robust of the two to outliers, which is the reason to expect it to
# matter on a target with fat tails - but expecting it is not the same as
# measuring it, and the sign of the gap is what the run reports. Declaring the
# loss a hyperparameter is what makes the question answerable at all.
# %% [markdown]
# ## 10. Hyperparameter Importance
# %%
importance_records = []
for lib_name, study_name, _, _ in LIBRARY_CONFIGS:
try:
study = optuna.load_study(study_name=study_name, storage=STORAGE_PATH)
if len(study.trials) > 1:
importance = optuna.importance.get_param_importances(study)
for param, imp in list(importance.items())[:6]:
importance_records.append(
{"model": lib_name, "param": param, "importance": round(imp, 4)}
)
except Exception as e:
print(f"Could not analyze {study_name}: {e}")
importance_df = pl.DataFrame(importance_records)
importance_df
# %% [markdown]
# Ranking the fANOVA importances per library puts the loss-type decision at or
# near the top of every panel: the categorical MSE/MAE switch explains more of
# the validation-IC variance than any continuous hyperparameter.
# %%
lib_order = [ln for ln, *_ in LIBRARY_CONFIGS]
fig, axes = plt.subplots(1, len(lib_order), figsize=(4.5 * len(lib_order), 5), sharex=False)
if len(lib_order) == 1:
axes = [axes]
for ax, lib_name in zip(axes, lib_order, strict=False):
sub = importance_df.filter(pl.col("model") == lib_name).sort("importance")
params = sub["param"].to_list()
imps = sub["importance"].to_list()
bar_colors = [COLORS["amber"] if p == "loss_type" else COLORS["slate"] for p in params]
ax.barh(params, imps, color=bar_colors)
for i, imp in enumerate(imps):
ax.text(imp + 0.01, i, f"{imp:.2f}", va="center", fontsize=8)
ax.set_xlim(0, 1.05)
ax.set_xlabel("fANOVA importance")
ax.set_title(lib_name)
fig.suptitle("fANOVA importance by hyperparameter, one panel per library", fontsize=13)
show_with_alt(
fig,
"One horizontal bar panel per library, ranking its hyperparameters by fANOVA "
"importance, with the loss-type bar picked out in gold.",
)
# %%
importance_df.write_csv(OUTPUT_DIR / "hpo_gbm_importance.csv")
# %% [markdown]
# ## 11. Optimization History
# %%
history = {}
best_trials = {}
for lib_name, study_name, _, _ in LIBRARY_CONFIGS:
study = optuna.load_study(study_name=study_name, storage=STORAGE_PATH)
best_num = study.best_trial.number
total = len(study.trials)
best_trials[lib_name] = best_num
loss_counts = {"mse": 0, "mae": 0}
values = []
for trial in study.trials:
if trial.state == optuna.trial.TrialState.COMPLETE:
loss_counts[trial.params["loss_type"]] += 1
values.append(trial.value if trial.value is not None else np.nan)
history[lib_name] = np.array(values, dtype=float)
print(
f"{lib_name}: best at trial #{best_num}/{total}, "
f"loss={study.best_params['loss_type'].upper()}, "
f"distribution: MSE={loss_counts['mse']}, MAE={loss_counts['mae']}"
)
# %% [markdown]
# Tracking the highest validation IC found so far against the trial index shows where
# TPE improved and where it sat still.
# %%
palette = [COLORS["slate"], COLORS["amber"], COLORS["copper"]]
fig, ax = plt.subplots(figsize=(9, 5))
for (lib_name, values), color in zip(history.items(), palette, strict=False):
running_best = np.fmax.accumulate(np.nan_to_num(values, nan=-np.inf))
running_best[running_best == -np.inf] = np.nan
trials_axis = np.arange(len(values))
ax.plot(trials_axis, running_best, color=color, linewidth=2, label=lib_name)
ax.axvline(N_WARMUP_TRIALS, linestyle="--", color="gray", linewidth=0.8, label="End of TPE warm-up")
ax.set_xlabel("Trial number")
ax.set_ylabel("Best validation IC so far")
ax.legend()
ax.set_title("Best validation IC so far, by trial, against the TPE warm-up")
show_with_alt(
fig,
"Best validation information coefficient found so far against trial number, one "
"line per library, with a dashed vertical line at the end of the TPE warm-up.",
)
earliest_late_best = min(best_trials.values())
_late = sum(1 for t in best_trials.values() if t > N_WARMUP_TRIALS)
print(
f"{_late} of {len(best_trials)} libraries found their best configuration after the "
f"{N_WARMUP_TRIALS}-trial warm-up; the earliest any of them found it was trial "
f"{earliest_late_best}."
)
# %%
best_overall = results_df.row(0, named=True)
print(
f"Best GBM: {best_overall['model']} ({best_overall['loss_type']}) IC={best_overall['test_ic']:.4f}"
)
# %% [markdown]
# ## Key Takeaways
#
# 1. **Loss type matters as a hyperparameter.** Putting MSE and MAE in the search
# space lets Optuna discover the better loss for each library. Here the
# categorical loss switch is the single most important hyperparameter in every
# study, and MAE wins across all three libraries on these heavy-tailed returns.
# 2. **Early stopping tames a large search range.** Setting `n_estimators` up to
# 1000 with an early-stopping patience of 50 lets the data pick the tree count
# instead of the search wasting compute on over-grown models.
# 3. **The library matters less than the tuning.** Read the spread of the three
# libraries' test IC in the summary table against the MSE-to-MAE gap in the loss
# comparison: the second is the larger of the two here, so a careful search over
# loss and depth buys more than switching implementations.
# 4. **These GBM results connect to Chapter 20.** The per-library, per-loss IC
# values illustrate the kind of tuned tabular baselines the cross-model
# synthesis in Chapter 20 builds on.
#
# **Next**: See `04_optuna_tuning` for the full Optuna workflow on the ETF case
# study, including walk-forward HPO and pruning effectiveness.
```Exibido na íntegra, com atribuição conforme a licença da fonte. Licença: MIT
Este resumo foi escrito pelo agente de pesquisa da Stratmill com base no original; não é uma cópia da fonte.