Passer au contenu
Tous les documents de la bibliothèque

Optimisation des hyperparamètres avec Optuna et trois bibliothèques de boosting de gradient

Code Machine Learning for Trading

Résumé

Ce notebook utilise Optuna pour optimiser XGBoost, LightGBM et CatBoost sur un jeu de données de caractéristiques d’entreprises réparti dans le temps. Pour chaque bibliothèque, la recherche traite la fonction de perte — erreur quadratique moyenne ou erreur absolue moyenne — comme un hyperparamètre catégoriel, au même titre que la complexité du modèle, le taux d’apprentissage, l’échantillonnage et la régularisation. L’objectif est le coefficient d’information transversal moyen sur les données de validation. L’échantillonnage TPE, un élagage médian, l’arrêt anticipé, des essais avec graine fixée et la persistance SQLite structurent la recherche ; le flux de travail reconstruit ensuite les modèles sélectionnés et compare les IC de validation et de test.

Le document indique que MAE a donné les meilleurs résultats parmi les trois bibliothèques sur ce jeu de données de rendements à queues épaisses, et souligne l’influence particulière du choix de la fonction de perte dans ces études. Il précise également que l’exécution sur GPU est stable d’une répétition à l’autre, sans être identique bit à bit, et suggère CPU ainsi que des paramètres déterministes lorsque la reproductibilité exacte est requise. Ces résultats dépendent du jeu de données, du découpage, de l’espace de recherche et de la configuration d’exécution ; ils ne démontrent pas qu’une fonction de perte ou une bibliothèque sera supérieure pour d’autres cibles. Le IC de test est inclus à titre de comparaison après l’optimisation ; son interprétation doit donc tenir compte de la sélection des modèles fondée sur la validation.

Idées clés

  • Le protocole d’optimisation compare trois bibliothèques de boosting de gradient sur une même tâche de prédiction financière.
  • Le type de fonction de perte est exploré avec les paramètres du modèle, plutôt que fixé à l’avance.
  • Le coefficient d’information transversal de validation guide l’optimisation, tandis que l’arrêt anticipé limite la croissance des arbres.
  • Le notebook indique que MAE est le meilleur choix de fonction de perte dans ces études particulières de rendements à queues épaisses.
  • L’exécution GPU peut légèrement varier d’une relance à l’autre, et les conclusions dépendent du jeu de données et du protocole de recherche.

Étiquettes

Texte intégral
# 05_cross_library_hpo.py


```py
# ---
# jupyter:
#   jupytext:
#     cell_metadata_filter: tags,-all
#     text_representation:
#       extension: .py
#       format_name: percent
#       format_version: '1.3'
#       jupytext_version: 1.19.3
#   kernelspec:
#     display_name: Python 3 (ipykernel)
#     language: python
#     name: python3
# ---

# %% [markdown]
# # Cross-Library HPO: XGBoost vs LightGBM vs CatBoost
#
# **Docker image**: `ml4t-gpu`
#
# **Chapter 12, Section 12.4**: Advanced Hyperparameter Tuning with Optuna
#
# > **GPU recommended**: This notebook trains models with PyTorch/CUDA. It will run on CPU
# > but training may be very slow. For GPU acceleration:
# > ```bash
# > docker compose run --rm ml4t-gpu python 12_gradient_boosting/05_cross_library_hpo.py
# > ```
#
#
# ## Purpose
# Optuna-based HPO for XGBoost, LightGBM, and CatBoost on the Chen-Pelger-Zhu
# (2020) firm characteristics dataset. **Loss type is a hyperparameter**
# within each library's study, treating the choice of MSE vs MAE as a tunable
# decision rather than a fixed prior.
#
# ## Learning Objectives
# - Tune XGBoost, LightGBM, and CatBoost with Optuna under a shared protocol
# - Treat loss function selection as a first-class hyperparameter
# - Compare best-valid and best-test IC after model reconstruction
# - Diagnose which hyperparameters drive gains in noisy financial targets
#
# **Prerequisites**: Requires the firm characteristics dataset and supporting
# loaders from `data/__init__.py`.
#
# ## Key Concepts
# - TPE with configurable warm-up trials for each library
# - MedianPruner for early stopping of unpromising trials
# - SQLite persistence for reproducibility
# - Loss function as a categorical hyperparameter
# - Hyperparameter importance analysis (including loss_type)
#
# ## Cross-References
# - **Section 12.4**: GBM-specific tuning strategy, pruning
# - **Chapter 20**: These results feed into the cross-model synthesis
# - **Related**: `04_optuna_tuning` (ETF case study), `07_hpo_comparison` (grid vs Optuna)

# %% [markdown]
# ## 1. Setup

# %%
"""Cross-Library HPO: tune XGBoost, LightGBM, and CatBoost with Optuna on financial data."""

# Import torch before ml4t.diagnostic. ml4t.diagnostic transitively dlopens
# the older system libcudart.so.12; loading torch first makes its bundled
# CUDA runtime win symbol resolution.
import warnings

import catboost as cb
import lightgbm as lgb
import matplotlib.pyplot as plt
import numpy as np
import optuna
import polars as pl
import torch  # noqa: F401
import xgboost as xgb
from optuna.pruners import MedianPruner
from optuna.samplers import TPESampler

from data import load_firm_characteristics
from utils.modeling import cross_sectional_ic_mean
from utils.paths import get_output_dir
from utils.reproducibility import set_global_seeds
from utils.style import COLORS, show_with_alt

# LightGBM records synthetic feature names when fitted on an array with an eval_set,
# and sklearn then warns at every predict on an array that has none to compare. One
# message, not the category: the fit and the predictions are unaffected.
warnings.filterwarnings(
    "ignore",
    message="X does not have valid feature names",
    category=UserWarning,
    module="sklearn.utils.validation",
)
optuna.logging.set_verbosity(optuna.logging.WARNING)

# %% tags=["parameters"]
N_TRIALS = 50
N_WARMUP_TRIALS = 5
EARLY_STOPPING_ROUNDS = 50
# 0 = all libraries, 1-3 = first N only
N_LIBRARIES = 0
SEED = 42


# %%
set_global_seeds(SEED)
# %%
OUTPUT_DIR = get_output_dir(12, "us_firm_characteristics")
CATBOOST_LOG_DIR = OUTPUT_DIR / "catboost_info"
OUTPUT_DIR.mkdir(parents=True, exist_ok=True)

# %% [markdown]
# ### Device Detection
#
# XGBoost and CatBoost support GPU acceleration. We detect availability
# at startup rather than hardcoding `device="cuda"`.

# %%
GPU_AVAILABLE = torch.cuda.is_available()

XGB_DEVICE = "cuda" if GPU_AVAILABLE else "cpu"
CB_TASK_TYPE = "GPU" if GPU_AVAILABLE else "CPU"
# CatBoost scores an MAE eval metric every fifth iteration on GPU, which it cannot
# compute on the device; declaring that period keeps the behaviour without the
# per-fit announcement on stderr.
CB_METRIC_PERIOD = 5 if GPU_AVAILABLE else 1
print(f"GPU available: {GPU_AVAILABLE} (XGB: {XGB_DEVICE}, CatBoost: {CB_TASK_TYPE})")

# %% [markdown]
# > **Reproducibility.** This search runs on a GPU for speed. GPU training is not
# > bitwise reproducible: fixed seeds pin each model's random choices but not the
# > order in which parallel hardware sums floating-point values, so the selected
# > configurations and IC values here are stable across reruns rather
# > than identical to the last decimal. For runs that must reproduce exactly, train
# > on CPU with deterministic settings; `02_gbm_comparison` measures the
# > GPU-versus-CPU prediction gap and gives the deterministic CPU recipe in full.

# %% [markdown]
# ## 2. Load Academic Dataset

# %%
train_df = load_firm_characteristics(split="train")
valid_df = load_firm_characteristics(split="valid")
test_df = load_firm_characteristics(split="test")

print(
    f"Train: {len(train_df):,} (1967-1989), Valid: {len(valid_df):,} (1990-1999), Test: {len(test_df):,} (2000-2016)"
)

# %%
feature_cols = [c for c in train_df.columns if c not in ["timestamp", "ret", "split"]]

X_train = train_df.select(feature_cols).to_numpy()
y_train = train_df["ret"].to_numpy()
X_valid = valid_df.select(feature_cols).to_numpy()
y_valid = valid_df["ret"].to_numpy()
X_test = test_df.select(feature_cols).to_numpy()
y_test = test_df["ret"].to_numpy()

# Per-date timestamps for cross-sectional IC. Firm IDs are anonymized in this
# dataset, so synthesize per-row entity IDs to make the cross-section join 1:1.
dates_valid = valid_df["timestamp"].to_numpy()
entities_valid = np.arange(len(y_valid))
dates_test = test_df["timestamp"].to_numpy()
entities_test = np.arange(len(y_test))

print(f"Features: {len(feature_cols)}, Shapes: {X_train.shape} / {X_valid.shape} / {X_test.shape}")


# %% [markdown]
# ## 3. Study Configuration
#
# SQLite storage enables persistent studies. The database-backed approach
# also supports parallel workers in production settings (PostgreSQL/MySQL
# for true parallelism; SQLite for single-worker reproducibility).
#
# Persistence has one sharp edge worth flagging: with `load_if_exists=True`,
# re-running the notebook against an existing database *appends* new trials to
# the old study rather than starting over, so the trial counts and best configs
# would creep upward on every execution. To keep the results reproducible we
# delete any prior database at the top of the run, so a fresh execution always
# optimizes exactly `N_TRIALS` seeded trials per library.

# %%
STORAGE_PATH = f"sqlite:///{OUTPUT_DIR / 'hpo_studies.db'}"

# Reset persisted studies so each full run reproduces from scratch (see above).
_db_file = OUTPUT_DIR / "hpo_studies.db"
if _db_file.exists():
    _db_file.unlink()


def create_study(study_name: str, direction: str = "maximize") -> optuna.Study:
    """Create or load an Optuna study with TPE sampler and pruning."""
    return optuna.create_study(
        study_name=study_name,
        storage=STORAGE_PATH,
        load_if_exists=True,
        direction=direction,
        sampler=TPESampler(seed=SEED, n_startup_trials=N_WARMUP_TRIALS),
        pruner=MedianPruner(n_startup_trials=5, n_warmup_steps=10),
    )


all_results = []

# %% [markdown]
# ## 4. Objective Functions with Loss as Hyperparameter
#
# Each library has one study where `loss_type` (MSE/MAE) is a categorical
# parameter. This treats the loss function as a tunable decision rather than
# a fixed choice.


# %%
def make_xgboost_objective():
    """Factory for XGBoost objective with loss as hyperparameter."""

    def objective(trial: optuna.Trial) -> float:
        loss_type = trial.suggest_categorical("loss_type", ["mse", "mae"])
        xgb_obj = "reg:squarederror" if loss_type == "mse" else "reg:absoluteerror"

        params = {
            "objective": xgb_obj,
            "n_estimators": trial.suggest_int("n_estimators", 100, 1000),
            "max_depth": trial.suggest_int("max_depth", 2, 10),
            "learning_rate": trial.suggest_float("learning_rate", 0.005, 0.3, log=True),
            "subsample": trial.suggest_float("subsample", 0.5, 1.0),
            "colsample_bytree": trial.suggest_float("colsample_bytree", 0.5, 1.0),
            "min_child_weight": trial.suggest_int("min_child_weight", 1, 300),
            "reg_alpha": trial.suggest_float("reg_alpha", 1e-5, 100.0, log=True),
            "reg_lambda": trial.suggest_float("reg_lambda", 1e-5, 100.0, log=True),
            "tree_method": "hist",
            "device": XGB_DEVICE,
            "early_stopping_rounds": EARLY_STOPPING_ROUNDS,
            "random_state": SEED,
            "verbosity": 0,
        }
        model = xgb.XGBRegressor(**params)
        model.fit(X_train, y_train, eval_set=[(X_valid, y_valid)], verbose=False)
        return cross_sectional_ic_mean(y_valid, model.predict(X_valid), dates_valid, entities_valid)

    return objective


# %% [markdown]
# ### LightGBM Objective Factory


# %%
def make_lightgbm_objective():
    """Factory for LightGBM objective with loss as hyperparameter."""

    def objective(trial: optuna.Trial) -> float:
        loss_type = trial.suggest_categorical("loss_type", ["mse", "mae"])
        lgb_obj = "regression" if loss_type == "mse" else "mae"

        params = {
            "objective": lgb_obj,
            "n_estimators": trial.suggest_int("n_estimators", 100, 1000),
            "max_depth": trial.suggest_int("max_depth", 2, 10),
            "learning_rate": trial.suggest_float("learning_rate", 0.005, 0.3, log=True),
            "num_leaves": trial.suggest_int("num_leaves", 8, 128),
            "subsample": trial.suggest_float("subsample", 0.5, 1.0),
            "colsample_bytree": trial.suggest_float("colsample_bytree", 0.5, 1.0),
            "min_child_samples": trial.suggest_int("min_child_samples", 5, 300),
            "reg_alpha": trial.suggest_float("reg_alpha", 1e-5, 100.0, log=True),
            "reg_lambda": trial.suggest_float("reg_lambda", 1e-5, 100.0, log=True),
            "random_state": SEED,
            "verbose": -1,
        }
        model = lgb.LGBMRegressor(**params)
        model.fit(
            X_train,
            y_train,
            eval_set=[(X_valid, y_valid)],
            callbacks=[lgb.early_stopping(EARLY_STOPPING_ROUNDS, verbose=False)],
        )
        return cross_sectional_ic_mean(y_valid, model.predict(X_valid), dates_valid, entities_valid)

    return objective


# %% [markdown]
# ### CatBoost Objective Factory


# %%
def make_catboost_objective():
    """Factory for CatBoost objective with loss as hyperparameter."""

    def objective(trial: optuna.Trial) -> float:
        loss_type = trial.suggest_categorical("loss_type", ["mse", "mae"])
        cb_loss = "RMSE" if loss_type == "mse" else "MAE"

        params = {
            "loss_function": cb_loss,
            "iterations": trial.suggest_int("iterations", 100, 1000),
            "depth": trial.suggest_int("depth", 2, 10),
            "learning_rate": trial.suggest_float("learning_rate", 0.005, 0.3, log=True),
            "l2_leaf_reg": trial.suggest_float("l2_leaf_reg", 1e-5, 100.0, log=True),
            "random_strength": trial.suggest_float("random_strength", 0.0, 10.0),
            "bagging_temperature": trial.suggest_float("bagging_temperature", 0.0, 1.0),
            "task_type": CB_TASK_TYPE,
            "metric_period": CB_METRIC_PERIOD,
            "random_seed": SEED,
            "verbose": False,
            "early_stopping_rounds": EARLY_STOPPING_ROUNDS,
            "train_dir": str(CATBOOST_LOG_DIR),
        }
        model = cb.CatBoostRegressor(**params)
        model.fit(X_train, y_train, eval_set=(X_valid, y_valid), verbose=False)
        return cross_sectional_ic_mean(y_valid, model.predict(X_valid), dates_valid, entities_valid)

    return objective


# %% [markdown]
# ## 5. Model Building Helper
#
# Shared function to reconstruct and evaluate models from Optuna trial
# parameters. This eliminates the ~80 lines of duplicated model-building
# code that would otherwise appear in the optimization and per-loss sections.


# %%
def build_model(lib_code, params):
    """Build model instance from Optuna params and selected library."""
    loss_type = params.get("loss_type", "mse")
    eval_params = {k: v for k, v in params.items() if k not in ["loss_type"]}

    if lib_code == "xgb":
        xgb_obj = "reg:squarederror" if loss_type == "mse" else "reg:absoluteerror"
        return xgb.XGBRegressor(
            **eval_params,
            objective=xgb_obj,
            tree_method="hist",
            device=XGB_DEVICE,
            early_stopping_rounds=EARLY_STOPPING_ROUNDS,
            random_state=SEED,
            verbosity=0,
        )

    if lib_code == "lgb":
        lgb_obj = "regression" if loss_type == "mse" else "mae"
        return lgb.LGBMRegressor(
            **eval_params,
            objective=lgb_obj,
            random_state=SEED,
            verbose=-1,
        )

    cb_loss = "RMSE" if loss_type == "mse" else "MAE"
    return cb.CatBoostRegressor(
        **eval_params,
        loss_function=cb_loss,
        task_type=CB_TASK_TYPE,
        metric_period=CB_METRIC_PERIOD,
        early_stopping_rounds=EARLY_STOPPING_ROUNDS,
        random_seed=SEED,
        verbose=False,
        train_dir=str(CATBOOST_LOG_DIR),
    )


# %% [markdown]
# ### Evaluate Best-Config Model


# %%
def build_and_evaluate(
    lib_code,
    params,
    X_tr,
    y_tr,
    X_va,
    y_va,
    X_te,
    y_te,
    dates_va,
    entities_va,
    dates_te,
    entities_te,
):
    """Fit model from Optuna params and return (model, valid_ic, test_ic)."""
    model = build_model(lib_code, params)
    if lib_code == "xgb":
        model.fit(X_tr, y_tr, eval_set=[(X_va, y_va)], verbose=False)
    elif lib_code == "lgb":
        model.fit(
            X_tr,
            y_tr,
            eval_set=[(X_va, y_va)],
            callbacks=[lgb.early_stopping(EARLY_STOPPING_ROUNDS, verbose=False)],
        )
    else:
        model.fit(X_tr, y_tr, eval_set=(X_va, y_va), verbose=False)

    valid_ic = cross_sectional_ic_mean(y_va, model.predict(X_va), dates_va, entities_va)
    test_ic = cross_sectional_ic_mean(y_te, model.predict(X_te), dates_te, entities_te)
    return model, valid_ic, test_ic


# %% [markdown]
# ## 6. Run Optimization

# %%
LIBRARY_CONFIGS = [
    ("XGBoost", "academic_xgboost", make_xgboost_objective(), "xgb"),
    ("LightGBM", "academic_lightgbm", make_lightgbm_objective(), "lgb"),
    ("CatBoost", "academic_catboost", make_catboost_objective(), "cb"),
]
if N_LIBRARIES > 0:
    LIBRARY_CONFIGS = LIBRARY_CONFIGS[:N_LIBRARIES]

for lib_name, study_name, objective, lib_code in LIBRARY_CONFIGS:
    print(f"\nOptimizing {lib_name} ({N_TRIALS} trials)...")

    study = create_study(study_name)
    study.optimize(objective, n_trials=N_TRIALS, show_progress_bar=True)

    _, valid_ic, test_ic = build_and_evaluate(
        lib_code,
        study.best_params,
        X_train,
        y_train,
        X_valid,
        y_valid,
        X_test,
        y_test,
        dates_valid,
        entities_valid,
        dates_test,
        entities_test,
    )

    loss_type = study.best_params["loss_type"]
    print(f"  Best: Valid IC={valid_ic:.4f}, Test IC={test_ic:.4f}, Loss={loss_type.upper()}")

    all_results.append(
        {
            "model": lib_name,
            "library": lib_code.upper(),
            "loss_type": loss_type.upper(),
            "valid_ic": round(valid_ic, 4),
            "test_ic": round(test_ic, 4),
            "n_trials": len(study.trials),
        }
    )

# %% [markdown]
# ## 7. Per-Loss Analysis
#
# Extract each library's highest-scoring MSE and MAE configuration to
# assess whether loss function choice matters.

# %%
detailed_results = []

for lib_name, study_name, _, lib_code in LIBRARY_CONFIGS:
    study = optuna.load_study(study_name=study_name, storage=STORAGE_PATH)

    loss_results = {"mse": None, "mae": None}
    for trial in study.trials:
        if trial.state == optuna.trial.TrialState.COMPLETE:
            lt = trial.params["loss_type"]
            if loss_results[lt] is None or trial.value > loss_results[lt]["valid_ic"]:
                loss_results[lt] = {"valid_ic": trial.value, "params": trial.params}

    for lt, info in loss_results.items():
        if info is None:
            continue

        _, valid_ic, test_ic = build_and_evaluate(
            lib_code,
            info["params"],
            X_train,
            y_train,
            X_valid,
            y_valid,
            X_test,
            y_test,
            dates_valid,
            entities_valid,
            dates_test,
            entities_test,
        )
        detailed_results.append(
            {
                "model": f"{lib_name}_{lt.upper()}",
                "library": lib_code.upper(),
                "loss_type": lt.upper(),
                "valid_ic": round(valid_ic, 4),
                "test_ic": round(test_ic, 4),
            }
        )

# %% [markdown]
# ## 8. Results Summary

# %%
results_df = pl.DataFrame(all_results).sort("test_ic", descending=True)
results_df.select(["model", "loss_type", "valid_ic", "test_ic", "n_trials"])

# %%
results_df.write_csv(OUTPUT_DIR / "hpo_gbm_results.csv")

# %%
detailed_df = pl.DataFrame(detailed_results).sort("test_ic", descending=True)
detailed_df.select(["model", "loss_type", "valid_ic", "test_ic"])

# %%
detailed_df.write_csv(OUTPUT_DIR / "hpo_gbm_detailed.csv")

# %% [markdown]
# ## 9. MSE vs MAE Comparison

# %%
loss_comparison = []
for lib in detailed_df["library"].unique(maintain_order=True).to_list():
    lib_mse = detailed_df.filter((pl.col("library") == lib) & (pl.col("loss_type") == "MSE"))
    lib_mae = detailed_df.filter((pl.col("library") == lib) & (pl.col("loss_type") == "MAE"))
    if len(lib_mse) > 0 and len(lib_mae) > 0:
        mse_ic = lib_mse["test_ic"][0]
        mae_ic = lib_mae["test_ic"][0]
        loss_comparison.append(
            {
                "library": lib,
                "mse_test_ic": mse_ic,
                "mae_test_ic": mae_ic,
                "difference": round(mae_ic - mse_ic, 4),
                "loss_choice": "MAE" if mae_ic > mse_ic else "MSE",
            }
        )

loss_comparison_df = pl.DataFrame(
    loss_comparison,
    schema={
        "library": pl.Utf8,
        "mse_test_ic": pl.Float64,
        "mae_test_ic": pl.Float64,
        "difference": pl.Float64,
        "loss_choice": pl.Utf8,
    },
)
loss_comparison_df

# %% [markdown]
# A library appears below only when its study completed a trial under each loss,
# which a short search need not do: the sampler can spend every trial on one of the
# two. Each bar is the highest-scoring configuration the search found *conditional on*
# that
# loss, scored on the held-out 2000-2016 window. The two therefore differ in depth,
# learning rate and regularization as well as in the loss, and the sampler did not
# spend equal effort on each - the loss distribution printed further down says how
# unequal. So the gap is the difference between how far the search got under each
# loss, which is the practical question, and not the isolated effect of the loss
# function on an otherwise matched model.

# %%
if loss_comparison_df.is_empty():
    print(
        "No library completed a trial under both losses in this run, so there is no\n"
        "MSE-against-MAE pair to draw. The comparison needs one configuration of each."
    )
else:
    libs = loss_comparison_df["library"].to_list()
    mse_ics = loss_comparison_df["mse_test_ic"].to_list()
    mae_ics = loss_comparison_df["mae_test_ic"].to_list()

    x = np.arange(len(libs))
    width = 0.38

    fig, ax = plt.subplots(figsize=(8, 5))
    bars_mse = ax.bar(x - width / 2, mse_ics, width, label="MSE loss", color=COLORS["slate"])
    bars_mae = ax.bar(x + width / 2, mae_ics, width, label="MAE loss", color=COLORS["amber"])
    for bars in (bars_mse, bars_mae):
        for bar in bars:
            ax.text(
                bar.get_x() + bar.get_width() / 2,
                bar.get_height() + 0.0007,
                f"{bar.get_height():.4f}",
                ha="center",
                fontsize=9,
            )
    ax.set_xticks(x)
    ax.set_xticklabels(libs)
    ax.set_xlabel("Library")
    ax.set_ylabel("Held-out test IC (2000-2016)")
    ax.legend()
    ax.set_title("Held-out IC of each library's best MSE and MAE configuration")
    show_with_alt(
        fig,
        "Grouped bars of held-out test information coefficient, one pair per library, "
        "comparing its best MSE configuration against its best MAE configuration.",
    )

    # The count and the average gap are results, so they belong here rather than in
    # the title, where a reader would have to check them against the chart.
    n_mae = int((loss_comparison_df["loss_choice"] == "MAE").sum())
    mean_gain = float(loss_comparison_df["difference"].mean())
    print(
        f"MAE scores higher on the held-out window for {n_mae} of "
        f"{loss_comparison_df.height} libraries, by {mean_gain:+.4f} IC on average."
    )

# %% [markdown]
# **What to read off it.** The gap between the two bars of one library is the
# cost of fixing the loss by prior instead of letting the search choose it. MAE
# is the more robust of the two to outliers, which is the reason to expect it to
# matter on a target with fat tails - but expecting it is not the same as
# measuring it, and the sign of the gap is what the run reports. Declaring the
# loss a hyperparameter is what makes the question answerable at all.

# %% [markdown]
# ## 10. Hyperparameter Importance

# %%
importance_records = []

for lib_name, study_name, _, _ in LIBRARY_CONFIGS:
    try:
        study = optuna.load_study(study_name=study_name, storage=STORAGE_PATH)
        if len(study.trials) > 1:
            importance = optuna.importance.get_param_importances(study)
            for param, imp in list(importance.items())[:6]:
                importance_records.append(
                    {"model": lib_name, "param": param, "importance": round(imp, 4)}
                )
    except Exception as e:
        print(f"Could not analyze {study_name}: {e}")

importance_df = pl.DataFrame(importance_records)
importance_df

# %% [markdown]
# Ranking the fANOVA importances per library puts the loss-type decision at or
# near the top of every panel: the categorical MSE/MAE switch explains more of
# the validation-IC variance than any continuous hyperparameter.

# %%
lib_order = [ln for ln, *_ in LIBRARY_CONFIGS]
fig, axes = plt.subplots(1, len(lib_order), figsize=(4.5 * len(lib_order), 5), sharex=False)
if len(lib_order) == 1:
    axes = [axes]

for ax, lib_name in zip(axes, lib_order, strict=False):
    sub = importance_df.filter(pl.col("model") == lib_name).sort("importance")
    params = sub["param"].to_list()
    imps = sub["importance"].to_list()
    bar_colors = [COLORS["amber"] if p == "loss_type" else COLORS["slate"] for p in params]
    ax.barh(params, imps, color=bar_colors)
    for i, imp in enumerate(imps):
        ax.text(imp + 0.01, i, f"{imp:.2f}", va="center", fontsize=8)
    ax.set_xlim(0, 1.05)
    ax.set_xlabel("fANOVA importance")
    ax.set_title(lib_name)

fig.suptitle("fANOVA importance by hyperparameter, one panel per library", fontsize=13)
show_with_alt(
    fig,
    "One horizontal bar panel per library, ranking its hyperparameters by fANOVA "
    "importance, with the loss-type bar picked out in gold.",
)

# %%
importance_df.write_csv(OUTPUT_DIR / "hpo_gbm_importance.csv")

# %% [markdown]
# ## 11. Optimization History

# %%
history = {}
best_trials = {}
for lib_name, study_name, _, _ in LIBRARY_CONFIGS:
    study = optuna.load_study(study_name=study_name, storage=STORAGE_PATH)
    best_num = study.best_trial.number
    total = len(study.trials)
    best_trials[lib_name] = best_num

    loss_counts = {"mse": 0, "mae": 0}
    values = []
    for trial in study.trials:
        if trial.state == optuna.trial.TrialState.COMPLETE:
            loss_counts[trial.params["loss_type"]] += 1
        values.append(trial.value if trial.value is not None else np.nan)
    history[lib_name] = np.array(values, dtype=float)

    print(
        f"{lib_name}: best at trial #{best_num}/{total}, "
        f"loss={study.best_params['loss_type'].upper()}, "
        f"distribution: MSE={loss_counts['mse']}, MAE={loss_counts['mae']}"
    )

# %% [markdown]
# Tracking the highest validation IC found so far against the trial index shows where
# TPE improved and where it sat still.

# %%
palette = [COLORS["slate"], COLORS["amber"], COLORS["copper"]]

fig, ax = plt.subplots(figsize=(9, 5))
for (lib_name, values), color in zip(history.items(), palette, strict=False):
    running_best = np.fmax.accumulate(np.nan_to_num(values, nan=-np.inf))
    running_best[running_best == -np.inf] = np.nan
    trials_axis = np.arange(len(values))
    ax.plot(trials_axis, running_best, color=color, linewidth=2, label=lib_name)

ax.axvline(N_WARMUP_TRIALS, linestyle="--", color="gray", linewidth=0.8, label="End of TPE warm-up")
ax.set_xlabel("Trial number")
ax.set_ylabel("Best validation IC so far")
ax.legend()
ax.set_title("Best validation IC so far, by trial, against the TPE warm-up")
show_with_alt(
    fig,
    "Best validation information coefficient found so far against trial number, one "
    "line per library, with a dashed vertical line at the end of the TPE warm-up.",
)
earliest_late_best = min(best_trials.values())
_late = sum(1 for t in best_trials.values() if t > N_WARMUP_TRIALS)
print(
    f"{_late} of {len(best_trials)} libraries found their best configuration after the "
    f"{N_WARMUP_TRIALS}-trial warm-up; the earliest any of them found it was trial "
    f"{earliest_late_best}."
)

# %%
best_overall = results_df.row(0, named=True)
print(
    f"Best GBM: {best_overall['model']} ({best_overall['loss_type']}) IC={best_overall['test_ic']:.4f}"
)

# %% [markdown]
# ## Key Takeaways
#
# 1. **Loss type matters as a hyperparameter.** Putting MSE and MAE in the search
#    space lets Optuna discover the better loss for each library. Here the
#    categorical loss switch is the single most important hyperparameter in every
#    study, and MAE wins across all three libraries on these heavy-tailed returns.
# 2. **Early stopping tames a large search range.** Setting `n_estimators` up to
#    1000 with an early-stopping patience of 50 lets the data pick the tree count
#    instead of the search wasting compute on over-grown models.
# 3. **The library matters less than the tuning.** Read the spread of the three
#    libraries' test IC in the summary table against the MSE-to-MAE gap in the loss
#    comparison: the second is the larger of the two here, so a careful search over
#    loss and depth buys more than switching implementations.
# 4. **These GBM results connect to Chapter 20.** The per-library, per-loss IC
#    values illustrate the kind of tuned tabular baselines the cross-model
#    synthesis in Chapter 20 builds on.
#
# **Next**: See `04_optuna_tuning` for the full Optuna workflow on the ETF case
# study, including walk-forward HPO and pruning effectiveness.

```

Reproduit dans son intégralité avec attribution, conformément à la licence de la source. Licence: MIT

Ce résumé a été rédigé par l’agent de recherche de Stratmill à partir de la source originale ; il n’en est pas une copie.