Pular para o conteúdo
Todos os documentos da biblioteca

Ajuste de hiperparâmetros Optuna em três bibliotecas

Notebook Machine Learning for Trading

Resumo

Este notebook compara buscas de hiperparâmetros com Optuna para XGBoost, LightGBM e CatBoost em um conjunto de dados de características de empresas. Cada biblioteca recebe seu próprio estudo de busca, usando um estimador de Parzen com estrutura de árvore, poda mediana, parada antecipada e armazenamento persistente do estudo. A escolha entre erro quadrático médio e erro absoluto médio é incluída como hiperparâmetro categórico, enquanto o desempenho na validação é medido pelo coeficiente de informação transversal médio.

O notebook relata que MAE vence nas três bibliotecas para os retornos de cauda pesada deste conjunto de dados e que a escolha da função de perda é o parâmetro de maior influência em cada busca. Também compara IC de validação e teste após a reconstrução do modelo e considera a importância dos parâmetros. Essas constatações se aplicam ao conjunto de dados e à configuração de busca apresentados; execuções em GPU podem variar ligeiramente devido a operações paralelas de ponto flutuante, e selecionar configurações pelos resultados de validação não garante desempenho em outros períodos ou mercados.

Ideias principais

  • O notebook busca configurações para XGBoost, LightGBM e CatBoost em uma estrutura comum baseada em Optuna.
  • Ele trata a escolha da função de perda de regressão entre MSE e MAE como parâmetro ajustável.
  • O coeficiente de informação transversal médio na validação orienta a busca, com parada antecipada limitando o crescimento desnecessário das árvores.
  • Os resultados relatados favorecem MAE nas três bibliotecas para este conjunto de dados de cauda pesada.
  • Resultados específicos do conjunto de dados e pequenas variações de não determinismo em GPU limitam a generalização ampla da comparação.

Tags

Texto completo
# Cross-Library HPO: XGBoost vs LightGBM vs CatBoost


# Cross-Library HPO: XGBoost vs LightGBM vs CatBoost

**Docker image**: `ml4t-gpu`

**Chapter 12, Section 12.4**: Advanced Hyperparameter Tuning with Optuna

> **GPU recommended**: This notebook trains models with PyTorch/CUDA. It will run on CPU
> but training may be very slow. For GPU acceleration:
> ```bash
> docker compose run --rm ml4t-gpu python 12_gradient_boosting/05_cross_library_hpo.py
> ```


## Purpose
Optuna-based HPO for XGBoost, LightGBM, and CatBoost on the Chen-Pelger-Zhu
(2020) firm characteristics dataset. **Loss type is a hyperparameter**
within each library's study, treating the choice of MSE vs MAE as a tunable
decision rather than a fixed prior.

## Learning Objectives
- Tune XGBoost, LightGBM, and CatBoost with Optuna under a shared protocol
- Treat loss function selection as a first-class hyperparameter
- Compare best-valid and best-test IC after model reconstruction
- Diagnose which hyperparameters drive gains in noisy financial targets

**Prerequisites**: Requires the firm characteristics dataset and supporting
loaders from `data/__init__.py`.

## Key Concepts
- TPE with configurable warm-up trials for each library
- MedianPruner for early stopping of unpromising trials
- SQLite persistence for reproducibility
- Loss function as a categorical hyperparameter
- Hyperparameter importance analysis (including loss_type)

## Cross-References
- **Section 12.4**: GBM-specific tuning strategy, pruning
- **Chapter 20**: These results feed into the cross-model synthesis
- **Related**: `04_optuna_tuning` (ETF case study), `07_hpo_comparison` (grid vs Optuna)

## 1. Setup

```python
"""Cross-Library HPO: tune XGBoost, LightGBM, and CatBoost with Optuna on financial data."""

# Import torch before ml4t.diagnostic. ml4t.diagnostic transitively dlopens
# the older system libcudart.so.12; loading torch first makes its bundled
# CUDA runtime win symbol resolution.
import warnings

import catboost as cb
import lightgbm as lgb
import matplotlib.pyplot as plt
import numpy as np
import optuna
import polars as pl
import torch  # noqa: F401
import xgboost as xgb
from optuna.pruners import MedianPruner
from optuna.samplers import TPESampler

from data import load_firm_characteristics
from utils.modeling import cross_sectional_ic_mean
from utils.paths import get_output_dir
from utils.reproducibility import set_global_seeds
from utils.style import COLORS, show_with_alt

# LightGBM records synthetic feature names when fitted on an array with an eval_set,
# and sklearn then warns at every predict on an array that has none to compare. One
# message, not the category: the fit and the predictions are unaffected.
warnings.filterwarnings(
    "ignore",
    message="X does not have valid feature names",
    category=UserWarning,
    module="sklearn.utils.validation",
)
optuna.logging.set_verbosity(optuna.logging.WARNING)
```

```python
N_TRIALS = 50
N_WARMUP_TRIALS = 5
EARLY_STOPPING_ROUNDS = 50
# 0 = all libraries, 1-3 = first N only
N_LIBRARIES = 0
SEED = 42
```

```python
set_global_seeds(SEED)
```

```python
OUTPUT_DIR = get_output_dir(12, "us_firm_characteristics")
CATBOOST_LOG_DIR = OUTPUT_DIR / "catboost_info"
OUTPUT_DIR.mkdir(parents=True, exist_ok=True)
```

### Device Detection

XGBoost and CatBoost support GPU acceleration. We detect availability
at startup rather than hardcoding `device="cuda"`.

```python
GPU_AVAILABLE = torch.cuda.is_available()

XGB_DEVICE = "cuda" if GPU_AVAILABLE else "cpu"
CB_TASK_TYPE = "GPU" if GPU_AVAILABLE else "CPU"
# CatBoost scores an MAE eval metric every fifth iteration on GPU, which it cannot
# compute on the device; declaring that period keeps the behaviour without the
# per-fit announcement on stderr.
CB_METRIC_PERIOD = 5 if GPU_AVAILABLE else 1
print(f"GPU available: {GPU_AVAILABLE} (XGB: {XGB_DEVICE}, CatBoost: {CB_TASK_TYPE})")
```

> **Reproducibility.** This search runs on a GPU for speed. GPU training is not
> bitwise reproducible: fixed seeds pin each model's random choices but not the
> order in which parallel hardware sums floating-point values, so the selected
> configurations and IC values here are stable across reruns rather
> than identical to the last decimal. For runs that must reproduce exactly, train
> on CPU with deterministic settings; `02_gbm_comparison` measures the
> GPU-versus-CPU prediction gap and gives the deterministic CPU recipe in full.

## 2. Load Academic Dataset

```python
train_df = load_firm_characteristics(split="train")
valid_df = load_firm_characteristics(split="valid")
test_df = load_firm_characteristics(split="test")

print(
    f"Train: {len(train_df):,} (1967-1989), Valid: {len(valid_df):,} (1990-1999), Test: {len(test_df):,} (2000-2016)"
)
```

```python
feature_cols = [c for c in train_df.columns if c not in ["timestamp", "ret", "split"]]

X_train = train_df.select(feature_cols).to_numpy()
y_train = train_df["ret"].to_numpy()
X_valid = valid_df.select(feature_cols).to_numpy()
y_valid = valid_df["ret"].to_numpy()
X_test = test_df.select(feature_cols).to_numpy()
y_test = test_df["ret"].to_numpy()

# Per-date timestamps for cross-sectional IC. Firm IDs are anonymized in this
# dataset, so synthesize per-row entity IDs to make the cross-section join 1:1.
dates_valid = valid_df["timestamp"].to_numpy()
entities_valid = np.arange(len(y_valid))
dates_test = test_df["timestamp"].to_numpy()
entities_test = np.arange(len(y_test))

print(f"Features: {len(feature_cols)}, Shapes: {X_train.shape} / {X_valid.shape} / {X_test.shape}")
```

## 3. Study Configuration

SQLite storage enables persistent studies. The database-backed approach
also supports parallel workers in production settings (PostgreSQL/MySQL
for true parallelism; SQLite for single-worker reproducibility).

Persistence has one sharp edge worth flagging: with `load_if_exists=True`,
re-running the notebook against an existing database *appends* new trials to
the old study rather than starting over, so the trial counts and best configs
would creep upward on every execution. To keep the results reproducible we
delete any prior database at the top of the run, so a fresh execution always
optimizes exactly `N_TRIALS` seeded trials per library.

```python
STORAGE_PATH = f"sqlite:///{OUTPUT_DIR / 'hpo_studies.db'}"

# Reset persisted studies so each full run reproduces from scratch (see above).
_db_file = OUTPUT_DIR / "hpo_studies.db"
if _db_file.exists():
    _db_file.unlink()


def create_study(study_name: str, direction: str = "maximize") -> optuna.Study:
    """Create or load an Optuna study with TPE sampler and pruning."""
    return optuna.create_study(
        study_name=study_name,
        storage=STORAGE_PATH,
        load_if_exists=True,
        direction=direction,
        sampler=TPESampler(seed=SEED, n_startup_trials=N_WARMUP_TRIALS),
        pruner=MedianPruner(n_startup_trials=5, n_warmup_steps=10),
    )


all_results = []
```

## 4. Objective Functions with Loss as Hyperparameter

Each library has one study where `loss_type` (MSE/MAE) is a categorical
parameter. This treats the loss function as a tunable decision rather than
a fixed choice.

```python
def make_xgboost_objective():
    """Factory for XGBoost objective with loss as hyperparameter."""

    def objective(trial: optuna.Trial) -> float:
        loss_type = trial.suggest_categorical("loss_type", ["mse", "mae"])
        xgb_obj = "reg:squarederror" if loss_type == "mse" else "reg:absoluteerror"

        params = {
            "objective": xgb_obj,
            "n_estimators": trial.suggest_int("n_estimators", 100, 1000),
            "max_depth": trial.suggest_int("max_depth", 2, 10),
            "learning_rate": trial.suggest_float("learning_rate", 0.005, 0.3, log=True),
            "subsample": trial.suggest_float("subsample", 0.5, 1.0),
            "colsample_bytree": trial.suggest_float("colsample_bytree", 0.5, 1.0),
            "min_child_weight": trial.suggest_int("min_child_weight", 1, 300),
            "reg_alpha": trial.suggest_float("reg_alpha", 1e-5, 100.0, log=True),
            "reg_lambda": trial.suggest_float("reg_lambda", 1e-5, 100.0, log=True),
            "tree_method": "hist",
            "device": XGB_DEVICE,
            "early_stopping_rounds": EARLY_STOPPING_ROUNDS,
            "random_state": SEED,
            "verbosity": 0,
        }
        model = xgb.XGBRegressor(**params)
        model.fit(X_train, y_train, eval_set=[(X_valid, y_valid)], verbose=False)
        return cross_sectional_ic_mean(y_valid, model.predict(X_valid), dates_valid, entities_valid)

    return objective
```

### LightGBM Objective Factory

```python
def make_lightgbm_objective():
    """Factory for LightGBM objective with loss as hyperparameter."""

    def objective(trial: optuna.Trial) -> float:
        loss_type = trial.suggest_categorical("loss_type", ["mse", "mae"])
        lgb_obj = "regression" if loss_type == "mse" else "mae"

        params = {
            "objective": lgb_obj,
            "n_estimators": trial.suggest_int("n_estimators", 100, 1000),
            "max_depth": trial.suggest_int("max_depth", 2, 10),
            "learning_rate": trial.suggest_float("learning_rate", 0.005, 0.3, log=True),
            "num_leaves": trial.suggest_int("num_leaves", 8, 128),
            "subsample": trial.suggest_float("subsample", 0.5, 1.0),
            "colsample_bytree": trial.suggest_float("colsample_bytree", 0.5, 1.0),
            "min_child_samples": trial.suggest_int("min_child_samples", 5, 300),
            "reg_alpha": trial.suggest_float("reg_alpha", 1e-5, 100.0, log=True),
            "reg_lambda": trial.suggest_float("reg_lambda", 1e-5, 100.0, log=True),
            "random_state": SEED,
            "verbose": -1,
        }
        model = lgb.LGBMRegressor(**params)
        model.fit(
            X_train,
            y_train,
            eval_set=[(X_valid, y_valid)],
            callbacks=[lgb.early_stopping(EARLY_STOPPING_ROUNDS, verbose=False)],
        )
        return cross_sectional_ic_mean(y_valid, model.predict(X_valid), dates_valid, entities_valid)

    return objective
```

### CatBoost Objective Factory

```python
def make_catboost_objective():
    """Factory for CatBoost objective with loss as hyperparameter."""

    def objective(trial: optuna.Trial) -> float:
        loss_type = trial.suggest_categorical("loss_type", ["mse", "mae"])
        cb_loss = "RMSE" if loss_type == "mse" else "MAE"

        params = {
            "loss_function": cb_loss,
            "iterations": trial.suggest_int("iterations", 100, 1000),
            "depth": trial.suggest_int("depth", 2, 10),
            "learning_rate": trial.suggest_float("learning_rate", 0.005, 0.3, log=True),
            "l2_leaf_reg": trial.suggest_float("l2_leaf_reg", 1e-5, 100.0, log=True),
            "random_strength": trial.suggest_float("random_strength", 0.0, 10.0),
            "bagging_temperature": trial.suggest_float("bagging_temperature", 0.0, 1.0),
            "task_type": CB_TASK_TYPE,
            "metric_period": CB_METRIC_PERIOD,
            "random_seed": SEED,
            "verbose": False,
            "early_stopping_rounds": EARLY_STOPPING_ROUNDS,
            "train_dir": str(CATBOOST_LOG_DIR),
        }
        model = cb.CatBoostRegressor(**params)
        model.fit(X_train, y_train, eval_set=(X_valid, y_valid), verbose=False)
        return cross_sectional_ic_mean(y_valid, model.predict(X_valid), dates_valid, entities_valid)

    return objective
```

## 5. Model Building Helper

Shared function to reconstruct and evaluate models from Optuna trial
parameters. This eliminates the ~80 lines of duplicated model-building
code that would otherwise appear in the optimization and per-loss sections.

```python
def build_model(lib_code, params):
    """Build model instance from Optuna params and selected library."""
    loss_type = params.get("loss_type", "mse")
    eval_params = {k: v for k, v in params.items() if k not in ["loss_type"]}

    if lib_code == "xgb":
        xgb_obj = "reg:squarederror" if loss_type == "mse" else "reg:absoluteerror"
        return xgb.XGBRegressor(
            **eval_params,
            objective=xgb_obj,
            tree_method="hist",
            device=XGB_DEVICE,
            early_stopping_rounds=EARLY_STOPPING_ROUNDS,
            random_state=SEED,
            verbosity=0,
        )

    if lib_code == "lgb":
        lgb_obj = "regression" if loss_type == "mse" else "mae"
        return lgb.LGBMRegressor(
            **eval_params,
            objective=lgb_obj,
            random_state=SEED,
            verbose=-1,
        )

    cb_loss = "RMSE" if loss_type == "mse" else "MAE"
    return cb.CatBoostRegressor(
        **eval_params,
        loss_function=cb_loss,
        task_type=CB_TASK_TYPE,
        metric_period=CB_METRIC_PERIOD,
        early_stopping_rounds=EARLY_STOPPING_ROUNDS,
        random_seed=SEED,
        verbose=False,
        train_dir=str(CATBOOST_LOG_DIR),
    )
```

### Evaluate Best-Config Model

```python
def build_and_evaluate(
    lib_code,
    params,
    X_tr,
    y_tr,
    X_va,
    y_va,
    X_te,
    y_te,
    dates_va,
    entities_va,
    dates_te,
    entities_te,
):
    """Fit model from Optuna params and return (model, valid_ic, test_ic)."""
    model = build_model(lib_code, params)
    if lib_code == "xgb":
        model.fit(X_tr, y_tr, eval_set=[(X_va, y_va)], verbose=False)
    elif lib_code == "lgb":
        model.fit(
            X_tr,
            y_tr,
            eval_set=[(X_va, y_va)],
            callbacks=[lgb.early_stopping(EARLY_STOPPING_ROUNDS, verbose=False)],
        )
    else:
        model.fit(X_tr, y_tr, eval_set=(X_va, y_va), verbose=False)

    valid_ic = cross_sectional_ic_mean(y_va, model.predict(X_va), dates_va, entities_va)
    test_ic = cross_sectional_ic_mean(y_te, model.predict(X_te), dates_te, entities_te)
    return model, valid_ic, test_ic
```

## 6. Run Optimization

```python
LIBRARY_CONFIGS = [
    ("XGBoost", "academic_xgboost", make_xgboost_objective(), "xgb"),
    ("LightGBM", "academic_lightgbm", make_lightgbm_objective(), "lgb"),
    ("CatBoost", "academic_catboost", make_catboost_objective(), "cb"),
]
if N_LIBRARIES > 0:
    LIBRARY_CONFIGS = LIBRARY_CONFIGS[:N_LIBRARIES]

for lib_name, study_name, objective, lib_code in LIBRARY_CONFIGS:
    print(f"\nOptimizing {lib_name} ({N_TRIALS} trials)...")

    study = create_study(study_name)
    study.optimize(objective, n_trials=N_TRIALS, show_progress_bar=True)

    _, valid_ic, test_ic = build_and_evaluate(
        lib_code,
        study.best_params,
        X_train,
        y_train,
        X_valid,
        y_valid,
        X_test,
        y_test,
        dates_valid,
        entities_valid,
        dates_test,
        entities_test,
    )

    loss_type = study.best_params["loss_type"]
    print(f"  Best: Valid IC={valid_ic:.4f}, Test IC={test_ic:.4f}, Loss={loss_type.upper()}")

    all_results.append(
        {
            "model": lib_name,
            "library": lib_code.upper(),
            "loss_type": loss_type.upper(),
            "valid_ic": round(valid_ic, 4),
            "test_ic": round(test_ic, 4),
            "n_trials": len(study.trials),
        }
    )
```

## 7. Per-Loss Analysis

Extract each library's highest-scoring MSE and MAE configuration to
assess whether loss function choice matters.

```python
detailed_results = []

for lib_name, study_name, _, lib_code in LIBRARY_CONFIGS:
    study = optuna.load_study(study_name=study_name, storage=STORAGE_PATH)

    loss_results = {"mse": None, "mae": None}
    for trial in study.trials:
        if trial.state == optuna.trial.TrialState.COMPLETE:
            lt = trial.params["loss_type"]
            if loss_results[lt] is None or trial.value > loss_results[lt]["valid_ic"]:
                loss_results[lt] = {"valid_ic": trial.value, "params": trial.params}

    for lt, info in loss_results.items():
        if info is None:
            continue

        _, valid_ic, test_ic = build_and_evaluate(
            lib_code,
            info["params"],
            X_train,
            y_train,
            X_valid,
            y_valid,
            X_test,
            y_test,
            dates_valid,
            entities_valid,
            dates_test,
            entities_test,
        )
        detailed_results.append(
            {
                "model": f"{lib_name}_{lt.upper()}",
                "library": lib_code.upper(),
                "loss_type": lt.upper(),
                "valid_ic": round(valid_ic, 4),
                "test_ic": round(test_ic, 4),
            }
        )
```

## 8. Results Summary

```python
results_df = pl.DataFrame(all_results).sort("test_ic", descending=True)
results_df.select(["model", "loss_type", "valid_ic", "test_ic", "n_trials"])
```

```python
results_df.write_csv(OUTPUT_DIR / "hpo_gbm_results.csv")
```

```python
detailed_df = pl.DataFrame(detailed_results).sort("test_ic", descending=True)
detailed_df.select(["model", "loss_type", "valid_ic", "test_ic"])
```

```python
detailed_df.write_csv(OUTPUT_DIR / "hpo_gbm_detailed.csv")
```

## 9. MSE vs MAE Comparison

```python
loss_comparison = []
for lib in detailed_df["library"].unique(maintain_order=True).to_list():
    lib_mse = detailed_df.filter((pl.col("library") == lib) & (pl.col("loss_type") == "MSE"))
    lib_mae = detailed_df.filter((pl.col("library") == lib) & (pl.col("loss_type") == "MAE"))
    if len(lib_mse) > 0 and len(lib_mae) > 0:
        mse_ic = lib_mse["test_ic"][0]
        mae_ic = lib_mae["test_ic"][0]
        loss_comparison.append(
            {
                "library": lib,
                "mse_test_ic": mse_ic,
                "mae_test_ic": mae_ic,
                "difference": round(mae_ic - mse_ic, 4),
                "loss_choice": "MAE" if mae_ic > mse_ic else "MSE",
            }
        )

loss_comparison_df = pl.DataFrame(
    loss_comparison,
    schema={
        "library": pl.Utf8,
        "mse_test_ic": pl.Float64,
        "mae_test_ic": pl.Float64,
        "difference": pl.Float64,
        "loss_choice": pl.Utf8,
    },
)
loss_comparison_df
```

A library appears below only when its study completed a trial under each loss,
which a short search need not do: the sampler can spend every trial on one of the
two. Each bar is the highest-scoring configuration the search found *conditional on*
that
loss, scored on the held-out 2000-2016 window. The two therefore differ in depth,
learning rate and regularization as well as in the loss, and the sampler did not
spend equal effort on each - the loss distribution printed further down says how
unequal. So the gap is the difference between how far the search got under each
loss, which is the practical question, and not the isolated effect of the loss
function on an otherwise matched model.

```python
if loss_comparison_df.is_empty():
    print(
        "No library completed a trial under both losses in this run, so there is no\n"
        "MSE-against-MAE pair to draw. The comparison needs one configuration of each."
    )
else:
    libs = loss_comparison_df["library"].to_list()
    mse_ics = loss_comparison_df["mse_test_ic"].to_list()
    mae_ics = loss_comparison_df["mae_test_ic"].to_list()

    x = np.arange(len(libs))
    width = 0.38

    fig, ax = plt.subplots(figsize=(8, 5))
    bars_mse = ax.bar(x - width / 2, mse_ics, width, label="MSE loss", color=COLORS["slate"])
    bars_mae = ax.bar(x + width / 2, mae_ics, width, label="MAE loss", color=COLORS["amber"])
    for bars in (bars_mse, bars_mae):
        for bar in bars:
            ax.text(
                bar.get_x() + bar.get_width() / 2,
                bar.get_height() + 0.0007,
                f"{bar.get_height():.4f}",
                ha="center",
                fontsize=9,
            )
    ax.set_xticks(x)
    ax.set_xticklabels(libs)
    ax.set_xlabel("Library")
    ax.set_ylabel("Held-out test IC (2000-2016)")
    ax.legend()
    ax.set_title("Held-out IC of each library's best MSE and MAE configuration")
    show_with_alt(
        fig,
        "Grouped bars of held-out test information coefficient, one pair per library, "
        "comparing its best MSE configuration against its best MAE configuration.",
    )

    # The count and the average gap are results, so they belong here rather than in
    # the title, where a reader would have to check them against the chart.
    n_mae = int((loss_comparison_df["loss_choice"] == "MAE").sum())
    mean_gain = float(loss_comparison_df["difference"].mean())
    print(
        f"MAE scores higher on the held-out window for {n_mae} of "
        f"{loss_comparison_df.height} libraries, by {mean_gain:+.4f} IC on average."
    )
```

**What to read off it.** The gap between the two bars of one library is the
cost of fixing the loss by prior instead of letting the search choose it. MAE
is the more robust of the two to outliers, which is the reason to expect it to
matter on a target with fat tails - but expecting it is not the same as
measuring it, and the sign of the gap is what the run reports. Declaring the
loss a hyperparameter is what makes the question answerable at all.

## 10. Hyperparameter Importance

```python
importance_records = []

for lib_name, study_name, _, _ in LIBRARY_CONFIGS:
    try:
        study = optuna.load_study(study_name=study_name, storage=STORAGE_PATH)
        if len(study.trials) > 1:
            importance = optuna.importance.get_param_importances(study)
            for param, imp in list(importance.items())[:6]:
                importance_records.append(
                    {"model": lib_name, "param": param, "importance": round(imp, 4)}
                )
    except Exception as e:
        print(f"Could not analyze {study_name}: {e}")

importance_df = pl.DataFrame(importance_records)
importance_df
```

Ranking the fANOVA importances per library puts the loss-type decision at or
near the top of every panel: the categorical MSE/MAE switch explains more of
the validation-IC variance than any continuous hyperparameter.

```python
lib_order = [ln for ln, *_ in LIBRARY_CONFIGS]
fig, axes = plt.subplots(1, len(lib_order), figsize=(4.5 * len(lib_order), 5), sharex=False)
if len(lib_order) == 1:
    axes = [axes]

for ax, lib_name in zip(axes, lib_order, strict=False):
    sub = importance_df.filter(pl.col("model") == lib_name).sort("importance")
    params = sub["param"].to_list()
    imps = sub["importance"].to_list()
    bar_colors = [COLORS["amber"] if p == "loss_type" else COLORS["slate"] for p in params]
    ax.barh(params, imps, color=bar_colors)
    for i, imp in enumerate(imps):
        ax.text(imp + 0.01, i, f"{imp:.2f}", va="center", fontsize=8)
    ax.set_xlim(0, 1.05)
    ax.set_xlabel("fANOVA importance")
    ax.set_title(lib_name)

fig.suptitle("fANOVA importance by hyperparameter, one panel per library", fontsize=13)
show_with_alt(
    fig,
    "One horizontal bar panel per library, ranking its hyperparameters by fANOVA "
    "importance, with the loss-type bar picked out in gold.",
)
```

```python
importance_df.write_csv(OUTPUT_DIR / "hpo_gbm_importance.csv")
```

## 11. Optimization History

```python
history = {}
best_trials = {}
for lib_name, study_name, _, _ in LIBRARY_CONFIGS:
    study = optuna.load_study(study_name=study_name, storage=STORAGE_PATH)
    best_num = study.best_trial.number
    total = len(study.trials)
    best_trials[lib_name] = best_num

    loss_counts = {"mse": 0, "mae": 0}
    values = []
    for trial in study.trials:
        if trial.state == optuna.trial.TrialState.COMPLETE:
            loss_counts[trial.params["loss_type"]] += 1
        values.append(trial.value if trial.value is not None else np.nan)
    history[lib_name] = np.array(values, dtype=float)

    print(
        f"{lib_name}: best at trial #{best_num}/{total}, "
        f"loss={study.best_params['loss_type'].upper()}, "
        f"distribution: MSE={loss_counts['mse']}, MAE={loss_counts['mae']}"
    )
```

Tracking the highest validation IC found so far against the trial index shows where
TPE improved and where it sat still.

```python
palette = [COLORS["slate"], COLORS["amber"], COLORS["copper"]]

fig, ax = plt.subplots(figsize=(9, 5))
for (lib_name, values), color in zip(history.items(), palette, strict=False):
    running_best = np.fmax.accumulate(np.nan_to_num(values, nan=-np.inf))
    running_best[running_best == -np.inf] = np.nan
    trials_axis = np.arange(len(values))
    ax.plot(trials_axis, running_best, color=color, linewidth=2, label=lib_name)

ax.axvline(N_WARMUP_TRIALS, linestyle="--", color="gray", linewidth=0.8, label="End of TPE warm-up")
ax.set_xlabel("Trial number")
ax.set_ylabel("Best validation IC so far")
ax.legend()
ax.set_title("Best validation IC so far, by trial, against the TPE warm-up")
show_with_alt(
    fig,
    "Best validation information coefficient found so far against trial number, one "
    "line per library, with a dashed vertical line at the end of the TPE warm-up.",
)
earliest_late_best = min(best_trials.values())
_late = sum(1 for t in best_trials.values() if t > N_WARMUP_TRIALS)
print(
    f"{_late} of {len(best_trials)} libraries found their best configuration after the "
    f"{N_WARMUP_TRIALS}-trial warm-up; the earliest any of them found it was trial "
    f"{earliest_late_best}."
)
```

```python
best_overall = results_df.row(0, named=True)
print(
    f"Best GBM: {best_overall['model']} ({best_overall['loss_type']}) IC={best_overall['test_ic']:.4f}"
)
```

## Key Takeaways

1. **Loss type matters as a hyperparameter.** Putting MSE and MAE in the search
   space lets Optuna discover the better loss for each library. Here the
   categorical loss switch is the single most important hyperparameter in every
   study, and MAE wins across all three libraries on these heavy-tailed returns.
2. **Early stopping tames a large search range.** Setting `n_estimators` up to
   1000 with an early-stopping patience of 50 lets the data pick the tree count
   instead of the search wasting compute on over-grown models.
3. **The library matters less than the tuning.** Read the spread of the three
   libraries' test IC in the summary table against the MSE-to-MAE gap in the loss
   comparison: the second is the larger of the two here, so a careful search over
   loss and depth buys more than switching implementations.
4. **These GBM results connect to Chapter 20.** The per-library, per-loss IC
   values illustrate the kind of tuned tabular baselines the cross-model
   synthesis in Chapter 20 builds on.

**Next**: See `04_optuna_tuning` for the full Optuna workflow on the ETF case
study, including walk-forward HPO and pruning effectiveness.
![notebook output](figures/p1_1.png)
![notebook output](figures/p1_2.png)
![notebook output](figures/p1_3.png)

Exibido na íntegra, com atribuição conforme a licença da fonte. Licença: MIT

Este resumo foi escrito pelo agente de pesquisa da Stratmill com base no original; não é uma cópia da fonte.