مواد پر جائیں
لائبریری کی تمام دستاویزات

ETF گریڈینٹ بوسٹنگ: تعاملات، باہمی خطیت اور چیک پوائنٹ کا انتخاب

نوٹ بک Machine Learning for Trading

خلاصہ

یہ کیس اسٹڈی ETF منافع کے لیبلز پر گریڈینٹ بوسٹڈ درخت فٹ کرتی اور ان کا ایک ایسے خطی بنیادی ماڈل سے موازنہ کرتی ہے جس نے مضبوط رِج ریگولرائزیشن کے تحت اچھی کارکردگی دکھائی۔ یہ بتاتی ہے کہ درخت تعاملات کیسے دریافت کر سکتے ہیں، مثلاً مارکیٹ کے دباؤ کی مختلف حالتوں میں مومینٹم کا مختلف برتاؤ، جبکہ حریصانہ تقسیمیں باہم مربوط فیچرز میں من مانا انتخاب کر سکتی ہیں جنہیں رِج یکجا کر سکتا ہے۔ گرڈ درخت کی گنجائش اور لاس بدلتی اور تربیت کے دوران ہر ترتیب کو متعدد چیک پوائنٹس پر جانچتی ہے۔

رپورٹ کردہ لرننگ کروز بتاتے ہیں کہ بہت سی ترتیبیں مقررہ تربیتی مدت پر پہلے کے چیک پوائنٹس سے بدتر کارکردگی دکھاتی ہیں؛ کچھ کی بہترین معلوماتی نسبت تربیت کے بیچ میں آتی ہے۔ چیک پوائنٹس کے درمیان فرق، ترتیبوں کے فرق کے مقابلے میں خاصا بڑا ہے، اس لیے بعد میں پسندیدہ چیک پوائنٹ چننا ایک اور انتخابی فیصلہ ہے۔ نوٹ پیش گوئیاں شائع کرتا ہے، مگر ماڈل اور چیک پوائنٹ کا انتخاب شارپ پر مبنی بعد کی توثیقی بیک ٹیسٹنگ پر چھوڑتا ہے۔ اس کے IC نتائج درجہ بندی کی تشخیص ہیں، لاگتوں کے بعد منافع کا ثبوت نہیں۔ تجزیہ اوورلیپنگ منافع سے زمانی انحصار، توثیقی ڈیٹا کے بار بار استعمال، اور ایسے گرڈ کا بھی ذکر کرتا ہے جس میں سیکھنے کی شرح اور تربیتی مدت مقرر رکھی گئی ہیں۔

اہم خیالات

  • درخت پہلے سے فراہم کردہ تعاملاتی اصطلاحات کے بغیر فیچر تعاملات سیکھ سکتے ہیں۔
  • حریصانہ درختی تقسیمیں باہم خطی پیش گو کنندگان کو رِج ریگریشن کے مقابلے میں کم موزوں طریقے سے سنبھال سکتی ہیں۔
  • بوسٹنگ کا ہر چیک پوائنٹ ایک الگ امیدوار ہے اور اسے تلاش کے حساب میں شامل کرنا چاہیے۔
  • بہت سی ترتیبیں مقررہ تربیتی مدت سے پہلے عروج پر پہنچتی ہیں، جو ممکنہ اوورفٹنگ کی نشاندہی کرتی ہیں۔
  • معلوماتی نسبت پیش گوئیوں کی درجہ بندی کرتی ہے، مگر لاگت کے بعد قابلِ تجارت کارکردگی ثابت نہیں کرتی۔

ٹیگز

مکمل متن
# ETFs: what trees can find that a penalty cannot, and what they give up to look


# ETFs: what trees can find that a penalty cannot, and what they give up to look

[`06_linear`](06_linear.ipynb) is the one place in these case studies where the linear grid
worked. Ridge at a very large penalty ranked the 99 funds well ahead of zero across the full
validation period, and the shape of that sweep said why: the feature set is **collinear**,
several columns carry nearly the same information, and a dense penalty that spreads weight
across a group of near-duplicates recovers signal that an unregularized fit buries in
offsetting coefficients.

That result frames the question for gradient boosting differently than it would be framed on a
case study where the linear model found nothing. There is a working baseline here, boosting has
to beat it, and the two properties that decide whether it can pull in opposite directions.

**In its favour, trees represent interactions a linear model cannot.** A linear model sees an
interaction only if someone multiplied two columns together and named the product. A tree splits
on one feature inside a region already defined by others, so an interaction is something it
discovers rather than something it is handed. The one this case study has a reason to expect is
the market-stress regime probability from
[`04_model_based_features`](04_model_based_features.ipynb): momentum worth following in a calm
regime may be worth fading in a stressed one, and no single coefficient on momentum expresses
both.

**Against it, collinearity is hostile to trees in a way it is not to ridge.** Faced with several
near-identical columns, a tree picks one of them at each split, and which one is close to
arbitrary. Ridge's advantage on this data came precisely from *not* choosing - from spreading
weight over the correlated group and averaging their noise down. A greedy splitter makes that
choice at every split and remakes it independently on every fold. The same property of the
feature set that made a dense penalty the right answer makes boosting an awkward fit for it.

Three dials control how far the fit goes, and this notebook varies all three:

- **Capacity**, set by `num_leaves`: how many regions one tree may carve the feature space into.
  Seven leaves can express a handful of conditions; 63 can express a partition fine enough to
  describe the training window and nothing beyond it. On a cross-section of 99 funds the top of
  this axis is in the grid as a failure case as much as a candidate.
- **The loss function**, which decides what "got wrong" means. `mse` minimizes squared error,
  `mae` absolute error, and `huber` behaves like squared error for small residuals and like
  absolute error past a threshold derived from each fold's own label spread.
- **When to stop**, set by the number of trees. Unlike a linear fit, a boosted model has a
  meaningful state at every iteration, so each configuration is scored at ten points along its
  own training run rather than only at the end.

The third dial changes how the results must be read. **A checkpoint is part of a configuration,
not a detail of how it was fitted.** Scoring 15 declared configurations at ten checkpoints each
produces 150 candidate models, and treating that as 15 candidates while quietly keeping each
one's best iteration would be reporting the maximum of ten numbers as though it were one.

**Learning objectives.** By the end of this notebook you will be able to:

- Say what a tree ensemble can represent that a penalized linear model cannot, and what a
  collinear feature set costs a greedy splitter.
- Explain why a boosted model produces one result per checkpoint while a linear model produces
  one result in total, and what that implies for counting candidates.
- Read a learning curve of out-of-sample information coefficient against tree count, and tell
  apart a model still learning from one that has begun fitting the training window.
- Say why the choice of loss function is a statement about the label's tails, and relate that to
  what a rank-based metric rewards.
- Recognise that picking each configuration's best checkpoint after seeing the results is a
  selection decision, and locate where selection is actually made.

**Book reference**: Chapter 12, Section 12.2 (GBM libraries) and Section 12.3 (how to tune a
boosted model). Chapter 6, Section 6.7 (Search accounting and run logging) introduces the run
log this notebook writes to.

**Prerequisites**: [`03_financial_features`](03_financial_features.ipynb) and
[`04_model_based_features`](04_model_based_features.ipynb) have written the feature matrices,
[`05_evaluation`](05_evaluation.ipynb) has established the walk-forward folds, and
[`06_linear`](06_linear.ipynb) fitted the linear population this one is compared against.

**What it writes**: one training run per configuration and one complete validation prediction
set per configuration and checkpoint, in `run_log/registry.db` and under `run_log/training/` and
`run_log/predictions/`, grouped under a named population.
[`14_backtest`](14_backtest.ipynb) reads that population and selects on validation backtest
Sharpe. **Selection happens there, not here.**

```python
"""Fit the declared ETF gradient boosting population on the walk-forward folds."""

import numpy as np
import plotly.graph_objects as go
import polars as pl
import yaml
from plotly.subplots import make_subplots

from case_studies.research import (
    declared_labels,
    load_model_configs,
    model_requests,
    narrows_declared_catalog,
    open_study,
    primary_label,
    resolved_model_plan,
    run_model_population,
)
from utils.artifact_specs import resolve_label_horizon
from utils.style import COLORS, show_plotly_with_alt
```

```python
LABELS: list[str] = []
EXECUTION_TIER = "canonical"
WORKSPACE: str = ""
PREVIEW_REDUCTIONS: dict = {}
CONFIG_NAMES: list[str] = []
POPULATION_NAME = ""
SUPERSEDES_POPULATION: str = ""
```

```python
study = open_study("etfs", execution_tier=EXECUTION_TIER, workspace=WORKSPACE or None)
```

## 1. Which labels, and which models

The labels are the ones the linear notebook fitted: `fwd_ret_21d`, the total return over the 21
trading days after the decision date, and the five-day variant `fwd_ret_5d`. Fitting the same
two here is what makes the two populations comparable - the families differ, the targets do
not - and it lets the question `06_linear` raised about the horizon be asked again of a tree
ensemble. Each label carries its own training menu at `config/training/{label}.yaml` and this
notebook fits the union of them; `LABELS` restricts the run to a subset when you want one.

```python
declared_labels(study, "gbm")
```

The menu lists 15 named configurations under `gbm:`, and each resolves to a preset in
`case_studies/config/lgb/`. The grid is a product of two axes:

- **Five capacity profiles.** `default` uses the library's own leaf count; the rest fix it at 7,
  15, 31 and 63.
- **Three objectives**, as described above.

Every configuration runs the same number of boosting iterations at the same learning rate, so
the grid isolates capacity and loss rather than confounding them with training length.

```python
configs = load_model_configs(
    study,
    "gbm",
    labels=LABELS or None,
    config_names=CONFIG_NAMES or None,
)
configs
```

`LABELS` and `CONFIG_NAMES` both narrow what is fitted, and a narrowed run declares a different
set of members than the canonical population does. A population is immutable once written, so
such a run must publish under its own name. Comparing the loaded rows against the complete
declared catalog catches either knob, and says so here rather than several cells later in a
message about hashes.

```python
if narrows_declared_catalog(study, "gbm", configs) and not POPULATION_NAME:
    raise ValueError(
        f"this run declares {configs.height} label-configuration pairs, which is not the "
        f"complete declared catalog, so it cannot publish the canonical population; pass "
        f"POPULATION_NAME to give it its own"
    )
```

## 2. Binding the declarations to the data

Resolving reads the label and feature files, computes the fold boundaries, works out the exact
rows each fit must predict, and turns any data-dependent parameter into the number it will use.
Huber's threshold is one of those: it is a fraction of the training labels' standard deviation,
so it is a different number on every fold and is resolved from that fold's own data.

Nothing is fitted here, so the plan can be inspected first. Four things to check:

- **`feature_count`, `eligible_entities` and `eligible_rows` agree across every row.** A row
  that differs is a configuration measured on a different sample from its neighbours.
- **`folds` is the same everywhere**, and equals the number of walk-forward splits.
- **`validation_start` and `validation_end` bracket the development sample**, with none of the
  held-out tail visible.
- **`checkpoints` is where this differs from the linear plan.** It is the number of training
  states each configuration will publish predictions for. Multiply it by the number of rows to
  get the number of candidate models this notebook is about to create.

```python
requests = model_requests(
    study,
    configs,
    execution_tier=EXECUTION_TIER,
    preview_reductions=PREVIEW_REDUCTIONS,
)
resolved = tuple(request.resolve() for request in requests)

plan = resolved_model_plan(resolved)
plan.select(
    "config_name",
    "feature_count",
    "eligible_entities",
    "eligible_rows",
    "folds",
    "checkpoints",
    "validation_start",
    "validation_end",
)
```

## 3. Fitting the population

`run_model_population` fits every resolved request. For one request it walks the folds, and on
each one:

1. takes the rows inside that fold's training window,
2. casts the design matrix to the precision LightGBM works in and leaves missing values in
   place - a tree routes a missing value down its own branch, so imputing a median here would
   hand the model an observation nobody made,
3. fits the declared number of boosting iterations,
4. predicts the fold's validation rows at each checkpoint, using only the trees built up to that
   iteration.

Step 4 is what makes one fit produce many results. The fold predictions are concatenated into
one series per checkpoint covering the whole validation period, and each becomes its own
registered prediction set with its own identity.

Slicing the window and cleaning its rows depends on the data and not on the model, so it is
work several configurations could share. Whether they do depends on how the requests reach the
runner. Handed over unresolved, they go to a batch path that walks folds on the outside and
configurations on the inside, so one prepared fold serves every configuration and only one is
held at a time. Resolved first - as they are here, so that the plan above can be shown against
the real data - each configuration prepares its own. On a cross-section of this size that costs
seconds and buys a plan that can be read before anything is fitted. On a panel large enough
that one prepared fold set does not comfortably fit in memory the trade runs the other way,
which is why the call also accepts requests that have not been resolved.

**What the call publishes is a population**: a named, immutable list of the prediction sets it
will produce, written down before the first fit. Afterwards every member must exist and be
complete, which is what makes the downstream comparison well defined.

`SUPERSEDES_POPULATION` names the population hash this run replaces. A population is the set
of prediction identities, so anything that moves a training identity - a changed estimator
parameter as much as a changed configuration menu - produces a different population under the
same name, and the registry refuses to write it without being told which snapshot it
supersedes. That lineage is the only record of which generation is which.

It defaults to empty, and the run that published this population passed the predecessor hash
as a parameter instead. The hash is part of what the snapshot is hashed over
(the population registry refuses it), so no default is right for both readers: a re-run that means
to resolve the published population must be handed the same hash the published run used,
while a first run against a fresh registry must be handed nothing. `run_log/` is not in the
repository, so a reader's first run starts from an empty registry, where a non-empty
supersede is refused outright - which is why empty is the default and the hash travels with
the one run that needs it.

A reduced-scale run also leaves it empty. A population produced under a reduction is thrown
away with the workspace it was written to, so it has no lineage to extend.

The default name is the contract with the notebooks downstream - `13_model_analysis` and
`14_backtest` resolve this population by name - rather than a label of convenience, which is
why a run that narrows the member set has to pass its own.

```python
population_name = POPULATION_NAME or "etfs-gbm-validation-v1"
execution, population = run_model_population(
    study,
    resolved,
    population_name=population_name,
    supersedes=SUPERSEDES_POPULATION or None,
)

print(f"{len(execution.runs)} configurations fitted")
print(f"population {population.name}: {len(population.members)} prediction sets")
```

Re-running this notebook unchanged costs the time it takes to read the data. Every identity is
re-derived from the inputs, the registry already holds the matching rows, and the runner returns
the stored result rather than fitting again.

### Running configurations of your own

The published run log is read-only. To add runs, open the study against a workspace, which holds
its own registry and artifacts and reads the same labels and features:

```python
study = open_study("etfs", workspace="~/ml4t-experiments")
configs = load_model_configs(
    study, "gbm", labels=["fwd_ret_21d"], config_names=["leaves_15_huber", "leaves_31_huber"]
)
requests = model_requests(study, configs)
resolved = tuple(request.resolve() for request in requests)
execution, population = run_model_population(study, resolved, population_name="my-gbm-v1")
```

`CONFIG_NAMES` fits a subset of what the menu declares. To fit something new, add a preset at
`case_studies/config/lgb/leaves_127_huber.yaml` and list `leaves_127_huber` under `gbm:` in the
label's menu. Editing an existing preset changes that configuration's identity, so its result
registers as a new row beside the old one rather than replacing it.
[`RUN_LOG.md`](../RUN_LOG.md#running-your-own-configurations) covers the rest.

## 4. What came out

One row per configuration and checkpoint. `ic_mean` is the **information coefficient**: on each
validation date, rank the funds by the model's prediction, rank them by the return they went on
to earn, correlate the two rankings, and average that daily correlation over the validation
period.

The table is sorted by label and then by IC, and the top of each label's block is the trap this
notebook exists to describe. The leading row for a label is the maximum of 150 numbers - fifteen
configurations at ten checkpoints each. Reading it as the result of one experiment would
attribute to the model whatever the stopping point contributed, and the section below measures
how large that contribution is before anything is concluded from the ranking.

`ic_n_days` carries the second warning, and it is the one `06_linear` turned on: a configuration
that scored fewer dates than its neighbours is not comparable to them, because its IC is an
average over the dates where it stayed non-degenerate. Every comparison below is restricted to
full-coverage members for that reason.

Coverage is judged against each label's own maximum number of scorable validation dates. The
horizons do not offer the same number to begin with, since a longer forward window runs out
earlier, so one global maximum would mark a whole label incomplete for a reason unrelated to any
model.

```python
catalog = execution.catalog_rows.select(
    "config_name",
    "label",
    "complete",
    "checkpoint_value",
    "ic_mean",
    "ic_std",
    "ic_n_days",
    "n_folds",
    "training_hash",
    "prediction_hash",
).sort(["label", "ic_mean"], descending=[False, True])

if catalog.filter(~pl.col("complete")).height:
    raise RuntimeError("gbm execution returned a partial prediction set")

catalog = catalog.with_columns(
    full_coverage=pl.col("ic_n_days") == pl.col("ic_n_days").max().over("label")
)

primary = primary_label(study)
present = sorted(set(catalog.get_column("label")))
# The primary label leads when it was fitted. A subset run that leaves it out orders the panels
# by whichever label it did fit rather than by one that is not there.
panel_labels = [label for label in [primary] if label in present] + [
    label for label in present if label != primary
]
order_label = panel_labels[0]
print(f"{catalog.height} candidate models: {catalog.n_unique('config_name')} configurations")
print(f"at {catalog.n_unique('checkpoint_value')} checkpoints each, on {len(panel_labels)} labels")
catalog.select(
    "label",
    "config_name",
    "checkpoint_value",
    "ic_mean",
    "ic_std",
    "ic_n_days",
    "full_coverage",
).head(15)
```

### What more trees do

Each line traces one configuration's out-of-sample IC as trees are added to it. This is the
figure the checkpoint dimension exists to produce, and it separates two things a single
end-of-training number cannot.

A line that rises and then falls has an interior optimum: the model was still learning, then
began fitting the training window at the expense of the validation folds. A line that wanders
without trend around zero never had anything to learn in the first place, and its highest point
is wherever the noise happened to peak. The difference matters, because both produce a
respectable-looking maximum.

One panel per label, because a horizon that has something to learn and one that does not would
be averaged into a single indistinct band if they shared axes.

```python
curves = catalog.filter("full_coverage").sort("label", "config_name", "checkpoint_value")
objectives = {"mse": COLORS["blue"], "mae": COLORS["amber"], "huber": COLORS["copper"]}


def objective_of(name: str) -> str:
    """Read the loss function out of a declared configuration name."""
    match = next((key for key in objectives if name.endswith(key)), None)
    if match is None:
        raise ValueError(f"{name!r} does not end in a declared objective: {sorted(objectives)}")
    return match


fig_curves = make_subplots(
    rows=len(panel_labels),
    cols=1,
    shared_xaxes=True,
    vertical_spacing=0.05,
    subplot_titles=[
        f"{label} ({'primary' if label == primary else 'variant'})" for label in panel_labels
    ],
)
for row, label in enumerate(panel_labels, start=1):
    panel = curves.filter(pl.col("label") == label)
    for objective, color in objectives.items():
        members = [
            name
            for name in panel.get_column("config_name").unique(maintain_order=True)
            if objective_of(name) == objective
        ]
        for index, config_name in enumerate(members):
            series = panel.filter(pl.col("config_name") == config_name)
            fig_curves.add_trace(
                go.Scatter(
                    x=series.get_column("checkpoint_value").to_list(),
                    y=series.get_column("ic_mean").to_list(),
                    mode="lines",
                    name=objective,
                    legendgroup=objective,
                    # One legend entry per loss function, not per configuration: the colour is
                    # the claim, and thirty named lines would bury it.
                    showlegend=row == 1 and index == 0,
                    line=dict(color=color, width=1.5),
                    opacity=0.75,
                ),
                row=row,
                col=1,
            )
    fig_curves.add_hline(
        y=0, line_width=1, line_dash="dash", line_color=COLORS["neutral"], row=row, col=1
    )
    fig_curves.update_yaxes(title_text="Mean IC (validation)", row=row, col=1)
fig_curves.update_xaxes(title_text="Boosting iterations (trees kept)", row=len(panel_labels), col=1)
fig_curves.update_layout(
    title="Validation IC against boosting iteration, by loss function and label",
    height=330 * len(panel_labels),
    width=1000,
    margin=dict(t=90),
    legend=dict(title_text="Loss function"),
)
# Where each panel sits and how many of its lines dip below zero are facts about the frame, so
# the alt text reads them rather than asserting them. A description written once and left alone
# would go stale the first time the fits move.
panel_facts = (
    curves.group_by("label")
    .agg(
        lowest=pl.col("ic_mean").min(),
        highest=pl.col("ic_mean").max(),
        total=pl.col("config_name").n_unique(),
        below=pl.col("config_name").filter(pl.col("ic_mean") < 0).n_unique(),
    )
    .to_dicts()
)
panel_facts = {row["label"]: row for row in panel_facts}
panel_text = ". ".join(
    "The {} panel spans {:+.3f} to {:+.3f}, with {} of its {} lines dipping below zero at some "
    "checkpoint".format(
        label,
        panel_facts[label]["lowest"],
        panel_facts[label]["highest"],
        panel_facts[label]["below"],
        panel_facts[label]["total"],
    )
    for label in panel_labels
)
show_plotly_with_alt(
    fig_curves,
    "Line charts of mean validation information coefficient against boosting iteration, one line "
    "per configuration, coloured by loss function: dark navy for squared error, gold for absolute "
    "error, copper for Huber. One panel per label, each with its own vertical scale and a dashed "
    f"zero line. {panel_text}.",
)
```

Whether the lines trend, and which way, is not something to read off a page of overlapping
curves. The frame below measures it: for each configuration, its IC at the last checkpoint
minus its IC at the first, and whether its best checkpoint is an interior one rather than
either end. A family that is still learning would show positive changes and interior peaks; a
family that is overfitting would show negative changes; a family with nothing to learn would
show changes centred on zero with peaks scattered anywhere.

```python
drift = (
    curves.group_by("label", "config_name")
    .agg(
        first_ic=pl.col("ic_mean").sort_by("checkpoint_value").first(),
        last_ic=pl.col("ic_mean").sort_by("checkpoint_value").last(),
        peak_checkpoint=pl.col("checkpoint_value").sort_by("ic_mean", descending=True).first(),
        first_checkpoint=pl.col("checkpoint_value").min(),
        last_checkpoint=pl.col("checkpoint_value").max(),
    )
    .with_columns(
        change=pl.col("last_ic") - pl.col("first_ic"),
        interior_peak=pl.col("peak_checkpoint").is_between(
            pl.col("first_checkpoint"), pl.col("last_checkpoint"), closed="none"
        ),
    )
)
trees_effect = (
    drift.group_by("label")
    .agg(
        configurations=pl.len(),
        median_change=pl.col("change").median(),
        ended_lower=(pl.col("change") < 0).sum(),
        interior_peaks=pl.col("interior_peak").sum(),
    )
    .sort("label")
)
trees_effect
```

### Comparing every configuration at the same training length

The curves are coloured by objective because that is the axis with a mechanism behind it. If
heavy tails are steering the squared-error fits, the three colours should separate, and they
should separate more as trees are added, since each additional tree is fitted to the residuals
the previous ones left.

The chart below drops the checkpoint dimension by taking each configuration's final state, so
every configuration is compared at the same amount of training. That is the comparison that does
not require choosing anything after the fact. The configurations are held in one order across
the panels - their ranking on the primary label - so a panel that does not descend is a horizon
that orders the grid differently.

```python
final = (
    catalog.filter(pl.col("checkpoint_value") == pl.col("checkpoint_value").max().over("label"))
    .filter("full_coverage")
    .with_columns(objective=pl.col("config_name").map_elements(objective_of, return_dtype=pl.Utf8))
    .sort(["label", "ic_mean"], descending=[False, True])
)
final_iteration = int(final.get_column("checkpoint_value").max())
config_order = (
    final.filter(pl.col("label") == order_label)
    .sort("ic_mean", descending=True)
    .get_column("config_name")
    .to_list()
)

fig_obj = make_subplots(
    rows=len(panel_labels),
    cols=1,
    shared_xaxes=True,
    vertical_spacing=0.09,
    subplot_titles=[
        f"{label} ({'primary' if label == primary else 'variant'})" for label in panel_labels
    ],
)
for row, label in enumerate(panel_labels, start=1):
    panel = final.filter(pl.col("label") == label)
    fig_obj.add_trace(
        go.Bar(
            x=panel.get_column("config_name").to_list(),
            y=panel.get_column("ic_mean").to_list(),
            marker_color=[objectives[value] for value in panel.get_column("objective")],
        ),
        row=row,
        col=1,
    )
    fig_obj.add_hline(
        y=0, line_width=1, line_dash="dash", line_color=COLORS["neutral"], row=row, col=1
    )
    fig_obj.update_yaxes(title_text="Mean IC (validation)", row=row, col=1)
fig_obj.update_xaxes(
    categoryorder="array",
    categoryarray=config_order,
    tickangle=-45,
    title_text=f"Configuration (ordered by validation IC on {order_label})",
    row=len(panel_labels),
    col=1,
)
fig_obj.update_layout(
    title="Validation IC at the final iteration, by loss function and label",
    height=320 * len(panel_labels),
    width=1000,
    showlegend=False,
    margin=dict(t=90),
)
final_text = ". ".join(
    "The {} panel runs from {:+.3f} to {:+.3f} with {} of {} bars above zero, led by {}".format(
        label,
        final.filter(pl.col("label") == label).get_column("ic_mean").min(),
        final.filter(pl.col("label") == label).get_column("ic_mean").max(),
        final.filter((pl.col("label") == label) & (pl.col("ic_mean") > 0)).height,
        final.filter(pl.col("label") == label).height,
        final.filter(pl.col("label") == label)
        .sort("ic_mean", descending=True)
        .row(0, named=True)["config_name"],
    )
    for label in panel_labels
)
show_plotly_with_alt(
    fig_obj,
    "Stacked bar charts of mean validation information coefficient at the final boosting "
    "iteration, coloured by loss function: dark navy for squared error, gold for absolute error, "
    "copper for Huber. The configurations are in the same order in each panel, that order being "
    f"their ranking on {order_label}, and each panel carries a dashed zero line. {final_text}.",
)
```

### Whether the horizon does here what it did to the linear grid

The question this notebook took from `06_linear`, answered at the same training length for
every configuration: whether the horizon that ranked the penalty grid one way ranks a tree
ensemble the same way. `rank_correlation` is between the two labels' orderings of the
configurations they both charted, so one would mean the horizons agree on the whole grid.

```python
horizons = (
    final.group_by("label")
    .agg(
        configurations=pl.len(),
        best_config=pl.col("config_name").sort_by("ic_mean", descending=True).first(),
        best_ic=pl.col("ic_mean").max(),
        worst_ic=pl.col("ic_mean").min(),
        above_zero=(pl.col("ic_mean") > 0).sum(),
    )
    .sort("label")
)
horizons
```

```python
pair = (
    final.select("label", "config_name", "ic_mean")
    .pivot(on="label", index="config_name", values="ic_mean")
    .drop_nulls()
)
if len(panel_labels) > 1:
    print(f"{pair.height} configurations charted at both horizons")
    for index, label_a in enumerate(panel_labels):
        for label_b in panel_labels[index + 1 :]:
            correlation = pair.select(
                pl.corr(pl.col(label_a).rank(), pl.col(label_b).rank())
            ).item()
            print(f"rank correlation {label_a} against {label_b}: {correlation:+.3f}")
else:
    print("one label charted, so there is no horizon comparison to make")
```

### Whether the loss function is what separates them

The colours in the chart above are the claim, so here is the claim as numbers: at the final
iteration, each loss function's mean IC and how many of its configurations finished above zero,
per label. The ordering a heavy-tailed label produces is Huber highest, absolute error next,
squared error lowest, because a squared-error fit weights a residual by its square and one
extreme observation therefore moves it more than many ordinary ones, while a rank measure gives
nothing back for chasing that extreme.

```python
objective_summary = (
    final.group_by("label", "objective")
    .agg(
        configurations=pl.len(),
        mean_ic=pl.col("ic_mean").mean(),
        best_ic=pl.col("ic_mean").max(),
        above_zero=(pl.col("ic_mean") > 0).sum(),
    )
    .sort(["label", "mean_ic"], descending=[False, True])
)
objective_summary
```

### What the tails of each label look like

The mechanism above is a claim about the label, not the model, so it is checkable without
fitting anything. Excess kurtosis measures how much of a distribution's variance comes from
rare large moves; where the loss functions separate, the heavier-tailed label should separate
them further. That is the prediction; section 5 reports what the two frames actually show.

Each label is cut on the date it **resolves**, not the date it is observed, and against its
own development boundary rather than a boundary shared with the other label. Both parts
matter here. A 21-session label observed inside the development window can have its forward
window closing after the boundary, which makes it a holdout row that a filter on the
observation date would have kept - the rule `02_labels` states and applies. And the two
labels resolve their last validation fold on different dates, so one shared cut would
describe `fwd_ret_21d` rows the 21-day fits never saw.

```python
development_end = {
    label: plan.filter(pl.col("label") == label).get_column("validation_end").max()
    for label in panel_labels
}
setup = yaml.safe_load((study.root / "config" / "setup.yaml").read_text())
label_horizon = {
    label: int(resolve_label_horizon("etfs", label, setup).rstrip("Dd")) for label in panel_labels
}
tails = pl.DataFrame(
    [
        {
            "label": label,
            "observations": len(values),
            "std": float(values.std(ddof=1)),
            "excess_kurtosis": float(
                ((values - values.mean()) ** 4).mean() / values.var(ddof=0) ** 2 - 3.0
            ),
            "share_beyond_4_sd": float(
                (np.abs(values - values.mean()) > 4 * values.std(ddof=1)).mean()
            ),
        }
        for label, values in (
            (
                label,
                study.labels.get(label, execution_tier=EXECUTION_TIER)
                .load()
                .with_columns(
                    pl.col("timestamp")
                    .shift(-label_horizon[label])
                    .over("symbol")
                    .alias("_label_end")
                )
                .filter(pl.col("_label_end") <= development_end[label])
                .get_column(label)
                .drop_nulls()
                .to_numpy(),
            )
            for label in panel_labels
        )
    ]
).sort("label")
tails
```

### How much the checkpoint moves a configuration

One number per configuration: the range its IC covers across its own ten checkpoints. This is
the quantity that decides whether choosing a stopping point is a decision worth making carefully
or one being made by noise. A configuration whose IC varies more across its own training run
than the configurations vary among themselves is one where the checkpoint, not the model, is
doing the ranking. The comparison is within a label, since the two horizons do not produce ICs
of the same size.

```python
spread = (
    curves.group_by("label", "config_name")
    .agg(
        ic_min=pl.col("ic_mean").min(),
        ic_max=pl.col("ic_mean").max(),
        ic_final=pl.col("ic_mean").sort_by("checkpoint_value").last(),
    )
    .with_columns(checkpoint_range=pl.col("ic_max") - pl.col("ic_min"))
    .sort(["label", "checkpoint_range"], descending=[False, True])
)
checkpoint_vs_config = (
    spread.group_by("label")
    .agg(
        across_configs=pl.col("ic_final").max() - pl.col("ic_final").min(),
        median_within_config=pl.col("checkpoint_range").median(),
    )
    .with_columns(
        checkpoint_dominates=pl.col("median_within_config") > pl.col("across_configs"),
    )
    .sort("label")
)
checkpoint_vs_config
```

```python
spread
```

## 5. What to notice

**Almost every configuration ranks the cross-section the right way, and the loss function
orders the families.** `above_zero` in the `horizons` frame is 14 of 15 on both labels, so one
configuration on each ends training pointing the wrong way, by a margin (`worst_ic`) small
enough to be indistinguishable from zero. Read `objective_summary` for the ordering: on both
labels Huber has the highest `mean_ic`, absolute error is next and squared error is last, and
squared error is also the only objective whose `above_zero` is not 5 of 5. The families order
on their averages but they overlap on individual configurations, so this is a statement about
objectives and not a ranking of the fifteen. Squared error weights an observation by the square
of its error, so the largest moves dominate what each successive tree is fitted to, while the
information coefficient is a rank correlation that cares about order rather than magnitude:
effort spent getting the extremes right buys nothing on this metric. **An objective is a claim
about which errors matter, and it is worth choosing to match the metric the result will be
judged on.**

**The ordering is what the tails predict; the size of the gap is not.** The prediction in
section 4 was that the heavier-tailed label should separate the objectives further. `tails`
puts `fwd_ret_5d` at excess kurtosis 10.26 against `fwd_ret_21d`'s 7.00, so five-day is the
heavier-tailed label. But the Huber-minus-squared-error gap in `objective_summary` is 0.0143
on `fwd_ret_21d` and 0.0075 on `fwd_ret_5d` - the heavier-tailed label separates them by
about half as much. The direction survives on both labels and the magnitude runs the other
way, so tail heaviness is not what sets how far apart the objectives land. Whatever does is
not measured here: the two labels differ in horizon, in autocorrelation and in how much of
each one a fold contains, and this notebook holds none of those fixed.

**The boosted population does not beat the linear one.** The strongest full-coverage
configuration here ranks below the strongest full-coverage configuration in
[`06_linear`](06_linear.ipynb), on the same label, the same features and the same folds. This is
the outcome the collinearity argument at the top of this notebook implies, and it is worth
taking seriously rather than treating as a tuning failure. Ridge earned its result by refusing
to choose among near-duplicate columns - spreading weight across a correlated group averages
their noise down. A tree chooses one column at every split, and remakes that choice
independently on every fold, which is the opposite operation on the same feature set. The
interactions the trees can represent and a linear model cannot are real, but on this data they
do not pay for what greedy splitting gives up.

**Almost every configuration is past its peak by the time it is reported.** `ended_lower` in
the `trees_effect` frame is 14 of 15 on both labels, and `median_change` is negative on both,
so the typical configuration is worse at the declared training length than it was early on.
`interior_peaks` says the descent is not uniform: 3 configurations on fwd_ret_21d and 5 on
fwd_ret_5d reach their best IC somewhere in the middle of the run rather than at the first
checkpoint. So the declared training length is longer than this data supports, and the
fixed-iteration comparison above is a comparison of fifteen models mostly in their overfitted
regime.
That comparison is still the honest one, because each configuration is measured at the same
amount of training and nothing is chosen after the fact - but the ranking it produces is not the
ranking their best states would produce, and neither is a result until something selects on a
criterion that was fixed in advance.

**The checkpoint moves the IC nearly as far as the model does.** `checkpoint_vs_config`
compares the two: `across_configs` is the IC range over the fifteen configurations at the
declared training length, `median_within_config` is the median range a single configuration
covers over its own ten checkpoints. The second is about five eighths of the first on both
labels, which is why `checkpoint_dominates` is false and also why it is not reassuring. A
stopping point chosen after seeing the curves would be doing a comparable amount of work to the
choice of model, which is why reporting the leading row of the results table would be reporting
the maximum of 150 numbers as though it were one.

**None of this selects anything.** IC measures whether predictions rank funds correctly, not
whether a strategy trading them makes money after costs and turnover. A monthly-horizon signal
that reorders the portfolio at every rebalance is exactly where a small ranking advantage is
lost. Selection is on validation backtest Sharpe over the population just published, and it
happens in [`14_backtest`](14_backtest.ipynb), where the checkpoint is part of what is selected.

**Known limitations.** The IC is an average of daily rank correlations with no adjustment for
the serial dependence overlapping 21-day returns create, so it is a diagnostic rather than a
test, and it carries no interval that would say whether these configurations differ from each
other or from the linear ones. The grid varies capacity and loss at a fixed learning rate and a
fixed training length, so it says nothing about trading one against another - and the previous
point is the reason that matters here. And every number is measured on validation folds that
have been read many times over by the time a case study reaches this notebook.

**Next**: [`08_tabular_dl`](08_tabular_dl.ipynb) asks the same question of a neural network
built for tabular data, which represents interactions in a third way again. The useful thing to
watch there is whether it recovers whatever structure the trees found here, and whether a
collinear feature set costs it what it costs a greedy splitter.
![notebook output](figures/p1_1.png)
![notebook output](figures/p1_2.png)

ماخذ کا حوالہ دیتے ہوئے مکمل متن دکھایا گیا ہے، ماخذ کے لائسنس کے تحت۔ لائسنس: MIT

یہ خلاصہ اصل ماخذ سے Stratmill کے تحقیقی ایجنٹ نے لکھا ہے؛ یہ ماخذ کی نقل نہیں۔