सामग्री पर जाएं
लाइब्रेरी के सभी दस्तावेज़

ट्रेडिंग पाइपलाइन में डेटा कवरेज और विशेषता गुणवत्ता की जाँच

कोड Machine Learning for Trading

सारांश

यह उपयोगिता क्वांटिटेटिव रिसर्च पाइपलाइन से बने डेटा आर्टिफ़ैक्ट की निदान प्रक्रिया बताती है। यह संख्यात्मक कॉलम में रिक्त मान, गैर-सीमित मान, शून्य, अलग-अलग मानों की संख्या, क्वांटाइल और टेल व्यवहार का प्रोफ़ाइल बनाती है, फिर विन्यास योग्य सलाहकारी सीमाएँ पार करने वाले कॉलम चिह्नित करती है। यह तैयार की गई कुंजियों की तुलना स्पष्ट रूप से घोषित अपेक्षित समूह से भी करती है और अनुपस्थित तथा अप्रत्याशित कुंजियाँ रिपोर्ट करती है, ताकि मौजूदा पंक्तियाँ पूरी होने भर से कम आकार का आउटपुट स्वस्थ न दिखे।

रिपोर्ट घोषित अपेक्षाओं के आधार पर इकाई और सत्र के अनुसार अनुपस्थित अवलोकनों का वर्गीकरण कर सकती है, ताकि नियोजित अंतरालों को अस्पष्ट शेष से अलग किया जा सके। अलग क्रॉस-सेक्शनल फैलाव जाँच उन सत्रों की पहचान करती है जहाँ पूर्वानुमान में इकाइयों के बीच कोई भिन्नता नहीं है और इसलिए उनकी क्रमबद्धता नहीं हो सकती। ये जाँचें नोटबुक लेखक को व्याख्या के लिए प्रमाण देती हैं; वे तय नहीं करतीं कि कोई रिक्त मान या शून्य दोषपूर्ण है, और न ही सलाहकारी फ़्लैग पाइपलाइन रोकते हैं। दस्तावेज़ डेटा गुणवत्ता को गणना और वितरण से समर्थित संदर्भ-आधारित अनुमोदन के रूप में प्रस्तुत करता है, जबकि कठोर अस्वीकृति नियम अलग मान्यकरण द्वारों के लिए रखता है।

मुख्य विचार

  • अनुपस्थित और अप्रत्याशित पंक्तियाँ दिखाने के लिए तैयार कुंजियों की तुलना घोषित समूह से करें।
  • हर संख्यात्मक कॉलम में रिक्त मान, गैर-सीमित मान, शून्य, अलग-अलग मानों की गिनती, क्वांटाइल और टेल अनुपात का प्रोफ़ाइल बनाएँ।
  • सीमा फ़्लैग को स्पष्टीकरण माँगने वाले संकेत मानें, क्योंकि स्वीकार्य मान विशेषता पर निर्भर होते हैं।
  • अपेक्षित अंतरालों को अस्पष्ट अंतरालों से अलग करने के लिए अनुपस्थित कुंजियों को इकाई और सत्र के अनुसार वर्गीकृत करें।
  • सत्र के भीतर पूर्वानुमान का फैलाव जाँचें, क्योंकि इकाइयों की क्रमबद्धता न कर सकने वाला पूर्वानुमान अनुपयोगी हो सकता है।

टैग

पूरा पाठ
# artifact_quality.py


```py
"""Descriptive statistics and coverage for what a production stage just wrote.

Three questions, asked at the end of the notebook that produces an artifact, about
the frame it is about to write:

1. **Is it the size it should be?** Not "are the rows it has non-null" - a stage that
   emits a thousand rows where a million were expected reports no missing values and
   is wrong. ``coverage_against`` compares the produced keys to the universe the
   caller declares.
2. **What is in each column?** ``profile_columns`` gives one row per column: nulls,
   zeros, non-finites, distinct count, quantiles and a tail ratio.
3. **What is unusual?** ``flag_columns`` marks the profile rows that cross a
   threshold, with the reason attached.

Nothing here raises. A zero share of 0.5 is a defect in a return and correct in a
binary direction label; a 45% null share is a defect in a price and expected in a
term-structure feature that needs two expiries. Only the notebook author knows
which, so this surfaces the number and the author signs off on it in prose. The
gates that *do* refuse live in ``coverage.py`` and ``notebook_contracts.py``.

Polars in, polars out. Pandas only if a caller is at a visualization boundary.
"""

from __future__ import annotations

from collections.abc import Sequence
from pathlib import Path

import polars as pl

__all__ = [
    "coverage_against",
    "cross_sectional_dispersion",
    "explain_missing",
    "flag_columns",
    "label_universe",
    "profile_columns",
    "quality_report",
    "render_cross_sectional_dispersion",
    "render_quality_report",
]

#: Quantiles reported for every numeric column. The outer pair is what makes a
#: heavy tail visible; the inner three are the shape.
DEFAULT_QUANTILES: tuple[float, ...] = (0.001, 0.01, 0.25, 0.5, 0.75, 0.99, 0.999)

#: Advisory thresholds. Deliberately loose - these decide what a reader is asked to
#: look at, not what is acceptable. Tighten per notebook by passing your own.
DEFAULT_RULES: dict[str, float] = {
    "null_share": 0.20,
    "zero_share": 0.60,
    "tail_ratio": 25.0,
}


def _numeric_columns(frame: pl.DataFrame, exclude: Sequence[str]) -> list[str]:
    excluded = set(exclude)
    return [c for c, dtype in frame.schema.items() if dtype.is_numeric() and c not in excluded]


def profile_columns(
    frame: pl.DataFrame,
    *,
    key_columns: Sequence[str] = (),
    quantiles: Sequence[float] = DEFAULT_QUANTILES,
) -> pl.DataFrame:
    """One row per numeric column: how much is there, and what shape is it.

    ``key_columns`` are excluded - a symbol id or a fold index has no distribution
    worth reporting. Non-finite values are counted separately from nulls: a null is
    an absent observation and an infinity is a computed one that overflowed, and the
    two have different causes even though most downstream code treats them alike.

    ``tail_ratio`` is ``|q99.9| / |q75|``, which is a scale-free way to see a column
    whose extreme values are orders of magnitude past its body. It is null where the
    denominator is zero rather than infinite, so the column stays sortable.
    """
    columns = _numeric_columns(frame, key_columns)
    if not columns:
        return pl.DataFrame(schema={"column": pl.String})

    rows: list[dict[str, object]] = []
    height = frame.height
    for name in columns:
        series = frame[name]
        finite = series.drop_nulls()
        n_nonfinite = int(finite.is_infinite().sum() or 0) + int(finite.is_nan().sum() or 0)
        clean = finite.filter(finite.is_finite())
        row: dict[str, object] = {
            "column": name,
            "rows": height,
            "n_null": int(series.null_count()),
            "null_share": series.null_count() / height if height else None,
            "n_nonfinite": n_nonfinite,
            "n_zero": int((clean == 0).sum() or 0),
            "zero_share": (float((clean == 0).sum() or 0) / clean.len()) if clean.len() else None,
            "n_distinct": int(clean.n_unique()),
            "constant": clean.n_unique() <= 1 if clean.len() else None,
            "mean": float(clean.mean()) if clean.len() else None,
            "std": float(clean.std()) if clean.len() > 1 else None,
        }
        for q in quantiles:
            row[f"q{q:g}"] = float(clean.quantile(q)) if clean.len() else None
        body = row.get("q0.75")
        tail = row.get("q0.999")
        row["tail_ratio"] = (
            abs(tail) / abs(body) if body not in (None, 0) and tail is not None else None
        )
        rows.append(row)
    return pl.DataFrame(rows)


def coverage_against(
    produced: pl.DataFrame,
    expected: pl.DataFrame,
    *,
    keys: Sequence[str],
) -> dict[str, pl.DataFrame]:
    """How much of the declared universe the stage actually produced.

    ``expected`` is the universe the caller declares this stage owes a row for - the
    label artifact restricted to a fold's validation window, a trading calendar
    crossed with the symbol universe, the feature panel a model is scored on. Both
    frames are reduced to their distinct ``keys`` before comparison, so passing a
    frame with extra columns is fine.

    Returns ``summary`` with the counts, plus ``missing`` and ``unexpected`` carrying
    the key tuples themselves. The tuples are what make the number readable: a
    handful of symbols absent entirely and every symbol missing its first weeks give
    the same percentage and are a universe restriction and a warm-up respectively.
    Group ``missing`` by the entity column to tell them apart.
    """
    key_list = list(keys)
    mismatched = [
        (k, expected.schema[k], produced.schema[k])
        for k in key_list
        if expected.schema[k] != produced.schema[k]
    ]
    if mismatched:
        detail = "; ".join(
            f"{k}: expected {want} but produced {got}" for k, want, got in mismatched
        )
        raise ValueError(
            "key dtypes differ between the produced frame and the declared universe, so "
            f"every row would read as missing - {detail}. Cast explicitly in the notebook: "
            "a silent cast here is how a Date label artifact and a Datetime prediction set "
            "come to report full coverage of nothing."
        )
    got = produced.select(key_list).unique()
    want = expected.select(key_list).unique()
    missing = want.join(got, on=key_list, how="anti")
    unexpected = got.join(want, on=key_list, how="anti")
    summary = pl.DataFrame(
        [
            {
                "expected": want.height,
                "produced": got.height,
                "missing": missing.height,
                "unexpected": unexpected.height,
                "coverage": (want.height - missing.height) / want.height if want.height else None,
            }
        ]
    )
    return {"summary": summary, "missing": missing, "unexpected": unexpected}


def flag_columns(
    profile: pl.DataFrame,
    *,
    rules: dict[str, float] | None = None,
) -> pl.DataFrame:
    """The profile rows a reader should look at, with the reason attached.

    A flag is a request for a sentence of explanation, not a failure. Every rule is a
    ceiling on a column of ``profile``; a column crossing none of them is absent from
    the result. Non-finite values and constants are always flagged, because neither
    has a legitimate reading that does not need saying out loud.
    """
    active = DEFAULT_RULES | (rules or {})
    reasons: list[pl.Expr] = [
        pl.when(pl.col("n_nonfinite") > 0)
        .then(pl.format("{} non-finite values", pl.col("n_nonfinite")))
        .otherwise(None),
        pl.when(pl.col("constant")).then(pl.lit("constant over every row")).otherwise(None),
    ]
    for column, ceiling in active.items():
        if column not in profile.columns:
            continue
        reasons.append(
            pl.when(pl.col(column) > ceiling)
            .then(pl.format(f"{column} {{}} above {ceiling:g}", pl.col(column).round(4)))
            .otherwise(None)
        )
    flagged = profile.with_columns(pl.concat_list(reasons).list.drop_nulls().alias("flags")).filter(
        pl.col("flags").list.len() > 0
    )
    return flagged.select(
        "column", pl.col("flags").list.join("; ").alias("why"), pl.exclude("column", "flags")
    )


def quality_report(
    frame: pl.DataFrame,
    *,
    name: str,
    key_columns: Sequence[str] = (),
    expected: pl.DataFrame | None = None,
    keys: Sequence[str] | None = None,
    rules: dict[str, float] | None = None,
    entity: str | Sequence[str] | None = None,
    session: str | None = None,
    expected_missing: dict[str, MissingExpectation] | None = None,
) -> dict[str, pl.DataFrame]:
    """Profile, flags, coverage and (when the shortfall is declared) its explanation.

    ``entity`` and ``session`` name the two axes of the key; ``entity`` takes a sequence where
    the entity is compound, as it is for an option contract keyed by underlying and instrument or
    a futures curve keyed by product and position. Given them, the missing
    keys are classified by ``explain_missing`` and ``expected_missing`` declares which
    positions a mechanism accounts for, so the report distinguishes a shortfall that is
    the design from one that is a defect. Without them the report still carries the
    coverage counts, and the notebook is left asserting a percentage is acceptable
    without saying what produced it.

    Returns the frames rather than printing them, so the notebook decides how to
    display each and writes its own sign-off around them.
    """
    profile = profile_columns(frame, key_columns=key_columns)
    report: dict[str, pl.DataFrame] = {
        "profile": profile,
        "flags": flag_columns(profile, rules=rules),
    }
    if expected is not None:
        coverage = coverage_against(frame, expected, keys=keys or list(key_columns))
        report.update({f"coverage_{part}": value for part, value in coverage.items()})
        if entity is not None and session is not None and coverage["missing"].height:
            explanation = explain_missing(
                coverage["missing"],
                frame,
                entity=entity,
                session=session,
                expected=expected_missing,
            )
            report.update({f"missing_{part}": value for part, value in explanation.items()})
    report["name"] = name
    return report


def cross_sectional_dispersion(
    frame: pl.DataFrame,
    *,
    value: str,
    entity: str = "symbol",
    session: str = "timestamp",
    eps: float | None = None,
) -> dict:
    """Does this forecast order anything at each decision time?

    ``profile_columns`` cannot answer that. Its ``constant`` flag is
    ``n_unique() <= 1`` over the whole column, so a forecast that is constant
    within every timestamp and merely drifts across them profiles as varied. That
    is the exact shape that ranks nothing: a top-k book built from it fills once and
    never replaces, trades between zero and two times, pays almost no cost, and can
    top a cost-dominated leaderboard on the strength of not trading.

    So the dispersion has to be measured **within** ``session``, across ``entity``,
    and the test is a standard deviation at or below ``eps``. That covers both shapes
    with one clause: an exactly-constant cross-section has a standard deviation of
    zero, and a regularized fit that leaves float noise instead of a constant carries
    1.95e-17 to 2.02e-17 - measured on the 24 denormal rows across all nine
    registries, against a smallest genuine value of 1.69e-03.

    An ``n_distinct == 1`` clause was written beside it first and removed: mutation
    showed it dead, because every cross-section it catches the standard deviation
    already catches. ``n_distinct`` is still reported, because a reader looking at a
    flagged session wants to know which of the two shapes it is.

    ``eps`` defaults to the same constant the registry's degeneracy rule uses, imported
    rather than restated so the two cannot drift apart. The import is deferred because
    this module is loaded by stage-02 notebooks and deliberately depends on nothing but
    polars at module level.

    Returns ``per_session``, one row per decision time, and ``summary``. **It does not
    raise**, in keeping with the rest of this module: a prediction set that ranks
    nothing at 3 of 500 timestamps may be a defect or may be three holidays, and only
    the notebook author knows which. The gate that refuses lives in
    ``notebook_contracts.degenerate_prediction_sql``.
    """
    if eps is None:
        from case_studies.utils.notebook_contracts import _DEGENERATE_IC_EPS

        eps = _DEGENERATE_IC_EPS

    missing = [c for c in (value, entity, session) if c not in frame.schema]
    if missing:
        raise KeyError(f"frame has no column(s) {missing}; it carries {sorted(frame.schema)}")

    per_session = (
        frame.select(session, entity, value)
        .drop_nulls(value)
        .group_by(session)
        .agg(
            n_entities=pl.len(),
            n_distinct=pl.col(value).n_unique(),
            std=pl.col(value).std(),
            spread=pl.col(value).max() - pl.col(value).min(),
        )
        .with_columns(
            # A session carrying one entity is not a failure to rank, it is a session
            # with nothing to rank. Marking it degenerate would report a universe
            # shrinking to one name as a model defect.
            degenerate=(pl.col("n_entities") > 1) & (pl.col("std").fill_null(0.0) <= eps),
        )
        .sort(session)
    )

    rankable = per_session.filter(pl.col("n_entities") > 1)
    n_degenerate = int(per_session.get_column("degenerate").sum())
    summary = {
        "value": value,
        "n_sessions": per_session.height,
        "n_rankable_sessions": rankable.height,
        "n_degenerate": n_degenerate,
        "share_degenerate": (n_degenerate / rankable.height) if rankable.height else None,
        "min_std": float(rankable.get_column("std").min()) if rankable.height else None,
        "median_std": float(rankable.get_column("std").median()) if rankable.height else None,
        "eps": eps,
    }
    return {"per_session": per_session, "summary": summary}


#: Table width for every frame these renderers print. Polars defaults to 80 characters,
#: which wraps a 19-column profile into an unreadable stack of fragments - the table is
#: there to be read at a glance and at 80 it cannot be. Notebook output has the room.
WIDE = 240


def render_quality_report(report: dict, *, max_rows: int = 80) -> None:
    """Print a ``quality_report`` as a notebook reader should read it: shortfall first.

    Coverage leads because it is the question the other tables cannot answer - a frame
    whose every column profiles cleanly is still wrong if it is a tenth of the universe.
    Where the shortfall has been classified, the explanation prints next to it and the
    residual is stated as its own number, because "76% coverage" and "76% coverage, all
    of it a declared 252-session estimation burn-in, nothing unexplained" are the same
    percentage and different findings. Flags come next as the shortlist a sign-off has
    to speak to, and the full profile last.
    """
    print(f"=== {report['name']} ===")

    summary = report.get("coverage_summary")
    if summary is not None:
        row = summary.row(0, named=True)
        share = "n/a" if row["coverage"] is None else f"{row['coverage']:.2%}"
        print(
            f"coverage {share}: {row['produced']:,} of {row['expected']:,} declared keys, "
            f"{row['missing']:,} missing, {row['unexpected']:,} unexpected"
        )
        explained = report.get("missing_summary")
        if explained is not None and explained.height:
            print("where the missing keys sit, and what accounts for them:")
            with pl.Config(
                tbl_rows=8,
                tbl_cols=9,
                fmt_str_lengths=40,
                tbl_width_chars=WIDE,
                tbl_hide_dataframe_shape=True,
            ):
                print(explained)
            undeclared = int(explained.filter(pl.col("why") == "not declared")["n_keys"].sum() or 0)
            excess = int(explained["excess_keys"].sum() or 0)
            over = int(explained["entities_over"].sum() or 0)
            print(
                f"residual to explain: {undeclared:,} key(s) in positions no mechanism was "
                f"declared for, and {excess:,} spent past a declared budget by {over} entities"
            )
        else:
            missing = report.get("coverage_missing")
            if missing is not None and missing.height:
                # Every key dimension, not the first one. A shortfall concentrated in a
                # few entities and one falling on whole timestamps are different defects
                # - a universe restriction against a feed outage - and reporting only the
                # column written first shows one and hides the other.
                for key in missing.columns:
                    per_key = missing.group_by(key).len().sort("len", descending=True)
                    print(f"  {per_key.height} distinct {key} values are short; worst:")
                    with pl.Config(tbl_rows=8, tbl_hide_dataframe_shape=True):
                        print(per_key.head(8))

    flags = report["flags"]
    if flags.height == 0:
        print("no column crossed a threshold")
    else:
        print(f"{flags.height} column(s) to speak to:")
        with pl.Config(
            tbl_rows=max_rows,
            tbl_cols=4,
            fmt_str_lengths=110,
            tbl_width_chars=WIDE,
            tbl_hide_dataframe_shape=True,
        ):
            print(flags.select("column", "why", "rows", "n_null"))

    with pl.Config(
        tbl_rows=max_rows, tbl_cols=20, tbl_width_chars=WIDE, tbl_hide_dataframe_shape=True
    ):
        print(report["profile"])


#: What a declared expectation looks like: how many sessions per entity the mechanism
#: is allowed to cost, and the mechanism's name. ``None`` for the budget means the
#: mechanism has no per-entity ceiling worth stating - an entity absent from the
#: universe for its whole life, say - and the position is then explained but unbounded.
MissingExpectation = tuple[int | None, str]


def explain_missing(
    missing: pl.DataFrame,
    produced: pl.DataFrame,
    *,
    entity: str | Sequence[str],
    session: str,
    expected: dict[str, MissingExpectation] | None = None,
) -> dict[str, pl.DataFrame]:
    """Split the missing keys into the ones a mechanism accounts for and the residual.

    A coverage percentage on its own decides nothing. A stage whose features need a
    year of history before they resolve, against a five-year sample, owes no row for
    the first fifth of it and is complete at 80%; a stage losing the same fraction
    scattered through the middle of every symbol's quoted range is broken at the same
    number. The two are told apart by *where* the key sits relative to what the entity
    did produce, so that is what this classifies:

    ``leading``
        Before the entity's first produced row. A lookback, a burn-in, an estimation
        window, or a listing date later than the sample start.
    ``trailing``
        After its last produced row. A forward label's horizon running off the end of
        the sample, or a delisting.
    ``interior``
        Inside the entity's own produced span, with rows on both sides. No window
        length explains one of these.
    ``absent``
        The entity produced no row at all. A universe restriction, and the case a
        percentage misleads on hardest: it reads as thin coverage of everything rather
        than no coverage of some.

    ``expected`` declares the mechanism per position and the sessions per entity it may
    cost: ``{"trailing": (10, "10-session forward label horizon")}``. A position with no
    entry is unexplained in full. An entity spending more than its budget contributes
    its keys to ``residual`` and its excess to the summary, and that is what a sign-off
    speaks to - the rest is design, which is what declaring it says.
    """
    entity_cols = [entity] if isinstance(entity, str) else list(entity)
    span = produced.group_by(entity_cols).agg(
        pl.col(session).min().alias("_first"),
        pl.col(session).max().alias("_last"),
    )
    classified = (
        missing.join(span, on=entity_cols, how="left")
        .with_columns(
            pl.when(pl.col("_first").is_null())
            .then(pl.lit("absent"))
            .when(pl.col(session) < pl.col("_first"))
            .then(pl.lit("leading"))
            .when(pl.col(session) > pl.col("_last"))
            .then(pl.lit("trailing"))
            .otherwise(pl.lit("interior"))
            .alias("where")
        )
        .drop("_first", "_last")
    )

    declared = expected or {}
    per_entity = classified.group_by([*entity_cols, "where"]).len().rename({"len": "sessions"})
    rows: list[dict[str, object]] = []
    residual_frames: list[pl.DataFrame] = []
    for position in ("absent", "leading", "interior", "trailing"):
        part = per_entity.filter(pl.col("where") == position)
        if part.height == 0:
            continue
        budget, why = declared.get(position, (0, "not declared"))
        over = part.filter(pl.col("sessions") > budget) if budget is not None else part.head(0)
        excess = int((over["sessions"] - budget).sum()) if budget is not None and over.height else 0
        rows.append(
            {
                "where": position,
                "n_keys": int(part["sessions"].sum()),
                "n_entities": part.height,
                "sessions_p50": float(part["sessions"].median()),
                "sessions_max": int(part["sessions"].max()),
                "budget": budget,
                "entities_over": over.height,
                "excess_keys": excess,
                "why": why,
            }
        )
        if over.height:
            residual_frames.append(
                classified.filter(pl.col("where") == position).join(
                    over.select(entity_cols), on=entity_cols, how="semi"
                )
            )

    residual = pl.concat(residual_frames) if residual_frames else classified.head(0)
    return {
        "classified": classified,
        "per_entity": per_entity,
        "summary": pl.DataFrame(rows) if rows else pl.DataFrame(schema={"where": pl.String}),
        "residual": residual,
    }


def label_universe(
    case_dir: Path | str,
    *,
    labels: Sequence[str] | None = None,
    keys: Sequence[str] = ("symbol", "timestamp"),
) -> pl.DataFrame:
    """The distinct keys the case study's labels declare, as an ``expected`` frame.

    This is what a feature stage owes a row for: a key carrying a label and no feature
    row is one no model can be asked to score. Reading it here rather than trusting the
    stage's own inputs is the whole point - a panel can be internally consistent, carry
    no nulls, and still be missing a fifth of the universe the labels declare.

    ``labels`` names the artifacts to union; the default is every parquet in
    ``labels/``. A row is counted whether or not its label value is null, because a
    stage does not know which labels a later stage will filter on.

    No dtype coercion happens here or in ``coverage_against``. A label artifact keying
    ``timestamp`` as ``Date`` against a feature panel keying it as ``Datetime`` is a real
    disagreement, and casting it away is how a frame reports full coverage of nothing.
    """
    directory = Path(case_dir) / "labels"
    paths = (
        [directory / f"{name}.parquet" for name in labels]
        if labels is not None
        else sorted(directory.glob("*.parquet"))
    )
    frames = [
        pl.scan_parquet(path).select(list(keys)).unique().collect()
        for path in paths
        if path.is_file()
    ]
    if not frames:
        raise FileNotFoundError(f"no label artifact under {directory}; the universe is undefined")
    return pl.concat(frames).unique()


def explain_column_gaps(
    frame: pl.DataFrame,
    *,
    columns: Sequence[str],
    entity: str | Sequence[str],
    session: str,
    expected: dict[str, dict[str, MissingExpectation]] | None = None,
) -> dict[str, pl.DataFrame]:
    """Where each column's nulls sit relative to the rows that column did produce.

    ``coverage_against`` asks whether the stage wrote a row for a key. A stage that
    writes one row per key and leaves the estimation window null answers that question
    perfectly and still ships an artifact whose feature is absent for a tenth of the
    panel, because the key is present and only the value is not. This asks the same
    positional question of each column on its own: the keys where the column is null
    are its missing set, the keys where it carries a value are what it produced, and
    ``explain_missing`` classifies the difference. A fitted feature's burn-in and a hole
    in the middle of a fitted series then read as the two different things they are.

    Declaring per column is what makes the answer usable, because the families in one
    artifact do not lose the same rows: a rolling model owes its burn-in, a
    market-wide series owes everything before the series existed, and a spread against
    a quoted surface owes every key the surface does not quote. ``expected`` therefore
    maps a column name to the per-position declaration ``explain_missing`` takes. A
    column with no entry is still measured; nothing is declared for it, so all of its
    nulls land in the residual.
    """
    entity_cols = [entity] if isinstance(entity, str) else list(entity)
    key_cols = [*entity_cols, session]
    declared = expected or {}
    summaries: list[pl.DataFrame] = []
    residuals: list[pl.DataFrame] = []
    per_column: list[dict[str, object]] = []
    for column in columns:
        produced = frame.filter(pl.col(column).is_not_null()).select(key_cols)
        missing = frame.filter(pl.col(column).is_null()).select(key_cols)
        per_column.append(
            {
                "column": column,
                "rows": frame.height,
                "n_null": missing.height,
                "coverage": (frame.height - missing.height) / frame.height
                if frame.height
                else None,
                "declared": column in declared,
            }
        )
        if missing.height == 0:
            continue
        parts = explain_missing(
            missing,
            produced,
            entity=entity_cols,
            session=session,
            expected=declared.get(column),
        )
        summaries.append(parts["summary"].with_columns(pl.lit(column).alias("column")))
        if parts["residual"].height:
            # explain_missing already carries an undeclared position here: its budget
            # defaults to zero, so every entity in it is over budget and its keys are
            # residual. Adding them again would double every unexplained null.
            residuals.append(parts["residual"].with_columns(pl.lit(column).alias("column")))

    summary = (
        pl.concat(summaries).select("column", pl.exclude("column"))
        if summaries
        else pl.DataFrame(schema={"column": pl.String, "where": pl.String})
    )
    residual = (
        pl.concat(residuals, how="diagonal_relaxed")
        if residuals
        else pl.DataFrame(schema={"column": pl.String})
    )
    return {
        "per_column": pl.DataFrame(per_column),
        "summary": summary,
        "residual": residual,
    }


def render_column_gaps(gaps: dict[str, pl.DataFrame], *, max_rows: int = 60) -> None:
    """Print ``explain_column_gaps`` the way a sign-off has to read it.

    Per-column coverage first, because a column at 100% needs no further reading and a
    column at 55% is the whole question. The positional breakdown second, for the
    columns that lost anything. The residual last as a single number, so the sign-off
    speaks to what no declaration accounts for rather than to a percentage.
    """
    per_column = gaps["per_column"]
    dense = per_column.filter(pl.col("n_null") == 0)
    print(f"{per_column.height} feature column(s), {dense.height} with a value at every key")
    with pl.Config(
        tbl_rows=max_rows, tbl_cols=5, tbl_width_chars=WIDE, tbl_hide_dataframe_shape=True
    ):
        print(per_column.filter(pl.col("n_null") > 0).sort("coverage"))

    summary = gaps["summary"]
    if summary.height:
        print("where each column's nulls sit, and what accounts for them:")
        with pl.Config(
            tbl_rows=max_rows,
            tbl_cols=10,
            fmt_str_lengths=44,
            tbl_width_chars=WIDE,
            tbl_hide_dataframe_shape=True,
        ):
            print(summary)

    residual = gaps["residual"]
    if residual.height:
        by_column = residual.group_by("column").len().sort("len", descending=True)
        print(f"residual to explain: {residual.height:,} value(s) no declaration accounts for")
        with pl.Config(tbl_rows=max_rows, tbl_hide_dataframe_shape=True):
            print(by_column)
    else:
        print("residual to explain: none - every null sits where a declaration puts it")


def render_cross_sectional_dispersion(report: dict, *, max_rows: int = 20) -> None:
    """Print a ``cross_sectional_dispersion`` report: the verdict, then the offenders.

    The share leads because it is the number a reader acts on, and the degenerate
    sessions follow because a share alone does not say whether they cluster at one end
    of the sample - three at the start is a warmup, three scattered is something else.
    """
    s = report["summary"]
    per_session = report["per_session"]

    if not s["n_rankable_sessions"]:
        print(f"{s['value']}: no session carries more than one entity - nothing to rank")
        return

    share = s["share_degenerate"]
    print(
        f"{s['value']}: {s['n_degenerate']:,} of {s['n_rankable_sessions']:,} rankable "
        f"session(s) rank nothing ({share:.2%}), at eps={s['eps']:g}"
    )
    if s["n_sessions"] != s["n_rankable_sessions"]:
        skipped = s["n_sessions"] - s["n_rankable_sessions"]
        print(f"  {skipped:,} session(s) carry a single entity and are not counted either way")
    print(
        f"  cross-sectional std across rankable sessions: min {s['min_std']:.3e}, "
        f"median {s['median_std']:.3e}"
    )

    offenders = per_session.filter(pl.col("degenerate"))
    if offenders.height:
        print(f"  the {min(offenders.height, max_rows):,} shown of {offenders.height:,}:")
        with pl.Config(tbl_rows=max_rows, tbl_width_chars=WIDE, tbl_hide_dataframe_shape=True):
            print(offenders.head(max_rows))
    else:
        print("  every rankable session orders its cross-section")

```

स्रोत के लाइसेंस के तहत श्रेय सहित पूरा पाठ दिखाया गया है। लाइसेंस: MIT

यह सारांश मूल स्रोत के आधार पर Stratmill के शोध एजेंट ने लिखा है; यह स्रोत की प्रति नहीं है।