诊断交易研究流程中的数据覆盖与特征质量
代码 《交易机器学习》
总结
本工具介绍用于诊断量化研究流程所生成数据产物的方法。它会统计数值列中的空值、非有限值、零值、不同值数量、分位数和尾部行为,并标记超出可配置提示阈值的列。它还会将已生成的键与明确声明的预期范围进行比较,报告缺失和意外的键,避免仅因现有行完整,就误以为规模不足的输出状况良好。
报告可以按实体和交易时段对缺失观测分类,利用已声明的预期值区分设计允许的缺口与原因不明的残留缺失。另一个横截面离散度检查会找出预测在不同实体间没有变化、因而无法排序的时段。这些检查会提供证据供笔记作者解读;它们不会判定某个特定空值或零值是否有问题,提示性标记也不会阻止流程运行。本文将数据质量视为需要结合计数和分布作出判断的问题,同时将强制拦截规则留给单独的验证关卡。
核心观点
- 将已生成的键与声明的预期数据集合进行比较,以发现缺失行和意外行。
- 按数值列统计空值、非有限值、零值、不同值数量、分位数和尾部比率。
- 将阈值标记视为需要说明的提示,因为可接受的值取决于特征。
- 按实体和交易时段对缺失键分类,以区分预期缺口与原因不明的缺口。
- 检查各时段内预测的离散程度,因为无法对实体排序的预测可能无法使用。
标签
全文
# artifact_quality.py
```py
"""Descriptive statistics and coverage for what a production stage just wrote.
Three questions, asked at the end of the notebook that produces an artifact, about
the frame it is about to write:
1. **Is it the size it should be?** Not "are the rows it has non-null" - a stage that
emits a thousand rows where a million were expected reports no missing values and
is wrong. ``coverage_against`` compares the produced keys to the universe the
caller declares.
2. **What is in each column?** ``profile_columns`` gives one row per column: nulls,
zeros, non-finites, distinct count, quantiles and a tail ratio.
3. **What is unusual?** ``flag_columns`` marks the profile rows that cross a
threshold, with the reason attached.
Nothing here raises. A zero share of 0.5 is a defect in a return and correct in a
binary direction label; a 45% null share is a defect in a price and expected in a
term-structure feature that needs two expiries. Only the notebook author knows
which, so this surfaces the number and the author signs off on it in prose. The
gates that *do* refuse live in ``coverage.py`` and ``notebook_contracts.py``.
Polars in, polars out. Pandas only if a caller is at a visualization boundary.
"""
from __future__ import annotations
from collections.abc import Sequence
from pathlib import Path
import polars as pl
__all__ = [
"coverage_against",
"cross_sectional_dispersion",
"explain_missing",
"flag_columns",
"label_universe",
"profile_columns",
"quality_report",
"render_cross_sectional_dispersion",
"render_quality_report",
]
#: Quantiles reported for every numeric column. The outer pair is what makes a
#: heavy tail visible; the inner three are the shape.
DEFAULT_QUANTILES: tuple[float, ...] = (0.001, 0.01, 0.25, 0.5, 0.75, 0.99, 0.999)
#: Advisory thresholds. Deliberately loose - these decide what a reader is asked to
#: look at, not what is acceptable. Tighten per notebook by passing your own.
DEFAULT_RULES: dict[str, float] = {
"null_share": 0.20,
"zero_share": 0.60,
"tail_ratio": 25.0,
}
def _numeric_columns(frame: pl.DataFrame, exclude: Sequence[str]) -> list[str]:
excluded = set(exclude)
return [c for c, dtype in frame.schema.items() if dtype.is_numeric() and c not in excluded]
def profile_columns(
frame: pl.DataFrame,
*,
key_columns: Sequence[str] = (),
quantiles: Sequence[float] = DEFAULT_QUANTILES,
) -> pl.DataFrame:
"""One row per numeric column: how much is there, and what shape is it.
``key_columns`` are excluded - a symbol id or a fold index has no distribution
worth reporting. Non-finite values are counted separately from nulls: a null is
an absent observation and an infinity is a computed one that overflowed, and the
two have different causes even though most downstream code treats them alike.
``tail_ratio`` is ``|q99.9| / |q75|``, which is a scale-free way to see a column
whose extreme values are orders of magnitude past its body. It is null where the
denominator is zero rather than infinite, so the column stays sortable.
"""
columns = _numeric_columns(frame, key_columns)
if not columns:
return pl.DataFrame(schema={"column": pl.String})
rows: list[dict[str, object]] = []
height = frame.height
for name in columns:
series = frame[name]
finite = series.drop_nulls()
n_nonfinite = int(finite.is_infinite().sum() or 0) + int(finite.is_nan().sum() or 0)
clean = finite.filter(finite.is_finite())
row: dict[str, object] = {
"column": name,
"rows": height,
"n_null": int(series.null_count()),
"null_share": series.null_count() / height if height else None,
"n_nonfinite": n_nonfinite,
"n_zero": int((clean == 0).sum() or 0),
"zero_share": (float((clean == 0).sum() or 0) / clean.len()) if clean.len() else None,
"n_distinct": int(clean.n_unique()),
"constant": clean.n_unique() <= 1 if clean.len() else None,
"mean": float(clean.mean()) if clean.len() else None,
"std": float(clean.std()) if clean.len() > 1 else None,
}
for q in quantiles:
row[f"q{q:g}"] = float(clean.quantile(q)) if clean.len() else None
body = row.get("q0.75")
tail = row.get("q0.999")
row["tail_ratio"] = (
abs(tail) / abs(body) if body not in (None, 0) and tail is not None else None
)
rows.append(row)
return pl.DataFrame(rows)
def coverage_against(
produced: pl.DataFrame,
expected: pl.DataFrame,
*,
keys: Sequence[str],
) -> dict[str, pl.DataFrame]:
"""How much of the declared universe the stage actually produced.
``expected`` is the universe the caller declares this stage owes a row for - the
label artifact restricted to a fold's validation window, a trading calendar
crossed with the symbol universe, the feature panel a model is scored on. Both
frames are reduced to their distinct ``keys`` before comparison, so passing a
frame with extra columns is fine.
Returns ``summary`` with the counts, plus ``missing`` and ``unexpected`` carrying
the key tuples themselves. The tuples are what make the number readable: a
handful of symbols absent entirely and every symbol missing its first weeks give
the same percentage and are a universe restriction and a warm-up respectively.
Group ``missing`` by the entity column to tell them apart.
"""
key_list = list(keys)
mismatched = [
(k, expected.schema[k], produced.schema[k])
for k in key_list
if expected.schema[k] != produced.schema[k]
]
if mismatched:
detail = "; ".join(
f"{k}: expected {want} but produced {got}" for k, want, got in mismatched
)
raise ValueError(
"key dtypes differ between the produced frame and the declared universe, so "
f"every row would read as missing - {detail}. Cast explicitly in the notebook: "
"a silent cast here is how a Date label artifact and a Datetime prediction set "
"come to report full coverage of nothing."
)
got = produced.select(key_list).unique()
want = expected.select(key_list).unique()
missing = want.join(got, on=key_list, how="anti")
unexpected = got.join(want, on=key_list, how="anti")
summary = pl.DataFrame(
[
{
"expected": want.height,
"produced": got.height,
"missing": missing.height,
"unexpected": unexpected.height,
"coverage": (want.height - missing.height) / want.height if want.height else None,
}
]
)
return {"summary": summary, "missing": missing, "unexpected": unexpected}
def flag_columns(
profile: pl.DataFrame,
*,
rules: dict[str, float] | None = None,
) -> pl.DataFrame:
"""The profile rows a reader should look at, with the reason attached.
A flag is a request for a sentence of explanation, not a failure. Every rule is a
ceiling on a column of ``profile``; a column crossing none of them is absent from
the result. Non-finite values and constants are always flagged, because neither
has a legitimate reading that does not need saying out loud.
"""
active = DEFAULT_RULES | (rules or {})
reasons: list[pl.Expr] = [
pl.when(pl.col("n_nonfinite") > 0)
.then(pl.format("{} non-finite values", pl.col("n_nonfinite")))
.otherwise(None),
pl.when(pl.col("constant")).then(pl.lit("constant over every row")).otherwise(None),
]
for column, ceiling in active.items():
if column not in profile.columns:
continue
reasons.append(
pl.when(pl.col(column) > ceiling)
.then(pl.format(f"{column} {{}} above {ceiling:g}", pl.col(column).round(4)))
.otherwise(None)
)
flagged = profile.with_columns(pl.concat_list(reasons).list.drop_nulls().alias("flags")).filter(
pl.col("flags").list.len() > 0
)
return flagged.select(
"column", pl.col("flags").list.join("; ").alias("why"), pl.exclude("column", "flags")
)
def quality_report(
frame: pl.DataFrame,
*,
name: str,
key_columns: Sequence[str] = (),
expected: pl.DataFrame | None = None,
keys: Sequence[str] | None = None,
rules: dict[str, float] | None = None,
entity: str | Sequence[str] | None = None,
session: str | None = None,
expected_missing: dict[str, MissingExpectation] | None = None,
) -> dict[str, pl.DataFrame]:
"""Profile, flags, coverage and (when the shortfall is declared) its explanation.
``entity`` and ``session`` name the two axes of the key; ``entity`` takes a sequence where
the entity is compound, as it is for an option contract keyed by underlying and instrument or
a futures curve keyed by product and position. Given them, the missing
keys are classified by ``explain_missing`` and ``expected_missing`` declares which
positions a mechanism accounts for, so the report distinguishes a shortfall that is
the design from one that is a defect. Without them the report still carries the
coverage counts, and the notebook is left asserting a percentage is acceptable
without saying what produced it.
Returns the frames rather than printing them, so the notebook decides how to
display each and writes its own sign-off around them.
"""
profile = profile_columns(frame, key_columns=key_columns)
report: dict[str, pl.DataFrame] = {
"profile": profile,
"flags": flag_columns(profile, rules=rules),
}
if expected is not None:
coverage = coverage_against(frame, expected, keys=keys or list(key_columns))
report.update({f"coverage_{part}": value for part, value in coverage.items()})
if entity is not None and session is not None and coverage["missing"].height:
explanation = explain_missing(
coverage["missing"],
frame,
entity=entity,
session=session,
expected=expected_missing,
)
report.update({f"missing_{part}": value for part, value in explanation.items()})
report["name"] = name
return report
def cross_sectional_dispersion(
frame: pl.DataFrame,
*,
value: str,
entity: str = "symbol",
session: str = "timestamp",
eps: float | None = None,
) -> dict:
"""Does this forecast order anything at each decision time?
``profile_columns`` cannot answer that. Its ``constant`` flag is
``n_unique() <= 1`` over the whole column, so a forecast that is constant
within every timestamp and merely drifts across them profiles as varied. That
is the exact shape that ranks nothing: a top-k book built from it fills once and
never replaces, trades between zero and two times, pays almost no cost, and can
top a cost-dominated leaderboard on the strength of not trading.
So the dispersion has to be measured **within** ``session``, across ``entity``,
and the test is a standard deviation at or below ``eps``. That covers both shapes
with one clause: an exactly-constant cross-section has a standard deviation of
zero, and a regularized fit that leaves float noise instead of a constant carries
1.95e-17 to 2.02e-17 - measured on the 24 denormal rows across all nine
registries, against a smallest genuine value of 1.69e-03.
An ``n_distinct == 1`` clause was written beside it first and removed: mutation
showed it dead, because every cross-section it catches the standard deviation
already catches. ``n_distinct`` is still reported, because a reader looking at a
flagged session wants to know which of the two shapes it is.
``eps`` defaults to the same constant the registry's degeneracy rule uses, imported
rather than restated so the two cannot drift apart. The import is deferred because
this module is loaded by stage-02 notebooks and deliberately depends on nothing but
polars at module level.
Returns ``per_session``, one row per decision time, and ``summary``. **It does not
raise**, in keeping with the rest of this module: a prediction set that ranks
nothing at 3 of 500 timestamps may be a defect or may be three holidays, and only
the notebook author knows which. The gate that refuses lives in
``notebook_contracts.degenerate_prediction_sql``.
"""
if eps is None:
from case_studies.utils.notebook_contracts import _DEGENERATE_IC_EPS
eps = _DEGENERATE_IC_EPS
missing = [c for c in (value, entity, session) if c not in frame.schema]
if missing:
raise KeyError(f"frame has no column(s) {missing}; it carries {sorted(frame.schema)}")
per_session = (
frame.select(session, entity, value)
.drop_nulls(value)
.group_by(session)
.agg(
n_entities=pl.len(),
n_distinct=pl.col(value).n_unique(),
std=pl.col(value).std(),
spread=pl.col(value).max() - pl.col(value).min(),
)
.with_columns(
# A session carrying one entity is not a failure to rank, it is a session
# with nothing to rank. Marking it degenerate would report a universe
# shrinking to one name as a model defect.
degenerate=(pl.col("n_entities") > 1) & (pl.col("std").fill_null(0.0) <= eps),
)
.sort(session)
)
rankable = per_session.filter(pl.col("n_entities") > 1)
n_degenerate = int(per_session.get_column("degenerate").sum())
summary = {
"value": value,
"n_sessions": per_session.height,
"n_rankable_sessions": rankable.height,
"n_degenerate": n_degenerate,
"share_degenerate": (n_degenerate / rankable.height) if rankable.height else None,
"min_std": float(rankable.get_column("std").min()) if rankable.height else None,
"median_std": float(rankable.get_column("std").median()) if rankable.height else None,
"eps": eps,
}
return {"per_session": per_session, "summary": summary}
#: Table width for every frame these renderers print. Polars defaults to 80 characters,
#: which wraps a 19-column profile into an unreadable stack of fragments - the table is
#: there to be read at a glance and at 80 it cannot be. Notebook output has the room.
WIDE = 240
def render_quality_report(report: dict, *, max_rows: int = 80) -> None:
"""Print a ``quality_report`` as a notebook reader should read it: shortfall first.
Coverage leads because it is the question the other tables cannot answer - a frame
whose every column profiles cleanly is still wrong if it is a tenth of the universe.
Where the shortfall has been classified, the explanation prints next to it and the
residual is stated as its own number, because "76% coverage" and "76% coverage, all
of it a declared 252-session estimation burn-in, nothing unexplained" are the same
percentage and different findings. Flags come next as the shortlist a sign-off has
to speak to, and the full profile last.
"""
print(f"=== {report['name']} ===")
summary = report.get("coverage_summary")
if summary is not None:
row = summary.row(0, named=True)
share = "n/a" if row["coverage"] is None else f"{row['coverage']:.2%}"
print(
f"coverage {share}: {row['produced']:,} of {row['expected']:,} declared keys, "
f"{row['missing']:,} missing, {row['unexpected']:,} unexpected"
)
explained = report.get("missing_summary")
if explained is not None and explained.height:
print("where the missing keys sit, and what accounts for them:")
with pl.Config(
tbl_rows=8,
tbl_cols=9,
fmt_str_lengths=40,
tbl_width_chars=WIDE,
tbl_hide_dataframe_shape=True,
):
print(explained)
undeclared = int(explained.filter(pl.col("why") == "not declared")["n_keys"].sum() or 0)
excess = int(explained["excess_keys"].sum() or 0)
over = int(explained["entities_over"].sum() or 0)
print(
f"residual to explain: {undeclared:,} key(s) in positions no mechanism was "
f"declared for, and {excess:,} spent past a declared budget by {over} entities"
)
else:
missing = report.get("coverage_missing")
if missing is not None and missing.height:
# Every key dimension, not the first one. A shortfall concentrated in a
# few entities and one falling on whole timestamps are different defects
# - a universe restriction against a feed outage - and reporting only the
# column written first shows one and hides the other.
for key in missing.columns:
per_key = missing.group_by(key).len().sort("len", descending=True)
print(f" {per_key.height} distinct {key} values are short; worst:")
with pl.Config(tbl_rows=8, tbl_hide_dataframe_shape=True):
print(per_key.head(8))
flags = report["flags"]
if flags.height == 0:
print("no column crossed a threshold")
else:
print(f"{flags.height} column(s) to speak to:")
with pl.Config(
tbl_rows=max_rows,
tbl_cols=4,
fmt_str_lengths=110,
tbl_width_chars=WIDE,
tbl_hide_dataframe_shape=True,
):
print(flags.select("column", "why", "rows", "n_null"))
with pl.Config(
tbl_rows=max_rows, tbl_cols=20, tbl_width_chars=WIDE, tbl_hide_dataframe_shape=True
):
print(report["profile"])
#: What a declared expectation looks like: how many sessions per entity the mechanism
#: is allowed to cost, and the mechanism's name. ``None`` for the budget means the
#: mechanism has no per-entity ceiling worth stating - an entity absent from the
#: universe for its whole life, say - and the position is then explained but unbounded.
MissingExpectation = tuple[int | None, str]
def explain_missing(
missing: pl.DataFrame,
produced: pl.DataFrame,
*,
entity: str | Sequence[str],
session: str,
expected: dict[str, MissingExpectation] | None = None,
) -> dict[str, pl.DataFrame]:
"""Split the missing keys into the ones a mechanism accounts for and the residual.
A coverage percentage on its own decides nothing. A stage whose features need a
year of history before they resolve, against a five-year sample, owes no row for
the first fifth of it and is complete at 80%; a stage losing the same fraction
scattered through the middle of every symbol's quoted range is broken at the same
number. The two are told apart by *where* the key sits relative to what the entity
did produce, so that is what this classifies:
``leading``
Before the entity's first produced row. A lookback, a burn-in, an estimation
window, or a listing date later than the sample start.
``trailing``
After its last produced row. A forward label's horizon running off the end of
the sample, or a delisting.
``interior``
Inside the entity's own produced span, with rows on both sides. No window
length explains one of these.
``absent``
The entity produced no row at all. A universe restriction, and the case a
percentage misleads on hardest: it reads as thin coverage of everything rather
than no coverage of some.
``expected`` declares the mechanism per position and the sessions per entity it may
cost: ``{"trailing": (10, "10-session forward label horizon")}``. A position with no
entry is unexplained in full. An entity spending more than its budget contributes
its keys to ``residual`` and its excess to the summary, and that is what a sign-off
speaks to - the rest is design, which is what declaring it says.
"""
entity_cols = [entity] if isinstance(entity, str) else list(entity)
span = produced.group_by(entity_cols).agg(
pl.col(session).min().alias("_first"),
pl.col(session).max().alias("_last"),
)
classified = (
missing.join(span, on=entity_cols, how="left")
.with_columns(
pl.when(pl.col("_first").is_null())
.then(pl.lit("absent"))
.when(pl.col(session) < pl.col("_first"))
.then(pl.lit("leading"))
.when(pl.col(session) > pl.col("_last"))
.then(pl.lit("trailing"))
.otherwise(pl.lit("interior"))
.alias("where")
)
.drop("_first", "_last")
)
declared = expected or {}
per_entity = classified.group_by([*entity_cols, "where"]).len().rename({"len": "sessions"})
rows: list[dict[str, object]] = []
residual_frames: list[pl.DataFrame] = []
for position in ("absent", "leading", "interior", "trailing"):
part = per_entity.filter(pl.col("where") == position)
if part.height == 0:
continue
budget, why = declared.get(position, (0, "not declared"))
over = part.filter(pl.col("sessions") > budget) if budget is not None else part.head(0)
excess = int((over["sessions"] - budget).sum()) if budget is not None and over.height else 0
rows.append(
{
"where": position,
"n_keys": int(part["sessions"].sum()),
"n_entities": part.height,
"sessions_p50": float(part["sessions"].median()),
"sessions_max": int(part["sessions"].max()),
"budget": budget,
"entities_over": over.height,
"excess_keys": excess,
"why": why,
}
)
if over.height:
residual_frames.append(
classified.filter(pl.col("where") == position).join(
over.select(entity_cols), on=entity_cols, how="semi"
)
)
residual = pl.concat(residual_frames) if residual_frames else classified.head(0)
return {
"classified": classified,
"per_entity": per_entity,
"summary": pl.DataFrame(rows) if rows else pl.DataFrame(schema={"where": pl.String}),
"residual": residual,
}
def label_universe(
case_dir: Path | str,
*,
labels: Sequence[str] | None = None,
keys: Sequence[str] = ("symbol", "timestamp"),
) -> pl.DataFrame:
"""The distinct keys the case study's labels declare, as an ``expected`` frame.
This is what a feature stage owes a row for: a key carrying a label and no feature
row is one no model can be asked to score. Reading it here rather than trusting the
stage's own inputs is the whole point - a panel can be internally consistent, carry
no nulls, and still be missing a fifth of the universe the labels declare.
``labels`` names the artifacts to union; the default is every parquet in
``labels/``. A row is counted whether or not its label value is null, because a
stage does not know which labels a later stage will filter on.
No dtype coercion happens here or in ``coverage_against``. A label artifact keying
``timestamp`` as ``Date`` against a feature panel keying it as ``Datetime`` is a real
disagreement, and casting it away is how a frame reports full coverage of nothing.
"""
directory = Path(case_dir) / "labels"
paths = (
[directory / f"{name}.parquet" for name in labels]
if labels is not None
else sorted(directory.glob("*.parquet"))
)
frames = [
pl.scan_parquet(path).select(list(keys)).unique().collect()
for path in paths
if path.is_file()
]
if not frames:
raise FileNotFoundError(f"no label artifact under {directory}; the universe is undefined")
return pl.concat(frames).unique()
def explain_column_gaps(
frame: pl.DataFrame,
*,
columns: Sequence[str],
entity: str | Sequence[str],
session: str,
expected: dict[str, dict[str, MissingExpectation]] | None = None,
) -> dict[str, pl.DataFrame]:
"""Where each column's nulls sit relative to the rows that column did produce.
``coverage_against`` asks whether the stage wrote a row for a key. A stage that
writes one row per key and leaves the estimation window null answers that question
perfectly and still ships an artifact whose feature is absent for a tenth of the
panel, because the key is present and only the value is not. This asks the same
positional question of each column on its own: the keys where the column is null
are its missing set, the keys where it carries a value are what it produced, and
``explain_missing`` classifies the difference. A fitted feature's burn-in and a hole
in the middle of a fitted series then read as the two different things they are.
Declaring per column is what makes the answer usable, because the families in one
artifact do not lose the same rows: a rolling model owes its burn-in, a
market-wide series owes everything before the series existed, and a spread against
a quoted surface owes every key the surface does not quote. ``expected`` therefore
maps a column name to the per-position declaration ``explain_missing`` takes. A
column with no entry is still measured; nothing is declared for it, so all of its
nulls land in the residual.
"""
entity_cols = [entity] if isinstance(entity, str) else list(entity)
key_cols = [*entity_cols, session]
declared = expected or {}
summaries: list[pl.DataFrame] = []
residuals: list[pl.DataFrame] = []
per_column: list[dict[str, object]] = []
for column in columns:
produced = frame.filter(pl.col(column).is_not_null()).select(key_cols)
missing = frame.filter(pl.col(column).is_null()).select(key_cols)
per_column.append(
{
"column": column,
"rows": frame.height,
"n_null": missing.height,
"coverage": (frame.height - missing.height) / frame.height
if frame.height
else None,
"declared": column in declared,
}
)
if missing.height == 0:
continue
parts = explain_missing(
missing,
produced,
entity=entity_cols,
session=session,
expected=declared.get(column),
)
summaries.append(parts["summary"].with_columns(pl.lit(column).alias("column")))
if parts["residual"].height:
# explain_missing already carries an undeclared position here: its budget
# defaults to zero, so every entity in it is over budget and its keys are
# residual. Adding them again would double every unexplained null.
residuals.append(parts["residual"].with_columns(pl.lit(column).alias("column")))
summary = (
pl.concat(summaries).select("column", pl.exclude("column"))
if summaries
else pl.DataFrame(schema={"column": pl.String, "where": pl.String})
)
residual = (
pl.concat(residuals, how="diagonal_relaxed")
if residuals
else pl.DataFrame(schema={"column": pl.String})
)
return {
"per_column": pl.DataFrame(per_column),
"summary": summary,
"residual": residual,
}
def render_column_gaps(gaps: dict[str, pl.DataFrame], *, max_rows: int = 60) -> None:
"""Print ``explain_column_gaps`` the way a sign-off has to read it.
Per-column coverage first, because a column at 100% needs no further reading and a
column at 55% is the whole question. The positional breakdown second, for the
columns that lost anything. The residual last as a single number, so the sign-off
speaks to what no declaration accounts for rather than to a percentage.
"""
per_column = gaps["per_column"]
dense = per_column.filter(pl.col("n_null") == 0)
print(f"{per_column.height} feature column(s), {dense.height} with a value at every key")
with pl.Config(
tbl_rows=max_rows, tbl_cols=5, tbl_width_chars=WIDE, tbl_hide_dataframe_shape=True
):
print(per_column.filter(pl.col("n_null") > 0).sort("coverage"))
summary = gaps["summary"]
if summary.height:
print("where each column's nulls sit, and what accounts for them:")
with pl.Config(
tbl_rows=max_rows,
tbl_cols=10,
fmt_str_lengths=44,
tbl_width_chars=WIDE,
tbl_hide_dataframe_shape=True,
):
print(summary)
residual = gaps["residual"]
if residual.height:
by_column = residual.group_by("column").len().sort("len", descending=True)
print(f"residual to explain: {residual.height:,} value(s) no declaration accounts for")
with pl.Config(tbl_rows=max_rows, tbl_hide_dataframe_shape=True):
print(by_column)
else:
print("residual to explain: none - every null sits where a declaration puts it")
def render_cross_sectional_dispersion(report: dict, *, max_rows: int = 20) -> None:
"""Print a ``cross_sectional_dispersion`` report: the verdict, then the offenders.
The share leads because it is the number a reader acts on, and the degenerate
sessions follow because a share alone does not say whether they cluster at one end
of the sample - three at the start is a warmup, three scattered is something else.
"""
s = report["summary"]
per_session = report["per_session"]
if not s["n_rankable_sessions"]:
print(f"{s['value']}: no session carries more than one entity - nothing to rank")
return
share = s["share_degenerate"]
print(
f"{s['value']}: {s['n_degenerate']:,} of {s['n_rankable_sessions']:,} rankable "
f"session(s) rank nothing ({share:.2%}), at eps={s['eps']:g}"
)
if s["n_sessions"] != s["n_rankable_sessions"]:
skipped = s["n_sessions"] - s["n_rankable_sessions"]
print(f" {skipped:,} session(s) carry a single entity and are not counted either way")
print(
f" cross-sectional std across rankable sessions: min {s['min_std']:.3e}, "
f"median {s['median_std']:.3e}"
)
offenders = per_session.filter(pl.col("degenerate"))
if offenders.height:
print(f" the {min(offenders.height, max_rows):,} shown of {offenders.height:,}:")
with pl.Config(tbl_rows=max_rows, tbl_width_chars=WIDE, tbl_hide_dataframe_shape=True):
print(offenders.head(max_rows))
else:
print(" every rankable session orders its cross-section")
```在遵守原作品许可的前提下,附作者信息全文展示。 许可协议: MIT
此摘要由 Stratmill 研究智能体根据原文撰写,并非原文副本。