跳至正文
返回文库全部文档

审计不同交易策略间的LEAN一致性

代码 《交易机器学习》

总结

本笔记从一项审计中单独分析 LEAN 引擎的结果。该审计将保留的真实策略运行结果与匹配的 ML4T Backtest配置进行比较。笔记列出四类受支持的工作负载:ETF配置、加密货币永续合约资金费率、以 USD 报价的 FX 配置,以及 US 股票面板。冻结的 CME 期货数据包因缺少创建相应原生 LEAN 订阅所需的带日期合约链和展期映射而被排除。比较报告成交、估值、时间戳一致性、权益及终值差异,以及负向对照的检出情况。

单独的计时部分仅测量引擎调用,并使用预热和进程隔离样本;数据加载、推理、准备和报告均不在测量区间内。笔记明确将计时结论限定于本次审计使用的固定版本、数据包和测量边界。合成压力测试结果也与真实策略的一致性分开处理。所提供文本显示受支持的行通过了状态检查,但没有包含底层审计数值,因此无法支持更广泛的性能结论。

核心观点

  • 一致性审计使用共同的冻结输入,将保留的 LEAN 结果与匹配的回测配置进行比较。
  • 所提供的数据包支持本次比较中的 ETF、加密货币永续合约、以 USD 报价的 FX 和 US 股票工作负载。
  • 由于数据包缺少带日期的合约和展期映射,CME 期货工作负载被省略。
  • 通过成交、估值和时间戳、权益差异、终值以及负向对照检出情况评估一致性。
  • 仅引擎计时不包括工作流中的其他步骤,结论只适用于指定版本和数据包。

标签

全文
# 15_lean_engine_parity.py


```py
# ---
# jupyter:
#   jupytext:
#     cell_metadata_filter: tags,-all
#     text_representation:
#       extension: .py
#       format_name: percent
#       format_version: '1.3'
#       jupytext_version: 1.19.3
#   kernelspec:
#     display_name: Python 3
#     language: python
#     name: python3
# ---

# %% [markdown]
# # LEAN Parity on Current Case-Study Strategies
#
# This notebook isolates the LEAN rows from the current real-strategy audit. It uses retained results
# generated by the native LEAN engine and the matching ML4T Backtest profiles. The shared inputs are
# frozen before engine execution.
#
# **Learning objectives**
#
# - Identify which selected asset classes are valid LEAN comparisons
# - Read LEAN parity across fills, valuations, and terminal value
# - Interpret LEAN engine-only timing on the measured strategies
# - Keep synthetic stress evidence separate from real-strategy equivalence
#
# **Book reference**: Chapter 16, Section 16.3

# %% [markdown]
# ## Setup

# %%
"""Current LEAN parity evidence."""

import json

import polars as pl
from IPython.display import Markdown, display

from utils.paths import get_chapter_dir

# %% tags=["parameters"]
# Production defaults - Papermill injects overrides after this cell
ROUND_SECONDS = 3

# %% tags=["results"]
AUDIT_PATH = get_chapter_dir(16) / "resources" / "framework_parity_audit.json"
audit = json.loads(AUDIT_PATH.read_text(encoding="utf-8"))
lean = audit["frameworks"]["lean"]
LEAN_NAME = f"{lean['display_name']} {lean['version']}"
CASE_NAMES = {
    "etfs": "ETF allocation",
    "cme_futures": "CME futures",
    "crypto_perps_funding": "Crypto perpetual funding",
    "fx_pairs": "FX allocation (USD-quoted pairs)",
    "us_equities_panel": "US equity panel",
}

# %% tags=["results"]
display(Markdown(f"**Pinned engine:** {LEAN_NAME} with ML4T profile `{lean['profile']}`"))

# %% [markdown]
# ## 1. Supported real strategies
#
# LEAN is required for the ETF, crypto-perpetual, USD-quoted foreign-exchange, and US equity-panel
# workloads. The CME row is unsupported for this particular frozen bundle: it contains continuous
# root prices but lacks the dated contract chain and roll map needed to construct a native LEAN
# futures subscription.

# %% tags=["results"]
lean_results = (
    pl.DataFrame(audit["real_strategy_records"])
    .filter(pl.col("framework") == "lean")
    .with_columns(pl.col("case_study").replace_strict(CASE_NAMES).alias("strategy"))
    .select(
        "strategy",
        "status",
        "fills",
        "valuations",
        "valuation_timestamps_match",
        "equity_gap",
        "equity_raw_gap",
        "terminal_gap",
        "terminal_raw_gap",
        "negative_control_detected",
    )
)

assert lean_results.height == 4
assert lean_results["status"].to_list() == ["pass"] * 4
assert lean_results["valuation_timestamps_match"].all()
assert lean_results["negative_control_detected"].all()

display(lean_results)

# %% tags=["results"]
lean_unsupported = (
    pl.DataFrame(audit["unsupported_records"])
    .filter(pl.col("framework") == "lean")
    .with_columns(pl.col("case_study").replace_strict(CASE_NAMES).alias("strategy"))
    .select("strategy", "reason")
)
display(lean_unsupported)

# %% [markdown]
# The table reports the complete fill and valuation counts for each supported workload. LEAN uses
# native equity, crypto-future, and foreign-exchange securities. The audit does not convert the
# continuous CME roots into a different instrument merely to add a LEAN row.

# %% [markdown]
# ## 2. Engine-only timing
#
# The timer starts immediately before the engine call and stops when it returns. One warmup and ten
# process-isolated samples are used. Input loading, model inference, target construction, adapter
# preparation, result extraction, and reporting are outside the timed region.

# %% tags=["results"]
lean_timing = (
    pl.DataFrame(audit["performance_records"])
    .filter(pl.col("framework") == "lean")
    .with_columns(
        pl.col("case_study").replace_strict(CASE_NAMES).alias("strategy"),
        pl.col("framework_median_seconds").round(ROUND_SECONDS).alias("lean_seconds"),
        pl.col("ml4t_median_seconds").round(ROUND_SECONDS).alias("ml4t_seconds"),
        pl.col("framework_to_ml4t_ratio").round(2).alias("lean_div_ml4t"),
    )
    .select("strategy", "lean_seconds", "ml4t_seconds", "lean_div_ml4t")
)
display(lean_timing)

# %% [markdown]
# The timing result applies to these pinned versions, bundles, and engine boundaries. It is not a
# general LEAN performance claim.

# %% [markdown]
# ## 3. Synthetic stress remains diagnostic

# %% tags=["results"]
lean_stress = (
    pl.DataFrame(audit["synthetic_stress"]["records"])
    .filter(pl.col("framework") == "lean")
    .select("intents", "fills", "trades", "terminal_value", "status")
)
display(lean_stress)

# %% [markdown]
# The retained stress row tests scale and the calibrated LEAN profile on generated inputs. The
# supported rows above provide the real-data evidence.

```

在遵守原作品许可的前提下,附作者信息全文展示。 许可协议: MIT

此摘要由 Stratmill 研究智能体根据原文撰写,并非原文副本。