跳至正文
返回文库全部文档

构建与评估交易策略的研究框架

文章 《交易机器学习》

总结

本章将策略研究描述为对可执行决策流程的设计和评估,涵盖从最初的经济构想到仓位规模、约束、成本和模拟实盘测试。章节建议先对策略系列和潜在优势来源分类,再构建模型;随后定义带版本的设置,以确保比较有意义。章节还将模型诊断、信号诊断和投资组合结果区分开来,以降低所有研究选择都针对回测表现进行优化的风险。

本章强调按时间顺序进行时间序列评估,包括滚动划分、针对重叠标签和特征的缓冲区、封存的留出集、嵌套评估和组合变体。章节建议先采用范围有限的基准,检查时点、覆盖范围和交易强度,再扩大搜索;同时要求记录试验,以便审查实验规模。笔记通过跨资产动量、验证方法和案例研究比较展示这一流程。材料提供的是研究流程,并非任何特定策略能够盈利的证据;结果仍取决于数据质量、实现假设、成本和市场变化。

核心观点

  • 策略必须明确可执行的决策流程,而不只是预测信号。
  • 按策略系列、合理的优势来源、可行性约束和失效模式对构想分类。
  • 比较参数选择时固定设置机制,并将参数选择与结构变化区分开来。
  • 采用按时间顺序的验证并控制重叠,同时为最终表现估计保留留出集。
  • 扩大搜索前先检查范围有限的基准,并记录每次试验。

标签

全文
# Chapter 6: Strategy Research Framework


# Chapter 6: Strategy Research Framework

The chapter establishes the chapter's core claim: a trading strategy is not just a signal or model, but an executable decision process that has to be defined at decision time and evaluated as if it were live. It distinguishes the live trading loop from the research loop and shows why disciplined iteration matters if historical testing is supposed to say anything about future behavior. The case studies make the workflow concrete across asset classes, cadences, and market structures, so readers see early that the same research logic must survive very different implementation environments.

## Learning Objectives

* Place a strategy idea on the strategy map by linking it to a strategy family, a plausible source of edge, and the dominant feasibility constraints and failure modes.
* Define a versioned trading setup in decision-time terms: what is tradable, when decisions are made, what information is admissible, how scores become positions, and which constraints and costs are treated as material.
* Define "better" economically and keep model diagnostics, signal diagnostics, and strategy outcomes in distinct roles during research and evaluation.
* Design a time-series evaluation protocol that preserves chronology, prevents overlap leakage, and separates model selection from final performance estimation.
* Establish a narrow baseline checkpoint with timing, coverage, and trading-intensity sanity checks before expanding the search space.
* Keep search auditable, reproducible, and countable using a simple trial taxonomy and automatic run logging.

## Sections

### 6.1 From Idea to Evidence with the ML4T Workflow

This section establishes the chapter's core claim: a trading strategy is not just a signal or model, but an executable decision process that has to be defined at decision time and evaluated as if it were live. It distinguishes the live trading loop from the research loop and shows why disciplined iteration matters if historical testing is supposed to say anything about future behavior. The case studies make the workflow concrete across asset classes, cadences, and market structures, so readers see early that the same research logic must survive very different implementation environments.

### 6.2 Mapping Strategies and Sources of Edge

This section gives readers a way to classify ideas before they start building models. Strategy families act as feasibility filters, while sources of edge act as durability filters: together they force the reader to ask not only whether a pattern can be backtested, but whether it is economically plausible, implementable, and likely to persist after costs, constraints, and competition. That makes this section important because it shifts strategy design away from loose narratives and toward testable economic hypotheses with explicit failure modes.

### 6.3 Defining the Trading Setup

Here the chapter turns the strategy map into a versioned trading setup. The key contribution is that comparability requires fixed invariants: tradability rules, decision schedule, score-to-trade mapping, constraints, and material cost components. The distinction between parameter tuning and mechanics changes is especially valuable because it gives readers a principled boundary for when they are still refining one strategy versus when they have quietly changed the structure of the strategy.

### 6.4 Setting Objectives and Evaluation Metrics

This section clarifies what "better" means in strategy research. Its main contribution is separating model diagnostics, signal diagnostics, and strategy outcomes so readers do not use one metric to answer incompatible questions. That separation matters because it reduces the temptation to optimize every micro-decision directly on simulated portfolio outcomes, which is one of the easiest ways to overfit a backtest.

### 6.5 Evaluation Protocol for Time Series

This is the chapter's methodological center. It explains why standard iid validation fails for financial time series, then introduces walk-forward evaluation, label and feature buffers, sealed holdouts, nested walk-forward, and combinatorial variants. Readers should care because this is the section that turns "out-of-sample" from a slogan into an actual protocol with admissibility rules, chronology, and governance around model selection versus final performance estimation.

### 6.6 Establishing a Baseline Checkpoint

This section argues that a narrow baseline is not a weak start but a governance tool. By insisting on timing, coverage, and trading-intensity sanity checks before large searches, it teaches readers how to rule out brittle setups early and earn the right to broaden the feature set or model class later. That is editorially strong because it frames baseline design as a way to avoid wasting effort on invalid or economically implausible research lines.

### 6.7 Search Accounting and Run Logging

This section makes experimentation auditable. Its value is not just reproducibility in the software-engineering sense, but countable search in the statistical sense: readers need to know what was tried, what was selected, and what was reserved for confirmation if they want performance claims to remain credible after iteration. The trial taxonomy is especially useful because it gives the book a concrete language for strategy, trial family, trial, and run.

## Notebooks

| Notebook                                                       | What it teaches                                                                                                          | Section |
|----------------------------------------------------------------|--------------------------------------------------------------------------------------------------------------------------|---------|
| [`01_where_ideas_come_from`](01_where_ideas_come_from.ipynb)   | From a story to a measured footprint: name a strategy's source of edge (the SLOW/WRONG/RISK taxonomy), then test whether cross-asset momentum leaves a footprint via a quintile conditional-return sort, an era stress test, and a turnover check. | §6.1–§6.2 |
| [`02_cv_foundations`](02_cv_foundations.ipynb)                 | Walk-forward CV from first principles: decision-time admissibility, label buffer (purging), feature buffer (embargo), calendar-aware splits, nested walk-forward, and CPCV. | §6.5    |
| [`03_case_study_overview`](03_case_study_overview.ipynb)       | Cross-strategy summary of the nine case studies — asset classes, universes, cost classes, evaluation protocols, and prediction-coverage timeline (Figure 6.5). | §6.3    |

## Running the Notebooks

```bash
# From the repository root
uv run python 06_strategy_definition/<notebook>.py

# Test mode (reduced data via Papermill)
uv run pytest tests/test_chapter_notebooks.py -v -k "06_strategy_definition"
```

## References

- **Brian Hurst et al.** A Century of Evidence on Trend-Following Investing.
- **Christoph Bergmeir et al.** (2018). [A note on the validity of cross-validation for evaluating autoregressive time series prediction](https://doi.org/10.1016/j.csda.2017.11.003). *Computational Statistics & Data Analysis*.
- **Clifford S. Asness et al.** (2013). [Value and Momentum Everywhere](https://www.jstor.org/stable/42002613). *The Journal of Finance*.
- **David H. Bailey and Marcos Lopez de Prado** (2014). [The Deflated Sharpe Ratio: Correcting for Selection Bias, Backtest Overfitting and Non-Normality](https://doi.org/10.2139/ssrn.2460551).
- **David H. Bailey et al.** (2015). [The Probability of Backtest Overfitting](https://doi.org/10.2139/ssrn.2326253).
- **Giuseppe A. Paleologo** (2025). The Elements of Quantitative Investing. *John Wiley & Sons*.
- **Kent Daniel and Tobias J. Moskowitz** (2016). [Momentum crashes](https://doi.org/10.1016/j.jfineco.2015.12.002). *Journal of Financial Economics*.
- **Marcos Lopez de Prado** (2018). Advances in Financial Machine Learning. *John Wiley & Sons*.
- **Narasimhan Jegadeesh and Sheridan Titman** (1993). [Returns to Buying Winners and Selling Losers: Implications for Stock Market Efficiency](https://doi.org/10.1111/j.1540-6261.1993.tb04702.x). *The Journal of Finance*.
- **R. David McLean and Jeffrey Pontiff** (2016). [Does Academic Research Destroy Stock Return Predictability?](https://doi.org/10.1111/jofi.12365). *Journal of Finance*.
- **Ron Kohavi** (1995). A study of cross-validation and bootstrap for accuracy estimation and model selection. *Morgan Kaufmann Publishers Inc.*.
- **Stephen Bates et al.** (2021). [Cross-validation: what does it estimate and how well does it do it?](https://doi.org/10.1080/01621459.2023.2197686).
- **Tobias J. Moskowitz et al.** (2011). [Time Series Momentum](https://doi.org/10.2139/ssrn.2089463).

在遵守原作品许可的前提下,附作者信息全文展示。 许可协议: MIT

此摘要由 Stratmill 研究智能体根据原文撰写,并非原文副本。