مواد پر جائیں
لائبریری کی تمام دستاویزات

ٹریڈنگ خصوصیات اور سگنل جانچ کے لیے قابلِ تکرار لائبریری طریقۂ کار

نوٹ بک Machine Learning for Trading

خلاصہ

دستاویز معیاری بازار ڈیٹا لوڈرز، پہلے سے تیار خصوصیات کے رجسٹر، اور سگنل کی تشخیص کو ملانے والا طریقۂ کار پیش کرتی ہے۔ خصوصیات کا حساب ناموں، پیرامیٹر لغات، یا YAML کے ذریعے متعین کیا جا سکتا ہے، جس سے قابلِ تکرار پائپ لائنز اور اشاروں کے متعدد افق ممکن ہوتے ہیں۔ یہ ایک دستی RSI نفاذ کا موازنہ کرتی ہے، جو ایکسپونینشل اسموڈنگ استعمال کرتا ہے، لائبریری کے نفاذ سے، جو وائلڈر اسموڈنگ استعمال کرتا ہے، اور بتاتی ہے کہ لائبریری خصوصیات یا حسبِ ضرورت کوڈ کب مناسب ہیں۔ پولارس کی مثالوں میں گروپ شدہ رولنگ حساب، وقت کے نقطے کے مطابق as-of جوائن، اور بڑے ڈیٹا سیٹس کے لیے سست پروسیسنگ شامل ہیں۔

جانچ کے لیے طریقۂ کار شمار کردہ مومینٹم خصوصیت تشخیصی آلے کو دیتا ہے، جو معلوماتی کوائف، ٹی شماریات، اور کوانٹائل اسپریڈ رپورٹ کرتا ہے۔ مثال میں خام مومینٹم عامل کا کراس سیکشنل پیش گوئی سے تعلق تقریباً صفر ہے اور ETF کائنات میں روایتی اہمیت تک نہیں پہنچتا، جو ظاہر کرتا ہے کہ ایک اشارہ اکیلا کمزور ہو سکتا ہے۔ یہ پورے خصوصیت انتخاب یا اسٹریٹیجی کے جائزے کے بجائے نظام اور طریقۂ کار کا تعارف ہے؛ دستاویز وسیع تر تشخیص اور ماڈل امتزاج کے لیے بعد کے مواد کی طرف رہنمائی کرتی ہے۔

اہم خیالات

  • متحد ڈیٹا لوڈرز خصوصیات کی تیاری اور جانچ کے لیے ان پٹس کو معیاری بناتے ہیں۔
  • خصوصیات کا رجسٹر دریافت، میٹا ڈیٹا کے معائنے، پیرامیٹر بندی، اور قابلِ تکرار ترتیب میں مدد دیتا ہے۔
  • دستی اور لائبریری اشاروں میں فرق اس لیے ہو سکتا ہے کہ ان کے اسموڈنگ یا تخمینے کے طریقے مختلف ہوں۔
  • گروپ شدہ رولنگ عمل، as-of جوائن، اور سست سوالات قابلِ توسیع اور وقت کے مطابق محفوظ پائپ لائنز کے مفید طریقے ہیں۔
  • خام عامل کو فارورڈ منافع کے مقابل جانچنا چاہیے؛ کسی ایک اشارے کی اکیلے پیش گوئی کی قوت کم ہو سکتی ہے۔

ٹیگز

مکمل متن
# The ml4t Library Ecosystem


```python
"""The ml4t Library Ecosystem - unified data, feature, and diagnostic libraries for Chapters 7-12."""
```

# The ml4t Library Ecosystem

**Docker image**: `ml4t`

**Chapter 7: Defining the Learning Task**
**Section Reference**: 7.1 - Data Preprocessing and Encodings

## Purpose

This notebook introduces the **ml4t library ecosystem** used throughout
Chapters 7-12: data loaders (`ml4t-data`), feature computation
(`ml4t-engineer`), and evaluation tools (`ml4t-diagnostic`).

## Learning Objectives

1. Load datasets using the unified `ml4t-data` loaders
2. Discover and compute features via the `ml4t-engineer` registry
3. Configure features with lists, dicts, or YAML for reproducibility
4. Validate library features against manual implementations
5. Preview feature evaluation with `ml4t-diagnostic` (full treatment in
   `05_signal_evaluation`)
6. Understand when to use the library vs manual code

## Data Policy

All examples use **real ETF data** (no synthetic data).

## Prerequisites

- `01_data_quality_diagnostics` - establishes the ETF dataset shape used here.
- Polars basics (`with_columns`, `over`, `lazy`).

```python
from __future__ import annotations

from datetime import datetime

import numpy as np
import polars as pl
from IPython.display import display
from ml4t.diagnostic.signal import analyze_signal
from ml4t.engineer import compute_features
from ml4t.engineer.core.registry import get_registry

from data import load_etfs
```

```python
# Production defaults - Papermill injects overrides for CI
SPY_START_DATE = "2015-01-01"
```

## ml4t-data: Unified Data Loaders

The `data` module provides consistent interfaces for all seven
datasets introduced in Chapter 2. Each loader returns a Polars DataFrame
with standardized column names.

```python
etfs = load_etfs()
print(f"ETF universe: {len(etfs):,} rows, {etfs['symbol'].n_unique()} symbols")
print(f"Columns: {etfs.columns}")
print(f"Date range: {etfs['timestamp'].min()} to {etfs['timestamp'].max()}")
```

```python
# For demos: SPY only
spy = etfs.filter(
    (pl.col("symbol") == "SPY")
    & (pl.col("timestamp") >= datetime.fromisoformat(SPY_START_DATE).date())
).sort("timestamp")

print(f"SPY: {len(spy):,} rows from {spy['timestamp'].min()} to {spy['timestamp'].max()}")
```

## ml4t-engineer: Feature Registry

The `ml4t-engineer` library provides 120+ pre-built features with
consistent naming, validation against reference implementations (TA-Lib
where applicable), and self-documenting metadata.

```python
registry = get_registry()
all_features = registry.list_all()

# Features by category
categories = {}
for name in all_features:
    metadata = registry.get(name)
    categories.setdefault(metadata.category, []).append(name)

print(f"Total features available: {len(all_features)}\n")
print("Features by category:")
for cat, feats in sorted(categories.items()):
    examples = ", ".join(feats[:4])
    suffix = ", ..." if len(feats) > 4 else ""
    print(f"  {cat:20s}: {len(feats):3d}  ({examples}{suffix})")
```

### Feature Metadata

Each registry entry carries its default parameters, input requirements, and
a description, plus a closed-form formula where a standard one exists (about
a quarter of the catalog: momentum and price transforms tend to have one,
while estimator-based volatility and microstructure features do not). This
makes features discoverable without reading source code.

```python
for feature_name in ["rsi", "atr", "garman_klass_volatility"]:
    meta = registry.get(feature_name)
    print(f"\n{'=' * 50}")
    print(f"Feature:     {meta.name}")
    print(f"Category:    {meta.category}")
    print(f"Description: {meta.description}")
    print(f"Formula:     {meta.formula or '(not recorded)'}")
    print(f"Parameters:  {meta.parameters}")
    print(f"Input type:  {meta.input_type}")
```

## Config-Driven Feature Computation

`compute_features()` accepts three input formats, from simplest to most
reproducible:

1. **List of names** - default parameters
2. **List of dicts** - custom parameters per feature
3. **YAML config** - stored configuration for pipelines

### Simple Feature List

```python
result = compute_features(spy, ["rsi", "sma", "ema", "atr"])

new_cols = [c for c in result.columns if c not in spy.columns]
print(f"Computed {len(new_cols)} features: {new_cols}")
display(result.select(["timestamp", "close"] + new_cols).tail(5))
```

### Parameterized Features

A list of dicts sets explicit parameters per feature. Each feature name
resolves to one output column per call, so a single call holds one parameter
set per feature. To compute several horizons of the *same* indicator (a
common multi-timeframe setup), call `compute_features` once per parameter set
and suffix the columns before joining.

Each entry below names one feature with an explicit parameter set. Bollinger bands take
their two deviations separately, following TA-Lib, so the upper and lower band need not
sit the same distance from the moving average.

```python
parameterized = [
    {"name": "rsi", "params": {"period": 10}},
    {"name": "atr", "params": {"period": 20}},
    {"name": "bollinger_bands", "params": {"period": 20, "nbdevup": 2.0, "nbdevdn": 2.0}},
]

result = compute_features(spy, parameterized)

new_cols = [c for c in result.columns if c not in spy.columns]
print(f"Parameterized features ({len(new_cols)} columns): {', '.join(new_cols)}")
```

```python
# Multiple horizons of one indicator: one call per period, suffix, then join.
rsi_multi = spy.select(["timestamp"])
for period in (10, 21, 63):
    horizon = compute_features(spy, [{"name": "rsi", "params": {"period": period}}])
    rsi_multi = rsi_multi.join(
        horizon.select(["timestamp", pl.col("rsi").alias(f"rsi_{period}")]),
        on="timestamp",
    )

rsi_cols = [c for c in rsi_multi.columns if c != "timestamp"]
print(f"Multi-horizon RSI columns: {rsi_cols}")
display(rsi_multi.tail(5))
```

### YAML Configuration (Production)

For reproducibility across notebooks and case studies, store feature
configurations in YAML:

```yaml
features:
  - name: rsi
    params:
      period: 14

  - name: macd
    params:
      fast: 12
      slow: 26
      signal: 9

  - name: atr
    params:
      period: 14
```

Load with: `compute_features(df, config_path="features.yaml")`

## Validation: Library vs Manual

The library uses Wilder's smoothing (matching TA-Lib) while a naive
implementation might use EWM span. Let's compare to understand the
difference.

```python
def manual_rsi(df: pl.DataFrame, period: int = 14) -> pl.DataFrame:
    """Manual RSI using EWM span (not Wilder's smoothing)."""
    return (
        df.with_columns(pl.col("close").diff().alias("delta"))
        .with_columns(
            pl.when(pl.col("delta") > 0).then(pl.col("delta")).otherwise(0).alias("gain"),
            pl.when(pl.col("delta") < 0).then(-pl.col("delta")).otherwise(0).alias("loss"),
        )
        .with_columns(
            pl.col("gain").ewm_mean(span=period, adjust=False).alias("avg_gain"),
            pl.col("loss").ewm_mean(span=period, adjust=False).alias("avg_loss"),
        )
        .with_columns(
            (100 - 100 / (1 + pl.col("avg_gain") / pl.col("avg_loss"))).alias("rsi_manual"),
        )
    )


manual_result = manual_rsi(spy.select(["timestamp", "close"]))
library_result = compute_features(spy, [{"name": "rsi", "params": {"period": 14}}])

comparison = manual_result.join(
    library_result.select(["timestamp", "rsi"]),
    on="timestamp",
    how="inner",
).select(["timestamp", "rsi_manual", "rsi"])

print("RSI comparison (manual EWM vs library Wilder's):")
display(comparison.filter(pl.col("rsi").is_not_null()).tail(10))

# Both columns can carry a leading NaN/null from the warm-up window; in Polars
# NaN is distinct from null, so guard against both before averaging.
diff = comparison.filter(
    pl.col("rsi").is_not_null()
    & pl.col("rsi").is_not_nan()
    & pl.col("rsi_manual").is_not_null()
    & pl.col("rsi_manual").is_not_nan()
).with_columns((pl.col("rsi") - pl.col("rsi_manual")).abs().alias("abs_diff"))
print(f"Mean absolute difference: {diff['abs_diff'].mean():.4f}")
print("Differences are due to smoothing method (EWM span vs Wilder's).")
```

### When to use the library vs manual code

| Use ml4t-engineer | Implement manually |
|---|---|
| Standard indicators (RSI, MACD, ATR) | Custom alpha factors |
| Production pipelines | Pedagogical demonstrations |
| Cross-validation with TA-Lib | Non-standard variations |

## ml4t-diagnostic: Feature Evaluation (Preview)

The third library, `ml4t-diagnostic`, closes the loop: once a feature is
computed, `analyze_signal()` measures whether it predicts forward returns in
the cross-section, reporting the Information Coefficient (rank correlation of
factor to forward return), its t-statistic, and quantile spreads. Here we run
a single call to show the `ml4t-engineer` to `ml4t-diagnostic` handoff on the
full ETF panel. Notebook `05_signal_evaluation` develops the full workflow
(ICIR, quantile monotonicity, turnover, half-life).

```python
# Compute one momentum factor across the whole ETF cross-section, then evaluate.
panel = etfs.sort(["symbol", "timestamp"])
factor = (
    panel.group_by("symbol", maintain_order=True)
    .map_groups(lambda g: compute_features(g, [{"name": "mom", "params": {"period": 21}}]))
    .select(["timestamp", "symbol", pl.col("mom").alias("factor")])
    .drop_nulls()
)
prices = panel.select(["timestamp", "symbol", pl.col("close").alias("price")])

signal = analyze_signal(
    factor,
    prices,
    periods=(1, 5, 21),
    quantiles=5,
    date_col="timestamp",
    asset_col="symbol",
    price_col="price",
    factor_col="factor",
)

print(f"Evaluated on {signal.n_assets} assets over {signal.n_dates:,} dates")
for horizon in ("1D", "5D", "21D"):
    print(
        f"  {horizon:>3s}  IC={signal.ic[horizon]:+.4f}  "
        f"t-stat={signal.ic_t_stat[horizon]:+.2f}  "
        f"quantile spread={signal.spread[horizon]:+.4f}"
    )
```

The 21-day momentum factor carries near-zero cross-sectional IC on this ETF
universe, and the t-statistics do not clear conventional significance. A
single raw indicator rarely predicts returns on its own; the feature
engineering and model-based combination methods in Chapters 8 through 12 are
what turn raw features into usable signals. The point here is the interface:
`ml4t-engineer` produces the factor and `ml4t-diagnostic` scores it, with no
glue code in between.

## Key Polars Patterns for Feature Engineering

These three patterns account for most of the feature-engineering code in this book.

### GroupBy + Rolling via `.over()`

Polars' `.over()` expression is the window function syntax - parallel
and significantly faster than pandas' `groupby().transform()`.

```python
multi_symbol = etfs.filter(pl.col("symbol").is_in(["SPY", "QQQ", "IWM"])).sort(
    ["symbol", "timestamp"]
)

features = multi_symbol.with_columns(
    pl.col("close").pct_change().over("symbol").alias("ret_1d"),
    (pl.col("close").pct_change().rolling_std(21).over("symbol") * np.sqrt(252)).alias("vol_21d"),
    pl.col("close").rolling_mean(21).over("symbol").alias("sma_21"),
)

print("Window function features:")
display(features.select(["symbol", "timestamp", "close", "ret_1d", "vol_21d"]).tail(10))
```

**Key pattern**: All transformations in ONE `with_columns()` call
for parallel execution - never chain separate calls.

### ASOF Joins (Point-in-Time Matching)

ASOF joins match by the closest timestamp. Critical for:
- Trade-quote matching
- Fundamental data alignment (announcement date to trading date)
- Macro data alignment (release date to trading date)

```python
trades = spy.select(["timestamp", "close"]).rename({"close": "trade_price"}).head(1000)

quotes = (
    spy.select(["timestamp", "close"])
    .rename({"close": "quote_price"})
    .with_columns(pl.col("timestamp").cast(pl.Date))
    .head(1000)
)

matched = trades.sort("timestamp").join_asof(
    quotes.sort("timestamp"),
    on="timestamp",
    strategy="backward",
)

print("ASOF join result (most recent quote for each trade):")
display(matched.head(5))
```

**Requirements**: Both DataFrames sorted by join key; use
`strategy="backward"` for point-in-time safety.

### Lazy Evaluation (Large File Processing)

For large files, `scan_parquet()` pushes filters to the storage layer.
Here we demonstrate this using a loader to first get the data, then
showing the lazy API pattern with `LazyFrame`.

```python
# In production this would start from pl.scan_parquet on the file itself.
spy_lazy = (
    load_etfs(symbols=["SPY"], start_date="2020-01-01")
    .lazy()
    .select(["timestamp", "close", "volume"])
    .with_columns(pl.col("close").pct_change().alias("returns"))
)

result = spy_lazy.collect()
print(f"Lazy query: {len(result):,} rows")
print(f"\nQuery plan:\n{spy_lazy.explain()}")
```

## Key Takeaways

1. **ml4t-data** provides unified loaders for all seven datasets
2. **ml4t-engineer** offers 120+ validated features via a registry API
3. **Config-driven** computation (list, dict, YAML) ensures reproducibility
4. **ml4t-diagnostic** scores features against forward returns via
   `analyze_signal`; notebook `05_signal_evaluation` covers the full workflow
5. **Library for production**, manual code for teaching and custom factors
6. **Polars patterns** (`.over()`, ASOF joins, lazy scans) power the pipelines

**Next**: Chapter 8 notebooks build features manually to explain the
economics, then use the registry for case study pipelines.

ماخذ کا حوالہ دیتے ہوئے مکمل متن دکھایا گیا ہے، ماخذ کے لائسنس کے تحت۔ لائسنس: MIT

یہ خلاصہ اصل ماخذ سے Stratmill کے تحقیقی ایجنٹ نے لکھا ہے؛ یہ ماخذ کی نقل نہیں۔