עבור לתוכן
כל מסמכי הספרייה

בקטסטים למאפייני מניות US: תזמון, עלויות ומגבלות הקצאה

קוד Machine Learning for Trading

סיכום

תצורה זו מתארת תיקי מניות US חודשיים בלונג ובשורט, המדורגים לפי מאפייני חברות. היא מגדירה יקום של תצפיות מלאות בלבד עם 46 מאפיינים, החלטות בסוף החודש, ביצוע בפתיחת המושב הבא, משקל שווה לרגלי הלונג והשורט ועלויות עסקה משמעותיות, לרבות עלות השאלת ניירות ערך. מרשם המאפיינים מקבץ אותות כגון ערך, איכות, השקעה ומומנטום, ומתעד את השהיות העדכון המשוערות ואת אופני הכשל שלהם. ההגדרה קובעת גם תוויות תשואה חודשיות והערכה בווק-פורוורד עם תקופת 2016 שמורה לבדיקה.

דיון תכנוני מרכזי עוסק באמידת מנגנון ההקצאה. בחלון מבט לאחור של 12 חודשים, דרגת אומדני הקווריאנציה היא לכל היותר אחת עשרה, ולכן הקצאה מבוססת מטריצה אינה מזוהה עבור מיפויים גדולים יותר וניתנת לאמידה רק עבור המיפוי הקטן ביותר ברשימה. שיטות סקלריות של הופכי תנודתיות ושל שוויון בסיכון ניתנות לאמידה, אך הושמטו במכוון. ההערות מדגישות שעמודות התנודתיות שסופקו עשויות לתמוך בגישות אחרות. אלה בחירות מפרט והנמקות, ולא תוצאות בקטסט מדווחות; תזמון המאפיינים, העלויות ומוסכמות הנתונים עדיין מגבילים את הפרשנות.

רעיונות מרכזיים

  • האסטרטגיה מדרגת חברות מדי חודש ויוצרת רגלי לונג ושורט במשקל שווה.
  • מאפיינים חשבונאיים שנתיים משתמשים בהשהיה של שישה חודשים, ואילו למאפייני מחירים חודשיים לא הוגדרה השהיה.
  • לחלון קווריאנציה של 12 חודשים דרגה של עד אחת עשרה, דבר שמגביל הקצאה מבוססת מטריצה בתיקים גדולים יותר.
  • שיטות שקלול סקלריות לפי תנודתיות נותרות ניתנות לאמידה, אך אינן כלולות בתפריט ההקצאה המתוכנן.
  • ההערכה משתמשת בחלוקות ווק-פורוורד ובתקופת החזקה 2016 שהוגדרה בנפרד.

תגיות

הטקסט המלא
# setup.yaml


```yaml
strategy_id: us_firm_characteristics
setup_version: v2

universe:
  inclusion_rule: complete_characteristic_case
  identifiers: anonymous_split_scoped_firm_axis
  note: >-
    The authors retain observations with all 46 characteristics available.
    Anonymous identifiers are persistent within each released tensor block;
    the archive does not publish a mapping between blocks.
  # Approximate active cross-sectional breadth. Display metadata only; the
  # backtest counts assets from data.
  n_assets: 2500

decision:
  cadence: monthly_month_end
  snapshot: month_end_close
  execution_delay: next_bar_open
  characteristic_availability: provider_standard_conventions
  yearly_update: end_of_june
  monthly_update: month_end_for_next_month

# Engine-level execution defaults. Single source of truth - Ch16-19 notebooks
# read these via get_backtest_config(); never declare a local INITIAL_CASH
# or share_type constant. Changing values here invalidates every existing
# backtest_hash for this case study (cash + share_type are spec-hash inputs).
#
# THE ALLOCATOR EXCLUSIONS, AND THE REASON IS THE CADENCE RATHER THAN THE PANEL.
#
# An earlier version of this comment said the moment-based allocators were
# excluded because firm identities are not stable across periods. That is false
# of this dataset and its own columns refute it: ``r36_13`` is a return measured
# from month 36 to month 13 before the decision date, and it is non-null on all
# 804,530 rows. A feature computed from a 36-month window cannot be complete
# without 36 months of continuous history per firm. Whatever the anonymous
# split-scoped identifiers do across tensor blocks, they are stable enough within
# one to support a three-year lookback on every row.
#
# The real constraint is the observation count. ``allocator_lookback`` is 12 and
# the bars are monthly, so a covariance is estimated from twelve points and has
# rank at most eleven. What that rules out depends on how many names are held,
# and ``top_k_grid`` below is [5, 10, 20, 50] against a long-short mapping, so
# the name count is twice the top_k:
#
#   top_k=5  -> 10 names -> rank 11 covers it        -> IDENTIFIED
#   top_k=10 -> 20 names -> unidentified
#   top_k=20 -> 40 names -> unidentified
#   top_k=50 -> 100 names -> unidentified
#
# So ``hrp`` (allocation.py:516, rolling correlation matrix) and
# ``mvo_ledoit_wolf`` (allocation.py:362, Ledoit-Wolf covariance) are structurally
# unidentified on three of the four mappings and estimable on the smallest. Not on
# any of them for a shrinkage reason: `sweep_config.get_allocators` injects the
# case-study lookback as ``lookback`` for mvo, OVERRIDING `compute_mvo_weights`'
# own 126 default, and shrinkage does not manufacture the missing observations.
#
# That argument does NOT reach ``inverse_vol`` (allocation.py:295) or
# ``risk_parity`` (allocation.py:321). Neither builds a matrix: both take a
# per-asset rolling standard deviation and weight by 1/vol or 1/vol^1.5, one
# scalar per name, full rank by construction, and estimable from twelve points.
# They are excluded by decision rather than by identification: the allocation
# stage is kept to the equal-weight baseline and the two lookback-free
# alternatives, and building the rolling volatility they need is work this case
# study is not spending.
#
# Worth recording so it is not re-derived: the panel already carries supplied
# volatility measures - ``Variance``, ``IdioVol``, ``Resid_Var``, ``Beta`` and
# ``MktBeta`` are all columns of ``labels/prices.parquet``. An allocator weighting
# by one of those would need no rolling window and no lookback at all. Nobody has
# asked for one and it is not scheduled; the option exists and was not taken.
#
# ``allocator_lookback`` stays because ``get_backtest_config`` requires the key,
# and it is inert while the menu declares no moment-based allocator:
# sweep_config.get_allocators injects it as ``vol_window`` only into those.
#
# Cash defaults to 1_000_000 rather than 100_000: monthly-cadence
# long-short portfolios at top_k=50 carry ~$10K per leg per name at the
# 100k tier - below realistic round-trip granularity for institutional
# US equity. 1M keeps the spec-implied position sizes (notional per name
# / fixed-cost ratio) in a regime the backtest engine resolves cleanly.
execution:
  initial_cash: 1_000_000        # Monthly long-short top_k=50 needs $1M to size cleanly
  share_type: integer            # US equities trade in whole shares
  allocator_lookback: 12         # 1 year of monthly bars

mapping:
  class: long_short_top_k_rebalance
  position_state_space: long_short
  entry_logic: rank_top_k_long_bottom_k_short
  sizing: equal_weight_within_leg

costs:
  class: material
  components: [spread, commission, market_impact, borrow_cost]
  per_leg_cost_bps_range: [5, 20]
  borrow_cost_note: Long-short requires borrow for the short leg.
  era_note: Pre-2001 spreads 15-30 bps; post-2001 (decimalization) 5-15 bps.

evaluation:
  n_splits: 10
  train_size: 10YE
  val_size: 1YE
  holdout_start: '2016-01-01'
  holdout_end: '2016-12-31'
  calendar: null  # Monthly returns; calendar-aware splitting needs daily frequency.
  periods_per_year: 12

labels:
  primary: fwd_ret_1m
  buffer: 1M
  variants:
    - fwd_ret_1m_win
    - fwd_class_1m
  variant_buffers:
    fwd_ret_1m_win: 1M
    fwd_class_1m: 1M
  # How far past its own timestamp each label's outcome resolves. This is the quantity
  # generate_cv_splits seals the last validation fold on, and it is not the buffer above.
  # The release pairs the characteristics observed at the close of month t-1 with the return
  # earned over month t and dates the row by month t, so a row's return is already realised on
  # the timestamp the row carries, and no validation month has to be given back before the
  # holdout opens. 02_labels section D measures that alignment rather than assuming it: ST_REV,
  # a firm's own most recent monthly return, has rank correlation +0.904 with the previous
  # row's label and -0.028 with its own.
  # The buffer above stays 1M. Separating a training window from the validation window that
  # follows it is a different decision from when an outcome becomes known, and that one is
  # deliberately conservative here.
  horizons:
    fwd_ret_1m: 0D
    fwd_ret_1m_win: 0D
    fwd_class_1m: 0D
  # Vectorized-backtest thinning step per label: number of schedule slots
  # to advance per trade so holding periods don't overlap. Authored from
  # (schedule cadence, the span the label measures over); add an entry for any new label.
  rebalance_step:
    fwd_ret_1m: 1
    fwd_ret_1m_win: 1
    fwd_class_1m: 1
  # Continuous return that each classification label is derived from.
  # IC for classification predictions is computed against this column;
  # AUC/accuracy/log_loss are computed against the binary label itself.
  classification_eval_label:
    fwd_class_1m: fwd_ret_1m

features:
  # Window over which the feature matrix is built. Both endpoints are inclusive and
  # 03_financial_features binds them from here rather than retyping them.
  window:
    start: 1990-01-01
    end: 2016-12-31

  # The feature register. Rendered by case_studies.utils.feature_engineering
  # .register_frame and drawn as the timing contract by plot_timing_contract, so
  # `lookback` and `lag` are read by the figure as well as by the prose.
  #
  # `lookback` and `lag` are counted in months, this case study's bar.
  #
  # A caveat that belongs with the numbers rather than under them: the release does
  # not publish a per-characteristic estimation window. What is published is the
  # update convention (decision.yearly_update, decision.monthly_update), and some
  # characteristics name their own window - r36_13 reads 36 months back to 13. The
  # lookbacks below are those two sources and nothing else; where a characteristic
  # publishes neither, the entry is the span of one provider observation. The lag is
  # the load-bearing column for look-ahead and it is fully sourced: an annual variable
  # is published at the end of June against a December fiscal year end.
  families:
    - name: value
      pattern: "BEME|E2P|CF2P|D2P|S2P|A2ME"
      role: signal
      hypothesis: A firm priced low against its fundamentals earns the higher subsequent return
      inputs: released annual accounting characteristics, provider rank-transformed
      lookback: 12
      lag: 6
      frame: cross section
      representation: provider cross-sectional rank in [-0.5, 0.5]
      failure_mode: cheapness that reflects permanent impairment rather than mispricing

    - name: quality
      pattern: "PROF|ROE|ROA|OP|PM|PCM|RNA"
      role: signal
      hypothesis: A more profitable firm earns the higher subsequent return at the same price
      inputs: released annual accounting characteristics, provider rank-transformed
      lookback: 12
      lag: 6
      frame: cross section
      representation: provider cross-sectional rank in [-0.5, 0.5]
      failure_mode: margins mean-revert faster than the annual update reports them

    - name: investment
      pattern: "Investment|NOA|DPI2A|NI|OA|AC"
      role: signal
      hypothesis: A firm growing its asset base aggressively earns the lower subsequent return
      inputs: released annual accounting characteristics, provider rank-transformed
      lookback: 24
      lag: 6
      frame: cross section
      representation: provider cross-sectional rank in [-0.5, 0.5]
      failure_mode: a growth measure spans two annual observations, so one restatement moves both

    - name: momentum
      pattern: "r12_2|r2_1|r12_7|r36_13|ST_REV|LT_Rev|SUV|Rel2High"
      role: signal
      hypothesis: Recent relative price trends persist over a quarter to a year and reverse beyond it
      inputs: released monthly price and return characteristics, provider rank-transformed
      lookback: 36
      lag: 0
      frame: cross section
      representation: provider cross-sectional rank in [-0.5, 0.5]
      failure_mode: trends break at a reversal faster than a 12-month window unwinds

    - name: risk
      pattern: "Beta|MktBeta|IdioVol|Resid_Var|Variance|Spread|LTurnover|LME"
      role: state
      hypothesis: Volatility, liquidity and size describe the regime a signal is read in
      inputs: released monthly risk and liquidity characteristics, provider rank-transformed
      lookback: 12
      lag: 0
      frame: cross section
      representation: provider cross-sectional rank in [-0.5, 0.5]
      failure_mode: >-
        LME travels with this family in the provider's grouping but behaves as a signal,
        so a family-level reading of the group mixes a size premium with regime description

    - name: other
      pattern: "Q|C|CF|AT|ATO|CTO|D2A|FC2Y|Lev|OL|SGA2S"
      role: signal
      hypothesis: Leverage, turnover and cost structure carry information the five named families omit
      inputs: released accounting characteristics, provider rank-transformed
      lookback: 12
      lag: 6
      frame: cross section
      representation: provider cross-sectional rank in [-0.5, 0.5]
      failure_mode: a residual grouping has no single thesis, so a family average over it means little

    # The constructed columns are split by the timing of the members they read, not
    # bundled under one row per construction. A single `composite` entry would have to
    # claim one lookback and one lag for `composite_value` (annual accounting, published
    # end-June) and `composite_momentum` (monthly prices, no publication lag), and would
    # be wrong about one of them whichever pair it named.
    #
    # A construction that mixes the two gets its own row, and the row spans everything it
    # reads: lookback 18 and lag 0. An earlier version gave these the accounting member's
    # lag of 6, which made `plot_timing_contract` draw them as reading nothing from the six
    # months before the decision - and they do, through the momentum member. The composite
    # itself is knowable at the decision timestamp, because its accounting member was
    # published six months earlier and its price member is current, so the lag is zero and
    # the window reaches back to the oldest input.

    - name: composite accounting
      pattern: "composite_value|composite_quality|composite_value_quality"
      role: signal
      hypothesis: Averaging ranks within and across families cancels characteristic-specific noise
      inputs: the released value and quality characteristics, within the same row
      lookback: 12
      lag: 6
      frame: cross section
      representation: equal-weight mean of member ranks, on the members' own scale
      failure_mode: an equal weight asserts the members are equally informative, which is untested here

    - name: composite investment
      pattern: "composite_investment"
      role: signal
      hypothesis: Averaging the investment characteristics cancels measure-specific noise
      inputs: the released investment characteristics, within the same row
      lookback: 24
      lag: 6
      frame: cross section
      representation: equal-weight mean of member ranks, on the members' own scale
      failure_mode: its members span two annual observations, so one restatement moves the composite twice

    - name: composite momentum
      pattern: "composite_momentum"
      role: signal
      hypothesis: Averaging the two 12-month momentum measures cancels their formation-window differences
      inputs: r12_2 and r12_7, within the same row
      lookback: 12
      lag: 0
      frame: cross section
      representation: equal-weight mean of member ranks, on the members' own scale
      failure_mode: both members skip a recent month, so neither reflects the last few weeks

    - name: interaction accounting
      pattern: "interaction_value_x_quality|interaction_value_x_roe"
      role: signal
      hypothesis: Some theses are conditional - cheap is worth more when the firm is also profitable
      inputs: BEME with PROF or ROE, within the same row
      lookback: 12
      lag: 6
      frame: cross section
      representation: product of two member ranks, so the sign encodes agreement
      failure_mode: a product of two centred ranks is large at both extremes and cannot separate them

    - name: composite mixed
      pattern: "composite_value_momentum|composite_quality_momentum"
      role: signal
      hypothesis: Pairing a slow accounting view with a fast price view cancels noise in both
      inputs: an annual accounting composite and the 12-month momentum composite, within the same row
      lookback: 18
      lag: 0
      frame: cross section
      representation: equal-weight mean of two composites, on the members' own scale
      failure_mode: >-
        the two halves move at different speeds, so the average is stale against the price
        member and current against the accounting one at the same time

    - name: interaction mixed
      pattern: "interaction_size_x_value"
      role: signal
      hypothesis: The value premium is larger among smaller firms
      inputs: LME and BEME, within the same row
      lookback: 18
      lag: 0
      frame: cross section
      representation: product of two member ranks, so the sign encodes agreement
      failure_mode: size is current while book-to-market is six months old, so the product pairs two different dates

    - name: interaction momentum
      pattern: "interaction_momentum_x_ivol"
      role: signal
      hypothesis: Momentum reads differently at high idiosyncratic volatility than at low
      inputs: r12_2 and IdioVol, within the same row
      lookback: 12
      lag: 0
      frame: cross section
      representation: product of two member ranks, so the sign encodes agreement
      failure_mode: a product of two centred ranks is large at both extremes and cannot separate them

backtest:
  rebalance:
    # Per-asset rebalance thresholds. A rebalance is skipped when the
    # per-asset weight change is below min_weight_change AND the resulting
    # trade notional is below min_trade_value. At top_k=50 the equal-weight
    # per-asset weight is 1/50 = 2%, comfortably above the 0.5% threshold.
    default:
      min_weight_change: 0.005
      min_trade_value: 100.0
    benchmark:
      min_weight_change: 0.0
      min_trade_value: 0.0
  sweep:
    # Iteration controls per stage. ``signal: 0`` means "all predictions";
    # downstream stages take the top-N from the upstream stage's rank-1.
    # Notebooks read these via get_top_n_predictions(case_study, stage).
    top_n_predictions:
      signal: 0                 # all predictions at the equal-weight baseline
      allocation: 10            # top-10 model configs by equal-weight baseline Sharpe
      # cost_sensitivity takes no entry: the sweep runs the canonical rank-1 carrier,
      # resolved by resolve_solvent_carrier, so there is no top-N to declare.
      risk_overlay: 1           # top-1 of {signal+allocation} per label
    # Skip MVO/HRP when allocator runtime is the bottleneck (intraday CSes).
    expensive_allocators_skip: false
    # The two alternatives below take the cheap path. Equal weight is already
    # the baseline and must not be repeated as an allocation-stage method.
    # Ch16 backtest is equal-weight sized within the selected set. The
    # selection rule is one of: top-k (rank), percentile band (long-only
    # cutoff), or quantile buckets (e.g. quintile long-short, which combines
    # with the long-short mapping). Only top_k_grid is active here;
    # uncomment percentile_grid / quantile_grid to add other selection axes
    # for any label.
    top_k_grid:
      fwd_ret_1m:     [5, 10, 20, 50]
      fwd_ret_1m_win: [5, 10, 20, 50]
      fwd_class_1m:   [5, 10, 20, 50]
    # percentile_grid:
    #   fwd_ret_1m: [80, 90, 95]
    # quantile_grid:
    #   fwd_ret_1m: [5, 10]
    # Ch17 portfolio: reuses top_k_grid above and sweeps over allocators.
    # Equal weight is the baseline; these two are the alternatives, and both are
    # lookback-free. Why the other four are absent is in the execution block above
    # - identification for hrp and mvo_ledoit_wolf, decision for inverse_vol and
    # risk_parity. No max-weight cap.
    allocators:
      - {name: score_weighted,  method: score_weighted}
      - {name: conformal_weighted, method: conformal_weighted}
    # Ch18 cost sensitivity (bps regime).
    cost_grid_bps: [0, 1, 2, 3, 5, 7, 10, 15, 20, 30, 50]
    # Ch19 risk overlays.
    risk_controls:
      position:
        - {name: stop_loss_3pct,  type: stop_loss,     threshold: 0.03}
        - {name: stop_loss_5pct,  type: stop_loss,     threshold: 0.05}
        - {name: stop_loss_10pct, type: stop_loss,     threshold: 0.10}
        - {name: stop_loss_15pct, type: stop_loss,     threshold: 0.15}
        - {name: trailing_1pct,   type: trailing_stop, threshold: 0.01}
        - {name: trailing_2pct,   type: trailing_stop, threshold: 0.02}
        - {name: trailing_3pct,   type: trailing_stop, threshold: 0.03}
        - {name: trailing_5pct,   type: trailing_stop, threshold: 0.05}
        - {name: trailing_10pct,  type: trailing_stop, threshold: 0.10}
        - {name: trailing_15pct,  type: trailing_stop, threshold: 0.15}
        - {name: trailing_20pct,  type: trailing_stop, threshold: 0.20}
        - {name: time_exit_10,    type: time_exit,     bars: 10}
        - {name: time_exit_20,    type: time_exit,     bars: 20}
        - {name: time_exit_40,    type: time_exit,     bars: 40}

modeling:
  gbm:
    libraries: [lightgbm]
    preset: default
    # CPU is the reader-facing reproducible path. Numerical parameters stay
    # fixed when maintainers opt into a different execution backend.
    device: cpu
    # LightGBM's own CPU default. 63 is the GPU default and was carried over with the
    # device when these runs moved off the GPU, so every CPU fit was quantizing the design
    # matrix into a quarter of the bins the library would have used. Coarser bins are
    # faster and lose split points; the reader running this on a CPU gets what the
    # documentation describes.
    max_bin: 255
    num_threads: 8
  tabular_dl:
    # The members are declared per label in config/training/<label>.yaml, which is what
    # load_model_configs reads; a second list here would be a menu nothing consults.
    # Device and thread count are part of a TabM training identity, not provenance beside
    # it: a network's arithmetic depends on both, so a CUDA result and a CPU result of the
    # same configuration are different computations and hash differently.
    device: cuda
    num_threads: 8
  latent_factors:
    persistent_entities: true
    device: cuda
    num_threads: 8
    deterministic_algorithms: true
    # Only unrevised daily market observations enter the SDF context. The
    # one-day availability lag prevents same-close information from entering
    # the next-month forecast.
    macro_series:
      - dff
      - dgs1
      - dgs2
      - dgs3
      - dgs5
      - dgs7
      - dgs10
      - dgs20
      - dgs30
      - t10y2y
      - vixcls
    macro_availability_lag_days: 1
    model_kwargs:
      ipca:
        max_iter: 10000
        factor_ridge: 0.01
        gamma_ridge: 0.01
      sdf:
        checkpoint_epochs: [256, 512, 768, 1024]  # conditional-relative; publishes global 256..1280
        beta_checkpoint_epochs: [256]
        beta_default_checkpoint: 256

causal:
  treatment: r12_2
  # Bars the treatment's own construction window spans, which is what the placebo block has
  # to cover: permuting r12_2 in blocks shorter than this destroys the serial
  # dependence the refutation exists to preserve, and the resulting p-value reads like a
  # refutation without being one. Declared here rather than inferred, because guessing which
  # element of a window list a column was built from puts a wrong number behind a right-looking
  # one. Derived from the construction, not chosen:
  #
  # cumulative return from twelve months back to two months back, on a monthly panel
  # (03_financial_features). Twelve rows, because a row here is a month.
  treatment_window: 12
  confounders: [Beta, IdioVol, LME, Variance]
  method: walk_forward_dml

```

מוצג במלואו בציון המקור ובהתאם לרישיון שלו. רישיון: MIT

הסיכום נכתב בידי סוכן המחקר של Stratmill על סמך המקור; הוא אינו העתק של המקור.