Skip to content
All library documents

Preparing Momentum Features and Time-Based Data Splits

Code Stratmill research code

Summary

This code describes a data formatter for a momentum model. It defines target returns, normalized returns over several horizons, MACD features, and optional change-point, calendar, and ticker identity inputs. It also assigns columns roles such as target, known input, static input, identifier, and time index, then derives input locations and category counts for a model.

The formatter filters rows with missing values and dates before a configurable test boundary, then splits the earlier data into training and validation sets either ticker by ticker or by date. It fits scalers on training data and transforms the splits, with optional lagged sequences and a buffered test set. These are implementation details rather than evidence of predictive performance. The excerpt is incomplete, and it does not establish whether all features are available at prediction time, whether preprocessing avoids leakage in every branch, or whether the resulting model performs well.

Key ideas

  • The column definitions distinguish targets, known inputs, static features, identifiers, and time indexes.
  • Momentum inputs include normalized returns at multiple horizons and MACD values.
  • Training and validation data can be split per ticker or across dates.
  • Scalers are calibrated on training data before transforming the dataset splits.
  • Optional lag handling creates sequence batches and adds historical context to test data.

Tags

This summary was written by Stratmill's research agent from the original; it is not a copy of the source.