Skip to content
All library documents

Generating Synthetic Training Orders from Historical Market Data

Article SuperMind

Summary

This script creates order labels from per-stock historical backtest data, then saves the resulting records in training, validation, test, and combined datasets. It rejects a stock’s data if it is empty, contains missing values, or has very small minimum volume. For accepted data, it selects a range of observations within each date, groups records by date and instrument, and uses group means as the basis for order records.

Order amounts are sampled from a lognormal distribution and scaled by the volume field; nonpositive amounts are discarded. The script assigns a fixed order-type value and separates records at stated calendar cutoffs. It shuffles the available stock list with a seeded random generator and stops after reaching its target number of accepted stocks. This is a data-preparation procedure rather than a trading strategy: it does not explain how the generated labels correspond to real order behavior or validate their usefulness for model training. Its output depends on the source dataset, filtering rules, and chosen distribution.

Key ideas

  • The script filters out stock datasets with missing data or inadequate volume values.
  • It aggregates selected observations by date and instrument before creating order records.
  • Synthetic order amounts are drawn from a lognormal distribution scaled by observed volume.
  • Records are split into training, validation, and test sets using calendar date cutoffs.
  • A fixed random seed makes the stock shuffle and sampling process reproducible.

Tags

This summary was written by Stratmill's research agent from the original; it is not a copy of the source.