Skip to content
All library documents

A Reproducible Machine Learning Pipeline for Trading Research

Article MQL5 articles

Summary

The article presents an integrated research pipeline intended to make trading machine learning experiments reproducible and auditable, from tick data through model export for MetaTrader 5. Its data architecture keeps tick data in a session-level RAM cache and stores processed bars in Parquet, aiming to reuse loaded data efficiently while avoiding duplication of large raw datasets. The loader example handles fully cached requests, subsets, and partially overlapping date ranges, and tracks cache hits and misses.

The broader pipeline is described as combining caching, logging, saved artifacts, analysis reports, cross-validation, and ONNX export validation. The article also discusses bid/ask-aware separate models for long and short positions and high-frequency volatility estimation intended to reduce microstructure noise. These are presented as system design and code examples, not as evidence that a particular trading strategy is profitable. Several research decisions, including feature importance and selection of triple-barrier settings, are explicitly outside the article’s scope; the reader must complete those steps and validate data quality, leakage controls, and model performance for their own use.

Key ideas

  • A research pipeline can record data, parameters, outputs, and artifacts to make experiments repeatable.
  • RAM caching can reuse tick data across requests, while processed bars can be stored more compactly on disk.
  • Partial cache coverage can be extended by loading only missing date ranges.
  • Long and short models can account for different execution sides by using ask prices for longs and bid prices for shorts.
  • The article focuses on research infrastructure and omits choices such as feature importance and triple-barrier settings.

Tags

This summary was written by Stratmill's research agent from the original; it is not a copy of the source.