MacroHFT: Regime-Specialized Reinforcement Learning for Crypto Trading
Summary
The article explains MacroHFT, a reinforcement-learning framework for high-frequency cryptocurrency trading. Its original design classifies fixed-length market blocks by trend and volatility, then trains specialized agents for the resulting market regimes. A higher-level agent combines their decisions through a meta-policy and uses a memory module to incorporate recent state-action experience. The described memory stores key vectors, retrieves nearby records by distance, and contributes to value estimates and training alignment.
The article also sketches an MQL5 adaptation that changes the original training design: it removes manual regime labeling and separate training phases, trains the coordinator and sub-agents together, and replaces the table memory with a multi-layer memory object. These are design descriptions rather than a completed evaluation. The implementation and testing are unfinished in this installment, and performance on historical data is deferred to a later article. Thus, the framework's claimed adaptability and risk benefits are not established here by trading results.
Key ideas
- The original MacroHFT design specializes separate agents for combinations of market trend and volatility.
- A hyper-agent combines specialized policies and uses stored experience to inform its decisions.
- Regime thresholds are derived from training data and then applied to test data.
- The MQL5 adaptation changes the labeling, training, and memory designs described in the original framework.
- This installment does not report completed historical performance tests.
Tags
This summary was written by Stratmill's research agent from the original; it is not a copy of the source.