Skip to content
All library documents

Building a Time-MoE Trading Model with Sparse Experts

Article MQL5 articles

Summary

This article describes integrating a Time-MoE architecture into an Actor–Director–Critic trading system. It outlines a SwiGLU embedding, a sparse Mixture of Experts layer with a Top-K router and shared expert, and an attention module that combines a recent token with its historical context. The model is assembled from configurable layers, then trained offline and fine-tuned online in the MetaTrader 5 Strategy Tester.

The article reports testing on January 2025 quotes, where the agent produced profits alongside deep drawdowns. That result motivates further work on risk controls and trading criteria. The account is partly truncated, and it does not provide enough detail here to assess the training setup, benchmark the model, or judge whether the reported performance generalizes beyond that test period.

Key ideas

  • The architecture replaces a standard Transformer feed-forward block with sparse expert processing.
  • A Top-K router selects experts for each token, while a shared expert provides an additional processing path.
  • The attention module routes recent-token input alongside the full historical context.
  • The training pipeline combines offline learning with online fine-tuning in the Strategy Tester.
  • The reported test showed profits but also deep drawdowns, leaving risk management as an open issue.

Tags

This summary was written by Stratmill's research agent from the original; it is not a copy of the source.