Skip to content
All library documents

Time-MoE for Multiscale Time-Series Forecasting in Trading

Article MQL5 articles

Summary

The article explains Time-MoE, a decoder-only Transformer foundation model for time-series forecasting, and sketches how its components could be implemented in MQL5. It describes converting each time step into a token, embedding features with SwiGLU, and processing past-only context through causal attention and sparsely routed experts. Multiple output heads predict at different horizons, with head selection adapted to conditions such as volatility or forecast confidence.

Training combines Huber losses across forecast horizons with an auxiliary term to discourage imbalanced expert use. The article reports that the model was trained by its authors on a large multi-domain time-series dataset and describes its parameter scale and sparse inference design, but gives no trading-specific performance results. Its practical implementation discussion is incomplete: it focuses on the SwiGLU embedding stage and defers the sparse MoE implementation to a later installment. Applying the architecture to markets therefore remains a proposed design rather than evidence of a profitable strategy.

Key ideas

  • Point-wise tokenization aims to preserve each time step's information for downstream forecasting.
  • Causal attention restricts model context to historical observations.
  • Sparse expert routing activates a subset of specialized networks for each token.
  • Parallel forecast heads target multiple horizons, while inference may select heads based on market conditions.
  • Huber losses and an expert-balancing penalty support training across horizons and routed components.

Tags

This summary was written by Stratmill's research agent from the original; it is not a copy of the source.