Skip to content
All library documents

Sparse Mixture-of-Experts Routing for Time-Series Forecasting

Article MQL5 articles

Summary

The article describes a Time-MoE forecasting architecture that represents each time step as a token, processes temporal context with transformer blocks, and predicts across multiple horizons. Its focus is the sparse mixture-of-experts component: a router selects a subset of specialized models for each input, reducing the computations needed as the model scales. A shared expert remains active to provide a common processing path.

The implementation places sparse activation in the second expert layer and combines it with output aggregation. The article explains a training concern: routing repeatedly to familiar experts can prevent others from specializing. Its design sends learning signals to the router when a selection performs poorly, while not directly training experts that were inactive. It outlines an MQL5 and OpenCL implementation, but offers architectural reasoning rather than forecasting benchmarks or evidence of trading performance.

Key ideas

  • Time-MoE treats time steps as tokens and produces forecasts for multiple horizons.
  • A router activates only selected experts for each input to reduce computation as the architecture grows.
  • A shared expert remains active alongside the routed specialists.
  • The described design applies sparse activation in the second expert layer and aggregates the selected outputs.
  • Router training is intended to encourage exploration of alternative experts when current selections produce poor forecasts.

Tags

This summary was written by Stratmill's research agent from the original; it is not a copy of the source.