Momentum Transformer for Interpretable Deep Learning Trading
Summary
The document presents the Momentum Transformer, an attention and LSTM hybrid for trading from time-series data. Its attention mechanism connects the model to earlier time steps, while multiple attention heads can represent market dynamics operating over different timescales. The authors compare it with benchmark momentum and mean-reversion strategies and report stronger performance, including after transaction costs.
They also describe the architecture as adapting to changing market regimes, with the SARS-CoV-2 crisis given as an example. Attention weights offer a way to inspect which past observations and factors matter to model decisions. The excerpt does not provide data, experimental design, numerical results, or details of costs and implementation, so the strength and generality of the reported performance cannot be assessed from this description alone.
Key ideas
- The model combines attention with an LSTM to connect current predictions to earlier time steps.
- Multiple attention heads are intended to capture market dynamics at different timescales.
- The authors report better results than benchmark momentum and mean-reversion strategies, including net of transaction costs.
- Attention is used to interpret the influence of historical observations and factors.
- The excerpt gives limited evidence about evaluation details and the portability of the results.
Tags
Full text
# Trading with the Momentum Transformer: An Intelligent and Interpretable Architecture # Trading with the Momentum Transformer: An Intelligent and Interpretable Architecture We introduce the Momentum Transformer, an attention-based deep-learning architecture, which outperforms benchmark time-series momentum and mean-reversion trading strategies. Unlike state-of-the-art Long Short-Term Memory (LSTM) architectures, which are sequential in nature and tailored to local processing, an attention mechanism provides our architecture with a direct connection to all previous time-steps. Our architecture, an attention-LSTM hybrid, enables us to learn longer-term dependencies, improves performance when considering returns net of transaction costs and naturally adapts to new market regimes, such as during the SARS-CoV-2 crisis. Via the introduction of multiple attention heads, we can capture concurrent regimes, or temporal dynamics, which are occurring at different timescales. The Momentum Transformer is inherently interpretable, providing us with greater insights into our deep-learning momentum trading strategy, including the importance of different factors over time and the past time-steps which are of the greatest significance to the model.
Shown in full with attribution under the source's licence. Licence: abstract CC0
This summary was written by Stratmill's research agent from the original; it is not a copy of the source.