Attention and LSTM Momentum Trading in Equities
Summary
This research adapts a Momentum Transformer architecture to equities and compares its intended use with time-series momentum and mean-reversion strategies. The model combines attention, which can relate information across the training window, with an LSTM, which processes sequential patterns. The authors present the combination as a way to capture longer-term dependencies and respond to changing market conditions, including the Covid pandemic.
The reported average return is 4.14%, described as similar to results from the earlier paper, while the average Sharpe ratio is 1.12. The authors attribute the lower Sharpe, relative to that work, to higher volatility in stocks than in futures and equity indices. The brief description does not provide details on the test period, equity universe, transaction-cost assumptions, benchmark results, or statistical uncertainty, so the figures alone are insufficient to assess robustness or live-trading performance.
Key ideas
- The study applies a Momentum Transformer architecture to equities.
- The model combines attention with an LSTM to represent sequential and longer-range patterns.
- The authors compare the approach with time-series momentum and mean-reversion strategies.
- They report a 4.14% average return and a 1.12 average Sharpe ratio.
- The stated results lack details needed to assess robustness or implementation costs.
Tags
Full text
# Enhanced Momentum with Momentum Transformers # Enhanced Momentum with Momentum Transformers The primary objective of this research is to build a Momentum Transformer that is expected to outperform benchmark time-series momentum and mean-reversion trading strategies. We extend the ideas introduced in the paper Trading with the Momentum Transformer: An Intelligent and Interpretable Architecture to equities as the original paper primarily only builds upon futures and equity indices. Unlike conventional Long Short-Term Memory (LSTM) models, which operate sequentially and are optimized for processing local patterns, an attention mechanism equips our architecture with direct access to all prior time steps in the training window. This hybrid design, combining attention with an LSTM, enables the model to capture long-term dependencies, enhance performance in scenarios accounting for transaction costs, and seamlessly adapt to evolving market conditions, such as those witnessed during the Covid Pandemic. We average 4.14% returns which is similar to the original papers results. Our Sharpe is lower at an average of 1.12 due to much higher volatility which may be due to stocks being inherently more volatile than futures and indices.
Shown in full with attribution under the source's licence. Licence: abstract CC0
This summary was written by Stratmill's research agent from the original; it is not a copy of the source.