Relative Attention Transformers for Market Forecasting
Summary
The article adapts Relative Molecule Self-Attention Transformer ideas to financial forecasting. Its central method encodes relationships between pairs of input elements, including relative position, distance, and contextual or neighborhood information, then uses these pair embeddings to add biases and values to Transformer attention. The architecture stacks relative-attention layers with feed-forward networks, aggregates their output through attention pooling, and produces predictions with a multilayer network. The article also describes an MQL5 and OpenCL implementation and its use in model training and testing.
The trading application treats relative attention as a way to represent dependencies among market factors and temporal observations. The reported test had only 15 trades, with gains concentrated early in the period followed by a flat stretch. The author therefore presents the outcome as a sign of potential requiring further development, rather than evidence of a ready live system. The excerpt does not give enough detail to assess robustness, compare against baselines, or judge performance across broader periods and markets.
Key ideas
- Pairwise relative embeddings can represent relationships that ordinary absolute positional encoding does not capture.
- The proposed attention mechanism adds relationship-derived biases and values to standard query, key, and value processing.
- The model combines relative-attention layers, attention pooling, and a prediction network.
- The article implements a financial adaptation using MQL5 and OpenCL.
- A short test with few trades and uneven results is preliminary evidence, not validation for live use.
Tags
This summary was written by Stratmill's research agent from the original; it is not a copy of the source.