LSEAttention for More Stable Transformer Time-Series Forecasting
Summary
The article presents LSEAttention, a Transformer modification intended to reduce attention collapse and numerical instability in multivariate time-series forecasting. It uses the Log-Sum-Exp reformulation to stabilize Softmax calculations against overflow and underflow, and combines this with GELU to smooth activation values. The described architecture also uses PReLU in feed-forward layers and invertible normalization to address distribution differences between training and test data. The article then discusses adapting existing neural-network components in MQL5 to incorporate these techniques.
The motivation is that unstable attention weights can weaken generalization, especially in variable time-series data. The article reports a historical test comparison in which the updated model made one fewer trade than the baseline, with both recording the same count of profitable trades; it notes the absence of a drawdown in one month as the visible difference. These observations are limited: the text provides no broad evaluation across markets or conditions, and the small reported comparison does not establish a durable performance advantage. The main claim concerns training stability, not proven live profitability.
Key ideas
- Log-Sum-Exp normalization is used to make Softmax attention calculations more numerically stable.
- GELU smooths activation changes that may otherwise amplify extreme attention scores.
- PReLU is used to preserve gradient flow for negative feed-forward activations.
- Invertible normalization is intended to reduce distribution mismatch between training and test data.
- The reported historical comparison is narrow and does not establish a general trading advantage.
Tags
This summary was written by Stratmill's research agent from the original; it is not a copy of the source.