Transformer Stock Selection with Time2Vec and Attention
Summary
This article explains attention and Transformer components, then adapts them to predict stock returns from factor time series. Its proposed model adds Time2Vec features to factor inputs, applies self-attention and multi-head attention, and stacks encoder layers with feed-forward networks, residual connections, and normalization. The article describes preprocessing, rolling windows, training, and a backtest workflow for ranking Chinese A-share stocks. The example trains on historical data, predicts future returns, and allocates more capital to higher-ranked selections while limiting position sizes and requiring a minimum holding period.
The article supplies architecture settings and an example evaluation setup, but it does not report numerical backtest results in the provided text. It presents Transformer attention as a way to model dependencies across time and Time2Vec as a way to represent periodic and nonperiodic time features. The implementation is explicitly described as an older version for learning. Its sample settings and backtest period are illustrative; they do not establish out-of-sample profitability or robustness, and practical results depend on the chosen factors, data handling, costs, and execution assumptions.
Key ideas
- Time2Vec combines linear and periodic components to represent time in a sequence model.
- Self-attention weights relationships among time steps in a stock factor sequence.
- The proposed stock model combines time embeddings, attention encoders, and a regression output for return prediction.
- The example uses rolling factor windows, preprocessing, and a historical train-and-predict workflow.
- The document provides no numerical performance results, so the example does not establish strategy profitability.
Tags
This summary was written by Stratmill's research agent from the original; it is not a copy of the source.