Transfer Learning for Crypto Cross-Sectional Momentum Ranking
Summary
The paper addresses overfitting in cross-sectional strategies trained on assets with limited historical data. It introduces Fused Encoder Networks, which combine representations from an encoder-attention module trained on a source dataset with a separate module focused on the smaller target dataset. Sharing and fusing information aims to improve generalization, while self-attention models interactions among instruments both during training and at inference.
The demonstration ranks the ten largest cryptocurrencies by market capitalization for a momentum strategy. The reported results outperform most reference models on several performance measures, including a threefold Sharpe ratio relative to classical momentum and roughly 50% improvement over the strongest benchmark without transaction costs. The paper says the method retains an advantage after accounting for crypto trading costs, but the supplied summary gives no further test design, robustness analysis, or evidence beyond this use case.
Key ideas
- Limited histories can make sophisticated cross-sectional ranking models prone to overfitting.
- Fused Encoder Networks combine representations from source and target datasets to transfer useful information.
- Self-attention allows the model to represent interactions among instruments at inference time.
- The cryptocurrency momentum example reports stronger benchmark performance, including after transaction costs are considered.
Tags
Full text
# Transfer Ranking in Finance: Applications to Cross-Sectional Momentum with Data Scarcity # Transfer Ranking in Finance: Applications to Cross-Sectional Momentum with Data Scarcity Cross-sectional strategies are a classical and popular trading style, with recent high performing variants incorporating sophisticated neural architectures. While these strategies have been applied successfully to data-rich settings involving mature assets with long histories, deploying them on instruments with limited samples generally produce over-fitted models with degraded performance. In this paper, we introduce Fused Encoder Networks -- a novel and hybrid parameter-sharing transfer ranking model. The model fuses information extracted using an encoder-attention module operated on a source dataset with a similar but separate module focused on a smaller target dataset of interest. This mitigates the issue of models with poor generalisability that are a consequence of training on scarce target data. Additionally, the self-attention mechanism enables interactions among instruments to be accounted for, not just at the loss level during model training, but also at inference time. Focusing on momentum applied to the top ten cryptocurrencies by market capitalisation as a demonstrative use-case, the Fused Encoder Networks outperforms the reference benchmarks on most performance measures, delivering a three-fold boost in the Sharpe ratio over classical momentum as well as an improvement of approximately 50% against the best benchmark model without transaction costs. It continues outperforming baselines even after accounting for the high transaction costs associated with trading cryptocurrencies.
Shown in full with attribution under the source's licence. Licence: abstract CC0
This summary was written by Stratmill's research agent from the original; it is not a copy of the source.