Attention-Based Deep Q-Learning for Straddle Option Trading
Summary
The study describes an automated long-volatility approach that trades straddle options without forecasting whether prices will rise or fall. It applies a Transformer-based double deep Q-network, using attention over time-series inputs and information from multiple cycles. Its reward design emphasizes excess earnings over a longer horizon while accounting for losses beyond a stop threshold, and resistance levels provide additional context during uncertain market conditions.
Experiments cover Chinese stocks, Brent crude oil, and Bitcoin. The authors report that their attention-based model had the lowest maximum drawdown across the tested markets and higher average returns than comparison models when crude oil was excluded. The reported advantage is therefore not uniform across all markets. The document gives no detail on option selection, implementation costs, sample periods, or robustness beyond these experiments, so the results do not establish how the approach would perform in live trading.
Key ideas
- Straddle options seek to benefit from volatility without requiring a directional price forecast.
- The model combines a Transformer and double deep Q-learning with temporal and multi-cycle attention.
- Its reward function prioritizes longer-term excess earnings while incorporating a stop threshold.
- Resistance levels are used as reference information when price direction is uncertain.
- The reported drawdown and return comparisons come from three markets, with the return advantage excluding crude oil.
Tags
Full text
# Automated Trading System for Straddle-Option Based on Deep Q-Learning # Automated Trading System for Straddle-Option Based on Deep Q-Learning Straddle Option is a financial trading tool that explores volatility premiums in high-volatility markets without predicting price direction. Although deep reinforcement learning has emerged as a powerful approach to trading automation in financial markets, existing work mostly focused on predicting price trends and making trading decisions by combining multi-dimensional datasets like blogs and videos, which led to high computational costs and unstable performance in high-volatility markets. To tackle this challenge, we develop automated straddle option trading based on reinforcement learning and attention mechanisms to handle unpredictability in high-volatility markets. Firstly, we leverage the attention mechanisms in Transformer-DDQN through both self-attention with time series data and channel attention with multi-cycle information. Secondly, a novel reward function considering excess earnings is designed to focus on long-term profits and neglect short-term losses over a stop line. Thirdly, we identify the resistance levels to provide reference information when great uncertainty in price movements occurs with intensified battle between the buyers and sellers. Through extensive experiments on the Chinese stock, Brent crude oil, and Bitcoin markets, our attention-based Transformer-DDQN model exhibits the lowest maximum drawdown across all markets, and outperforms other models by 92.5\% in terms of the average return excluding the crude oil market due to relatively low fluctuation.
Shown in full with attribution under the source's licence. Licence: abstract CC0
This summary was written by Stratmill's research agent from the original; it is not a copy of the source.