DQN Reinforcement Learning for Daily Equity Index Timing
Summary
This research report explains reinforcement learning concepts, including Markov decision processes, state and action values, reward design, and value-based and policy-based algorithms. Its applied example uses a deep Q-network to time long positions in the Shanghai Composite Index. Market observations over a lookback window form the state; actions are to buy, sell, or hold; and rewards are based on returns over a forecast horizon. Multiple random-seed models are combined by majority vote, with trades executed at the next opening price.
The report describes an out-of-sample backtest from 2017 through June 2022, following training on earlier data, and reports results for both baseline and tuned hyperparameters. It also examines sensitivity to discounting, replay memory, lookback, and forecast horizon. These are historical backtest findings for one index, and the report warns that reinforcement learning can overfit, depends on random seeds and hyperparameters, and is difficult to interpret. It also identifies limited data and the lack of a realistic interactive market simulator as important constraints.
Key ideas
- The report frames reinforcement learning as learning a policy from state, action, and reward feedback over time.
- Its DQN example uses market data as states and buy, sell, or hold as the available actions.
- Signals from models trained with different random seeds are combined by majority vote.
- The report evaluates the strategy out of sample and tests sensitivity to several DQN hyperparameters.
- The findings are limited to one index and carry risks from overfitting, model instability, and limited market simulation.
Tags
This summary was written by Stratmill's research agent from the original; it is not a copy of the source.