Skip to content
All library documents

Reinforcement Learning Foundations and DQN Index Timing

Article BigQuant

Summary

The document introduces reinforcement learning through the interaction of an agent and an environment, where actions produce feedback and the objective is to maximize discounted future rewards. It frames the problem as a Markov decision process, explaining state and action values and the role of Bellman equations. It surveys value-based approaches, including Monte Carlo, temporal difference learning, SARSA, Q-learning, and deep Q-networks, alongside policy-based methods such as policy gradients, REINFORCE, and actor-critic algorithms.

Its application uses DQN for daily timing of the Shanghai Composite Index. Market data over a lookback window form the state; actions are to go long, close a long position, or hold, and rewards reflect returns over a forecast window. The report describes training on 2007–2016 data and testing on 2017 through June 2022, with a stated one-way fee assumption. It reports out-of-sample excess return, Sharpe ratio, and turnover, and says parameter changes improved results. These are historical backtest figures; the excerpt gives limited detail on robustness, implementation costs, or generalization beyond this index and period.

Key ideas

  • Reinforcement learning learns an action policy from rewards generated through interaction with an environment.
  • Markov decision processes represent states, actions, rewards, transitions, and discounted future returns.
  • DQN extends Q-learning with neural networks, replay memory, and a target network.
  • The index-timing example uses market history as state and long, close, or hold as available actions.
  • The reported out-of-sample backtest uses the Shanghai Composite and a specified train-test split, but does not establish future performance.

Tags

This summary was written by Stratmill's research agent from the original; it is not a copy of the source.