Skip to content
All library documents

A DQN Framework for Single-Stock Market Timing

Article BigQuant

Summary

This study outlines a Deep Q Network approach to timing trades in one stock. Instead of storing action values for every possible market state in a table, a neural network estimates values for available actions. The agent stores state, action, reward, and next-state experiences, samples them for training, and updates the estimated value for the action taken using the reward plus a discounted estimate of the best next action. The example state combines daily OHLCV data with seven factors; the actions are buying or selling, and changes in account value define rewards.

The document describes a three-layer fully connected network and a chronological training and test split using one Chinese stock, but provides no numerical backtest results in the results section. It calls the implementation a basic learning example. Suggested extensions include a separate target network, a hold action, different state definitions, and tuning network structure. Trading costs, reward design, and out-of-sample robustness remain important considerations, and the proposed extensions are not evaluated here.

Key ideas

  • DQN uses a neural network to estimate action values for continuous market states.
  • The example stores experiences and trains from sampled transitions using a discounted next-state value.
  • Daily OHLCV data and seven factors form the example state, with buy and sell as the actions.
  • Account value changes serve as rewards, while trading costs are included in environment design.
  • The article describes a single-stock train/test workflow but reports no numerical backtest outcome.
  • A target network and a hold action are suggested as possible improvements, not tested findings.

Tags

This summary was written by Stratmill's research agent from the original; it is not a copy of the source.