A DQN Framework for Single-Stock Market Timing
Summary
This study outlines a Deep Q Network approach to timing trades in one stock. Instead of storing action values for every possible market state in a table, a neural network estimates values for available actions. The agent stores state, action, reward, and next-state experiences, samples them for training, and updates the estimated value for the action taken using the reward plus a discounted estimate of the best next action. The example state combines daily OHLCV data with seven factors; the actions are buying or selling, and changes in account value define rewards.
The document describes a three-layer fully connected network and a chronological training and test split using one Chinese stock, but provides no numerical backtest results in the results section. It calls the implementation a basic learning example. Suggested extensions include a separate target network, a hold action, different state definitions, and tuning network structure. Trading costs, reward design, and out-of-sample robustness remain important considerations, and the proposed extensions are not evaluated here.
Key ideas
- DQN uses a neural network to estimate action values for continuous market states.
- The example stores experiences and trains from sampled transitions using a discounted next-state value.
- Daily OHLCV data and seven factors form the example state, with buy and sell as the actions.
- Account value changes serve as rewards, while trading costs are included in environment design.
- The article describes a single-stock train/test workflow but reports no numerical backtest outcome.
- A target network and a hold action are suggested as possible improvements, not tested findings.
Tags
This summary was written by Stratmill's research agent from the original; it is not a copy of the source.