Skip to content
All library documents

Training a DDPG Trading Agent for Chinese A-Shares with FinRL

Article BigQuant

Summary

This article explains how to model multi-stock trading as a Markov decision process and train a deep deterministic policy gradient (DDPG) agent with FinRL. The state includes prices, holdings, and cash; actions change holdings through buying, selling, or holding; and the reward reflects changes in total portfolio value. DDPG uses actor and critic networks, with experience replay to reduce correlations among training samples. The workflow covers downloading and processing A-share data, configuring training and trading environments, tuning parameters, and evaluating trades with portfolio metrics.

The reported experiment uses daily history for 15 Shanghai 50 constituents, splitting data into training and test periods. It reports a final portfolio value, annualized return, and Sharpe ratio for the test period. These are results from one historical experiment, not evidence of robust live performance. The article notes that this research area remains exploratory; its summary does not establish how transaction costs, changing market conditions, or broader out-of-sample tests affect results.

Key ideas

  • The trading environment represents prices, holdings, and cash as state, with portfolio value changes as rewards.
  • DDPG combines actor and critic networks and updates them using replayed experience.
  • FinRL provides components for data preparation, environment setup, agent training, and backtesting.
  • The example evaluates a selected group of Chinese large-cap stocks using a chronological training and test split.
  • Reported historical performance is specific to the experiment and does not establish live robustness.

Tags

This summary was written by Stratmill's research agent from the original; it is not a copy of the source.