DDPG for Continuous Trading Actions, Position Sizing, and Risk Controls
Summary
This article introduces Deep Deterministic Policy Gradient (DDPG) as a reinforcement-learning approach for trading agents that must choose continuous action values. Instead of expanding a discrete action list for every combination of direction, volume, stop loss, and take profit, DDPG uses an Actor to propose an action and a Critic to estimate its value. The described training loop stores state, action, next-state, and reward experiences, samples them for learning, adds noise to encourage exploration, and uses target networks with soft updates for stability.
The practical implementation discusses updating target-network parameters in an OpenCL context and applies the method to an agent that selects trading direction, volume, and exit levels. The article reports that the model made a profit on its training set, but did not achieve similar results outside that set. It identifies online training and limited parallel exploration as constraints. The example therefore illustrates implementation mechanics and a generalization problem, not evidence of a robust or deployable trading strategy.
Key ideas
- DDPG uses an Actor to select continuous actions and a Critic to evaluate state-action pairs.
- Continuous actions allow a trading agent to choose position volume and stop-loss and take-profit levels without enumerating every discrete combination.
- Experience replay, exploration noise, and target networks are part of the described training process.
- Soft updates blend target-network parameters toward trained-network parameters to support training stability.
- The reported profit was confined to the training set, exposing a generalization limitation in the example.
Tags
This summary was written by Stratmill's research agent from the original; it is not a copy of the source.