Training an Actor–Director–Critic Trading Agent Offline and Online
Summary
The article describes a reinforcement learning setup for trading that adds a Director to the usual Actor–Critic pair. The Actor selects actions, the Critic estimates their reward quality, and the Director classifies whether actions fit the learned strategy. The authors train both evaluators on latent features shared by agents in a multi-agent framework, aiming to give the Actor complementary continuous and categorical feedback.
Training is split into offline learning from stored trajectories and online fine-tuning through interaction with the environment. Offline batches sample trajectory segments, clear recurrent model state between segments, and process each segment in chronological order to preserve sequence context. The document explains these design choices but provides only a partial view of the training code and no detailed quantitative evaluation. It reports that testing supported the approach’s viability, while giving no metrics or comparative results in the excerpt. The authors describe the programs as demonstrations and recommend representative data and comprehensive validation before live use.
Key ideas
- The Director adds binary strategy-fit feedback to the Actor–Critic system’s reward-based evaluation.
- Offline training uses sampled trajectory segments to establish an initial policy before online adaptation.
- Recurrent model state is cleared between unrelated batches to prevent stale context from affecting learning.
- States within each historical segment are processed in order to preserve temporal information.
- The reported tests are qualitative, and the authors recommend further validation before live trading.
Tags
This summary was written by Stratmill's research agent from the original; it is not a copy of the source.