Skip to content
All library documents

Diagnosing PPO Against an Analytic Broker Trading Policy

Article arXiv papers · Author: Siu Tung Wong (Institute of Finance and Technology et al.

Summary

This study tests whether proximal policy optimization can learn a broker’s trading-speed decisions in a continuous-time broker–trader game with an analytical solution. It derives a discrete reward from the continuous-time objective and checks the implementation using grid refinement and a one-step identity, then compares learned policies with the analytic benchmark across settings with and without stochastic uninformed order flow and partial information.

With no uninformed flow, a feedforward PPO policy approaches the reference action. With stochastic flow, feedforward and recurrent PPO policies remain inaccurate, even though supervised learning indicates their actors can represent the target. Monte Carlo diagnostics point to critics that struggle to rank nearby actions, and reward shaping does not reliably help. A history-based certainty-equivalent controller performs closer to the benchmark under partial information. PPO adapts a frozen analytic policy to a changed execution cost, but the reported gain closes only a small part of the gap to the new reference. The results illustrate both the value of analytical benchmarks and limits of the tested RL methods.

Key ideas

  • An analytically solved trading game provides a direct benchmark for evaluating a reinforcement-learning policy.
  • PPO approaches the reference action without stochastic uninformed order flow but is inaccurate in the tested stochastic setting.
  • Actor representational capacity does not guarantee accurate decisions when the critic ranks actions unreliably.
  • A causal controller based on observable history stays closer to the reference under partial information.
  • Starting PPO from the analytical policy supports some adaptation to changed execution costs.

Tags

Full text
# When a Correct Reward Is Not Enough: Diagnosing and Guiding PPO in an Analytically Solved Broker-Trader Game


# When a Correct Reward Is Not Enough: Diagnosing and Guiding PPO in an Analytically Solved Broker-Trader Game









Reinforcement learning (RL) is increasingly used for financial optimal-control problems when complex dynamics make analytical strategies difficult to obtain. There are financial mathematics literactures which provides many solved models whose equations and controls could evaluate and guide learning; we ask whether RL can exploit these results. We place a proximal policy optimisation (PPO) agent in an analytically solved continuous-time broker--trader game. PPO replaces the broker and chooses its trading speed while interacting with an informed trader and stochastic uninformed order flow. We derive a finite-step reward from the broker's continuous-time payoff and verify its discrete implementation through grid refinement and an exact one-step identity. With zero uninformed flow, a validation-selected PPO--FFNN approaches the reference action. With stochastic uninformed flow, the tested PPO--FFNN and PPO--LSTM remain inaccurate, although supervised learning confirms that their actors can represent the action. Monte Carlo diagnostics show that their critics do not reliably rank nearby actions; potential-based reward shaping also gives no reliable improvement. Under partial information, a causal certainty-equivalent controller based on the broker's observable history remains close to the reference, while PPO has larger errors and lower payoffs. Finally, we freeze the analytical policy and train PPO to adjust it after the execution cost changes. Halving the cost yields a repeatable improvement that closes \(2.22\%\) of the gap to the changed-cost reference. The analytical solution therefore provides both a benchmark for diagnosing RL and a useful starting policy for adaptation.

Shown in full with attribution under the source's licence. Licence: abstract CC0

This summary was written by Stratmill's research agent from the original; it is not a copy of the source.