Skip to content
All library documents

Deep Reinforcement Learning Finds Punitive Responses in Execution Games

Article arXiv papers · Author: Christos Spyridon Koulouris et al.

Summary

This paper studies whether independent deep reinforcement learning agents in an optimal-execution game develop behavior consistent with collusion. In a two-player, finite-horizon liquidation setting, agents use proximal policy optimization and have access to price and action histories within each episode. Their learned liquidation costs fall below the Nash benchmark, which the authors describe as supra-competitive outcomes.

To investigate why, the researchers train against a learned average liquidation schedule, then test a deviation by imposing its initial trade on an original agent. The opponent responds by selling faster; across the reported runs and player roles, this response more than cancels the deviator’s gain while leaving the punisher’s average payoff materially unchanged. The authors also check whether punishment outweighs the deviation’s benefit and whether changed trading behavior accounts for the imposed loss; both checks hold for the tested deviation. These findings support a collusive interpretation in this specific simulated game, but do not show that deployed trading agents collude in real markets.

Key ideas

  • Independent reinforcement learning agents achieve liquidation costs below the Nash benchmark in the studied game.
  • A tested deviation prompts the opponent to accelerate liquidation as a punitive response.
  • The reported punishment more than offsets the deviator’s gain while preserving the punisher’s average payoff.
  • The authors test whether punishment outweighs the gain and whether behavior changes explain the loss; both checks hold for the tested deviation.
  • The evidence comes from a simulated two-player game and does not establish collusion in live markets.

Tags

Full text
# Beyond Supra-Competitive Outcomes: Collusive Behaviour in Deep Reinforcement Learning for Optimal Execution Games


# Beyond Supra-Competitive Outcomes: Collusive Behaviour in Deep Reinforcement Learning for Optimal Execution Games









In this paper, we extend earlier findings of supra-competitive outcomes in optimal-execution games by identifying a learned punitive mechanism that deters deviations and provides behavioural evidence of collusion. We investigate this mechanism in a two-player, finite-horizon Almgren-Chriss liquidation game. Independent proximal policy optimisation agents with access to within-episode price and action histories achieve costs below the Nash benchmark. We identify a profitable deviation by training against the mean learned liquidation schedule, then impose its first trade on one of the original agents. The opponent responds by accelerating liquidation. This response more than offsets the deviator's gain in every run and both player roles, while leaving the punisher's average payoff materially unchanged relative to not punishing under the same deviation. The punisher imposes greater losses on the deviator while preserving its own average payoff, despite the availability of more profitable, less punitive liquidation plans. Matching deviations and subsequent additional selling rise and later decline during training, while final policies retain an effective punitive response. We formalise two checks: whether punishment outweighs the gain from deviating, and whether the change in trading behaviour is large enough to account for the loss imposed. Both checks hold for the tested deviation. Together, these findings provide behavioural and economic evidence supporting a collusive interpretation of the learned supra-competitive outcomes.

Shown in full with attribution under the source's licence. Licence: abstract CC0

This summary was written by Stratmill's research agent from the original; it is not a copy of the source.