深度强化学习在执行博弈中发现惩罚性回应
文章 arXiv papers · 作者: Christos Spyridon Koulouris et al.
总结
本文研究独立的深度强化学习智能体是否会在最优执行博弈中形成与合谋相符的行为。在双人有限期限的清算场景中,智能体使用近端策略优化,并可在每个回合中获取价格和行动历史。它们学得的清算成本低于纳什基准,作者将其描述为超竞争性结果。
为探究原因,研究人员让智能体针对学得的平均清算计划进行训练,随后将该计划的初始交易强加给一个原始智能体,以测试偏离行为。对手以加快卖出的方式回应;在报告的各次运行和双方角色中,这一回应对偏离者造成的损失超过其所得收益,同时惩罚者的平均收益基本不变。作者还检验了惩罚是否超过偏离带来的收益,以及交易行为的改变是否能解释施加的损失;在所测试的偏离中,这两项检验均成立。这些发现支持对这一特定模拟博弈作出合谋解释,但并未表明已部署的交易智能体会在真实市场中合谋。
核心观点
- 在所研究的博弈中,独立强化学习智能体实现了低于纳什基准的清算成本。
- 一项受测偏离引发对手加快清算,作为惩罚性回应。
- 报告的惩罚超过了偏离者的收益,同时保留了惩罚者的平均收益。
- 作者检验了惩罚是否超过收益,以及行为变化是否解释了损失;在所测试的偏离中,两项检验均成立。
- 证据来自双人模拟博弈,无法证明实盘市场中存在合谋。
标签
全文
# Beyond Supra-Competitive Outcomes: Collusive Behaviour in Deep Reinforcement Learning for Optimal Execution Games # Beyond Supra-Competitive Outcomes: Collusive Behaviour in Deep Reinforcement Learning for Optimal Execution Games In this paper, we extend earlier findings of supra-competitive outcomes in optimal-execution games by identifying a learned punitive mechanism that deters deviations and provides behavioural evidence of collusion. We investigate this mechanism in a two-player, finite-horizon Almgren-Chriss liquidation game. Independent proximal policy optimisation agents with access to within-episode price and action histories achieve costs below the Nash benchmark. We identify a profitable deviation by training against the mean learned liquidation schedule, then impose its first trade on one of the original agents. The opponent responds by accelerating liquidation. This response more than offsets the deviator's gain in every run and both player roles, while leaving the punisher's average payoff materially unchanged relative to not punishing under the same deviation. The punisher imposes greater losses on the deviator while preserving its own average payoff, despite the availability of more profitable, less punitive liquidation plans. Matching deviations and subsequent additional selling rise and later decline during training, while final policies retain an effective punitive response. We formalise two checks: whether punishment outweighs the gain from deviating, and whether the change in trading behaviour is large enough to account for the loss imposed. Both checks hold for the tested deviation. Together, these findings provide behavioural and economic evidence supporting a collusive interpretation of the learned supra-competitive outcomes.
在遵守原作品许可的前提下,附作者信息全文展示。 许可协议: abstract CC0
此摘要由 Stratmill 研究智能体根据原文撰写,并非原文副本。