コンテンツへスキップ
ライブラリの全資料

深層強化学習が執行ゲームで見いだす懲罰的な応答

記事 arXiv papers · 著者: Christos Spyridon Koulouris et al.

サマリー

本論文は、独立した深層強化学習エージェントが最適執行ゲームで談合と整合する行動を身につけるかを検証します。2人のプレイヤーによる有限期間の清算設定で、エージェントは近接方策最適化を用い、各エピソード内の価格と行動の履歴を利用できます。学習後の清算コストはナッシュ均衡の基準を下回り、著者らはこれを競争水準を超える結果と説明しています。

理由を調べるため、研究者らは学習済みの平均清算スケジュールを相手に訓練し、元のエージェントに初回取引を強制して逸脱を検証します。相手は売却を速めて応答します。報告された実行とプレイヤーの役割を通じて、この応答は逸脱者の利益を上回って打ち消し、懲罰側の平均利得は実質的に変わりません。著者らは、懲罰が逸脱の利益を上回るか、取引行動の変化が課された損失を説明するかも確認し、検証した逸脱についてはいずれも成立しました。これらの結果は、この特定のシミュレーションゲームにおける談合という解釈を支持しますが、実際の市場で運用される取引エージェントが談合することを示すものではありません。

主なアイデア

  • 独立した強化学習エージェントは、対象ゲームでナッシュ均衡の基準を下回る清算コストを実現します。
  • 検証した逸脱に対し、相手は清算を速める懲罰的な応答をします。
  • 報告された懲罰は逸脱者の利益を上回って打ち消し、懲罰側の平均利得を保ちます。
  • 著者らは懲罰が利益を上回るか、行動変化が損失を説明するかを検証し、対象の逸脱では両方を確認しています。
  • 証拠は2人のシミュレーションゲームに基づき、実市場での談合を立証するものではありません。

タグ

全文
# Beyond Supra-Competitive Outcomes: Collusive Behaviour in Deep Reinforcement Learning for Optimal Execution Games


# Beyond Supra-Competitive Outcomes: Collusive Behaviour in Deep Reinforcement Learning for Optimal Execution Games









In this paper, we extend earlier findings of supra-competitive outcomes in optimal-execution games by identifying a learned punitive mechanism that deters deviations and provides behavioural evidence of collusion. We investigate this mechanism in a two-player, finite-horizon Almgren-Chriss liquidation game. Independent proximal policy optimisation agents with access to within-episode price and action histories achieve costs below the Nash benchmark. We identify a profitable deviation by training against the mean learned liquidation schedule, then impose its first trade on one of the original agents. The opponent responds by accelerating liquidation. This response more than offsets the deviator's gain in every run and both player roles, while leaving the punisher's average payoff materially unchanged relative to not punishing under the same deviation. The punisher imposes greater losses on the deviator while preserving its own average payoff, despite the availability of more profitable, less punitive liquidation plans. Matching deviations and subsequent additional selling rise and later decline during training, while final policies retain an effective punitive response. We formalise two checks: whether punishment outweighs the gain from deviating, and whether the change in trading behaviour is large enough to account for the loss imposed. Both checks hold for the tested deviation. Together, these findings provide behavioural and economic evidence supporting a collusive interpretation of the learned supra-competitive outcomes.

出典を明記したうえで、ライセンスに従って全文を掲載しています。 ライセンス: abstract CC0

この要約は原文をもとにStratmillのリサーチエージェントが作成したもので、出典の複製ではありません。