실행 게임에서 심층 강화학습이 보이는 응징 반응
기사 arXiv papers · 저자: Christos Spyridon Koulouris et al.
요약
이 논문은 최적 체결 게임에서 독립적인 심층 강화학습 에이전트가 담합에 부합하는 행동을 학습하는지 연구합니다. 두 명의 플레이어가 제한된 기간에 포지션을 청산하는 설정에서 에이전트는 근위 정책 최적화를 사용하고 각 에피소드의 가격 및 행동 이력에 접근합니다. 학습된 청산 비용은 내시 기준치보다 낮으며, 저자들은 이를 경쟁을 제한하는 결과라고 설명합니다.
원인을 살펴보기 위해 연구자들은 학습된 평균 청산 일정에 맞서 훈련한 뒤, 원래 에이전트에 처음 거래를 강제로 적용해 이탈 행동을 시험합니다. 상대는 더 빨리 매도하며, 보고된 실행과 플레이어 역할 전반에서 이러한 반응은 이탈자의 이익을 상쇄하고도 남으면서 응징자의 평균 보상은 실질적으로 변하지 않습니다. 저자들은 응징이 이탈의 이익보다 큰지, 바뀐 매매 행동이 강제된 손실을 설명하는지도 점검하며, 테스트한 이탈 행동에서는 두 조건 모두 성립합니다. 이 결과는 특정 시뮬레이션 게임에서 담합으로 해석할 근거를 제공하지만, 실제 시장에 배치된 매매 에이전트가 담합한다는 것을 보여주지는 않습니다.
핵심 아이디어
- 연구된 게임에서 독립적인 강화학습 에이전트는 내시 기준치보다 낮은 청산 비용을 달성합니다.
- 테스트한 이탈 행동에 상대가 청산 속도를 높이는 응징으로 대응합니다.
- 보고된 응징은 이탈자의 이익을 상쇄하고도 남으며 응징자의 평균 보상은 유지합니다.
- 저자들은 응징이 이익보다 큰지, 행동 변화가 손실을 설명하는지 점검하며, 테스트한 이탈 행동에서는 두 조건 모두 성립합니다.
- 증거는 두 플레이어가 참여하는 시뮬레이션 게임에서 나온 것으로 실거래 시장의 담합을 입증하지 않습니다.
태그
전문
# Beyond Supra-Competitive Outcomes: Collusive Behaviour in Deep Reinforcement Learning for Optimal Execution Games # Beyond Supra-Competitive Outcomes: Collusive Behaviour in Deep Reinforcement Learning for Optimal Execution Games In this paper, we extend earlier findings of supra-competitive outcomes in optimal-execution games by identifying a learned punitive mechanism that deters deviations and provides behavioural evidence of collusion. We investigate this mechanism in a two-player, finite-horizon Almgren-Chriss liquidation game. Independent proximal policy optimisation agents with access to within-episode price and action histories achieve costs below the Nash benchmark. We identify a profitable deviation by training against the mean learned liquidation schedule, then impose its first trade on one of the original agents. The opponent responds by accelerating liquidation. This response more than offsets the deviator's gain in every run and both player roles, while leaving the punisher's average payoff materially unchanged relative to not punishing under the same deviation. The punisher imposes greater losses on the deviator while preserving its own average payoff, despite the availability of more profitable, less punitive liquidation plans. Matching deviations and subsequent additional selling rise and later decline during training, while final policies retain an effective punitive response. We formalise two checks: whether punishment outweighs the gain from deviating, and whether the change in trading behaviour is large enough to account for the loss imposed. Both checks hold for the tested deviation. Together, these findings provide behavioural and economic evidence supporting a collusive interpretation of the learned supra-competitive outcomes.
출처의 라이선스에 따라 출처를 표시하고 전문을 공개합니다. 라이선스: abstract CC0
이 요약은 원문을 바탕으로 Stratmill의 리서치 에이전트가 작성했으며, 원문을 복사한 것이 아닙니다.