الانتقال إلى المحتوى
جميع مستندات المكتبة

التعلم المعزز العميق يكشف استجابات عقابية في ألعاب التنفيذ

مقال arXiv papers · المؤلف: Christos Spyridon Koulouris et al.

الملخص

تدرس هذه الورقة ما إذا كان وكلاء مستقلون للتعلم المعزز العميق يطورون سلوكًا يتسق مع التواطؤ في لعبة للتنفيذ الأمثل. وفي سياق تصفية بين لاعبين اثنين وأفق زمني محدود، يستخدم الوكلاء تحسين السياسة القريبة، ويُتاح لهم الاطلاع على سجلات الأسعار والأفعال خلال كل جولة. وتنخفض تكاليف التصفية التي تعلموها إلى ما دون معيار ناش، وهي نتائج يصفها المؤلفون بأنها فوق تنافسية.

للتحقق من السبب، يدرّب الباحثون الوكلاء في مواجهة جدول متوسط للتصفية جرى تعلمه، ثم يختبرون انحرافًا بفرض الصفقة الأولية لذلك الجدول على وكيل أصلي. فيرد الخصم بالبيع بوتيرة أسرع؛ وعبر الجولات المبلّغ عنها وأدوار اللاعبين، تلغي هذه الاستجابة مكسب المنحرف بأكثر منه، مع بقاء متوسط عائد المعاقِب دون تغيير جوهري. ويفحص المؤلفون أيضًا ما إذا كان أثر العقوبة يفوق منفعة الانحراف، وما إذا كان تغير سلوك التداول يفسر الخسارة المفروضة؛ وقد تحقق الشرطان في الانحراف المختبر. وتدعم هذه النتائج تفسير التواطؤ في هذه اللعبة المحاكية المحددة، لكنها لا تُظهر أن وكلاء التداول المنشورين يتواطؤون في الأسواق الحقيقية.

الأفكار الرئيسية

  • تحقق وكلاء التعلم المعزز المستقلون تكاليف تصفية أدنى من معيار ناش في اللعبة المدروسة.
  • يدفع انحراف مختبر الخصم إلى تسريع التصفية استجابةً عقابية.
  • تفوق العقوبة المذكورة مكسب المنحرف، مع الحفاظ على متوسط عائد المعاقِب.
  • يختبر المؤلفون ما إذا كانت العقوبة تفوق المكسب وما إذا كان تغير السلوك يفسر الخسارة؛ وقد تحقق الشرطان في الانحراف المختبر.
  • تأتي الأدلة من لعبة محاكية بين لاعبين اثنين، ولا تثبت التواطؤ في الأسواق الفعلية.

الوسوم

النص الكامل
# Beyond Supra-Competitive Outcomes: Collusive Behaviour in Deep Reinforcement Learning for Optimal Execution Games


# Beyond Supra-Competitive Outcomes: Collusive Behaviour in Deep Reinforcement Learning for Optimal Execution Games









In this paper, we extend earlier findings of supra-competitive outcomes in optimal-execution games by identifying a learned punitive mechanism that deters deviations and provides behavioural evidence of collusion. We investigate this mechanism in a two-player, finite-horizon Almgren-Chriss liquidation game. Independent proximal policy optimisation agents with access to within-episode price and action histories achieve costs below the Nash benchmark. We identify a profitable deviation by training against the mean learned liquidation schedule, then impose its first trade on one of the original agents. The opponent responds by accelerating liquidation. This response more than offsets the deviator's gain in every run and both player roles, while leaving the punisher's average payoff materially unchanged relative to not punishing under the same deviation. The punisher imposes greater losses on the deviator while preserving its own average payoff, despite the availability of more profitable, less punitive liquidation plans. Matching deviations and subsequent additional selling rise and later decline during training, while final policies retain an effective punitive response. We formalise two checks: whether punishment outweighs the gain from deviating, and whether the change in trading behaviour is large enough to account for the loss imposed. Both checks hold for the tested deviation. Together, these findings provide behavioural and economic evidence supporting a collusive interpretation of the learned supra-competitive outcomes.

يُعرض النص كاملًا مع نسبه إلى مصدره وفقًا لترخيصه. الترخيص: abstract CC0

أعدّ وكيل الأبحاث في Stratmill هذا الملخص استنادًا إلى المصدر الأصلي؛ وهو ليس نسخة منه.