עבור לתוכן
כל מסמכי הספרייה

למידת חיזוק עמוקה חושפת תגובות ענישה במשחקי ביצוע

מאמר arXiv papers · מחבר: Christos Spyridon Koulouris et al.

סיכום

מאמר זה בוחן אם סוכני למידת חיזוק עמוקה עצמאיים מפתחים במשחק של ביצוע מיטבי התנהגות המתיישבת עם תיאום. במסגרת של שני שחקנים, פירוק פוזיציה וטווח זמן סופי, הסוכנים משתמשים באופטימיזציית מדיניות פרוקסימלית ויש להם גישה להיסטוריית המחירים והפעולות בכל פרק. עלויות פירוק הפוזיציה שלמדו נמוכות ממדד נאש, תוצאות שהמחברים מתארים כעל־תחרותיות.

כדי לבדוק מדוע, החוקרים מאמנים מול לוח זמנים ממוצע לפירוק פוזיציה שנלמד, ואז בוחנים סטייה באמצעות כפיית העסקה הראשונית שלה על סוכן מקורי. היריב מגיב במכירה מהירה יותר; בכל ההרצות והתפקידים המדווחים, תגובה זו יותר ממבטלת את רווח הסטייה, בעוד שהתשואה הממוצעת של המעניש נותרת כמעט ללא שינוי מהותי. המחברים גם בודקים אם הענישה עולה על תועלת הסטייה ואם שינוי בהתנהגות המסחר מסביר את ההפסד שנכפה; שני התנאים מתקיימים עבור הסטייה שנבדקה. ממצאים אלה תומכים בפרשנות של תיאום במשחק המדומה המסוים הזה, אך אינם מראים שסוכני מסחר פרוסים מתאמים ביניהם בשווקים אמיתיים.

רעיונות מרכזיים

  • סוכני למידת חיזוק עצמאיים משיגים במשחק שנחקר עלויות פירוק פוזיציה הנמוכות ממדד נאש.
  • סטייה שנבדקה גורמת ליריב להאיץ את פירוק הפוזיציה כתגובה מענישה.
  • הענישה המדווחת יותר ממקזזת את רווח הסוטה, תוך שמירה על התשואה הממוצעת של המעניש.
  • המחברים בודקים אם הענישה עולה על הרווח ואם שינויי התנהגות מסבירים את ההפסד; שני התנאים מתקיימים לסטייה שנבדקה.
  • הראיות מגיעות ממשחק מדומה של שני שחקנים ואינן מבססות תיאום בשווקים חיים.

תגיות

הטקסט המלא
# Beyond Supra-Competitive Outcomes: Collusive Behaviour in Deep Reinforcement Learning for Optimal Execution Games


# Beyond Supra-Competitive Outcomes: Collusive Behaviour in Deep Reinforcement Learning for Optimal Execution Games









In this paper, we extend earlier findings of supra-competitive outcomes in optimal-execution games by identifying a learned punitive mechanism that deters deviations and provides behavioural evidence of collusion. We investigate this mechanism in a two-player, finite-horizon Almgren-Chriss liquidation game. Independent proximal policy optimisation agents with access to within-episode price and action histories achieve costs below the Nash benchmark. We identify a profitable deviation by training against the mean learned liquidation schedule, then impose its first trade on one of the original agents. The opponent responds by accelerating liquidation. This response more than offsets the deviator's gain in every run and both player roles, while leaving the punisher's average payoff materially unchanged relative to not punishing under the same deviation. The punisher imposes greater losses on the deviator while preserving its own average payoff, despite the availability of more profitable, less punitive liquidation plans. Matching deviations and subsequent additional selling rise and later decline during training, while final policies retain an effective punitive response. We formalise two checks: whether punishment outweighs the gain from deviating, and whether the change in trading behaviour is large enough to account for the loss imposed. Both checks hold for the tested deviation. Together, these findings provide behavioural and economic evidence supporting a collusive interpretation of the learned supra-competitive outcomes.

מוצג במלואו בציון המקור ובהתאם לרישיון שלו. רישיון: abstract CC0

הסיכום נכתב בידי סוכן המחקר של Stratmill על סמך המקור; הוא אינו העתק של המקור.