עבור לתוכן
כל מסמכי הספרייה

אבחון PPO מול מדיניות מסחר אנליטית של ברוקר

מאמר arXiv papers · מחבר: Siu Tung Wong (Institute of Finance and Technology et al.

סיכום

מחקר זה בוחן אם אופטימיזציית מדיניות קרובה (PPO) יכולה ללמוד החלטות של ברוקר לגבי קצב המסחר במשחק רציף בזמן בין ברוקר לסוחר, שלו פתרון אנליטי. הוא גוזר תגמול בדיד מהיעד הרציף בזמן ובודק את המימוש באמצעות עידון רשת וזהות בצעד אחד, ואז משווה בין מדיניות שנלמדו לבין אמת המידה האנליטית בתרחישים עם ובלי זרימת הוראות אקראית של משתתפים לא־מיודעים ועם מידע חלקי.

בהיעדר זרימה לא־מיודעת, מדיניות PPO של רשת הזנה קדימה מתקרבת לפעולת הייחוס. עם זרימה אקראית, מדיניות PPO של רשת הזנה קדימה ושל רשת חוזרת נותרות לא מדויקות, אף שלמידה מפוקחת מראה שהשחקנים שלהן מסוגלים לייצג את היעד. אבחוני מונטה קרלו מצביעים על מבקרים שמתקשים לדרג פעולות סמוכות, ועיצוב מחדש של התגמול אינו מסייע באופן מהימן. בקר המבוסס על היסטוריה ושקילות ודאית מתקרב יותר לאמת המידה תחת מידע חלקי. PPO מתאימה מדיניות אנליטית קפואה לעלות ביצוע ששונתה, אך השיפור המדווח סוגר רק חלק קטן מהפער לאמת המידה החדשה. התוצאות ממחישות הן את ערכן של אמות מידה אנליטיות והן את מגבלות שיטות RL שנבדקו.

רעיונות מרכזיים

  • משחק מסחר שנפתר אנליטית מספק אמת מידה ישירה להערכת מדיניות למידת חיזוק.
  • PPO מתקרבת לפעולת הייחוס ללא זרימת הוראות אקראית של משתתפים לא־מיודעים, אך אינה מדויקת בתרחיש האקראי שנבדק.
  • יכולת הייצוג של השחקן אינה מבטיחה החלטות מדויקות כשהמבקר מדרג פעולות באופן לא מהימן.
  • בקר סיבתי המבוסס על היסטוריה נצפית נשאר קרוב יותר לייחוס תחת מידע חלקי.
  • התחלה של PPO מהמדיניות האנליטית מאפשרת התאמה מסוימת לעלויות ביצוע שהשתנו.

תגיות

הטקסט המלא
# When a Correct Reward Is Not Enough: Diagnosing and Guiding PPO in an Analytically Solved Broker-Trader Game


# When a Correct Reward Is Not Enough: Diagnosing and Guiding PPO in an Analytically Solved Broker-Trader Game









Reinforcement learning (RL) is increasingly used for financial optimal-control problems when complex dynamics make analytical strategies difficult to obtain. There are financial mathematics literactures which provides many solved models whose equations and controls could evaluate and guide learning; we ask whether RL can exploit these results. We place a proximal policy optimisation (PPO) agent in an analytically solved continuous-time broker--trader game. PPO replaces the broker and chooses its trading speed while interacting with an informed trader and stochastic uninformed order flow. We derive a finite-step reward from the broker's continuous-time payoff and verify its discrete implementation through grid refinement and an exact one-step identity. With zero uninformed flow, a validation-selected PPO--FFNN approaches the reference action. With stochastic uninformed flow, the tested PPO--FFNN and PPO--LSTM remain inaccurate, although supervised learning confirms that their actors can represent the action. Monte Carlo diagnostics show that their critics do not reliably rank nearby actions; potential-based reward shaping also gives no reliable improvement. Under partial information, a causal certainty-equivalent controller based on the broker's observable history remains close to the reference, while PPO has larger errors and lower payoffs. Finally, we freeze the analytical policy and train PPO to adjust it after the execution cost changes. Halving the cost yields a repeatable improvement that closes \(2.22\%\) of the gap to the changed-cost reference. The analytical solution therefore provides both a benchmark for diagnosing RL and a useful starting policy for adaptation.

מוצג במלואו בציון המקור ובהתאם לרישיון שלו. רישיון: abstract CC0

הסיכום נכתב בידי סוכן המחקר של Stratmill על סמך המקור; הוא אינו העתק של המקור.