ارزیابی PPO در برابر سیاست تحلیلی معاملهگری کارگزار
خلاصه
این مطالعه میآزماید که آیا بهینهسازی سیاست مجاور میتواند تصمیمهای سرعت معامله کارگزار را در یک بازی پیوستهزمان کارگزار–معاملهگر با راهحل تحلیلی بیاموزد. پاداش گسستهای را از هدف پیوستهزمان استخراج میکند و پیادهسازی را با پالایش شبکه و یک همانی یکگامی بررسی میکند؛ سپس سیاستهای آموختهشده را در شرایط دارای جریان سفارش تصادفیِ معاملهگران ناآگاه و اطلاعات ناقص، و نیز بدون این عوامل، با معیار تحلیلی مقایسه میکند.
بدون جریان سفارش ناآگاه، سیاست PPO پیشخور به کنش مرجع نزدیک میشود. با جریان تصادفی، سیاستهای PPO پیشخور و بازگشتی همچنان دقیق نیستند، هرچند یادگیری نظارتشده نشان میدهد بازیگران آنها میتوانند هدف را بازنمایی کنند. عیبیابی مونتکارلو نشان میدهد منتقدها در رتبهبندی کنشهای نزدیک به هم مشکل دارند و شکلدهی پاداش هم بهطور قابلاتکا کمک نمیکند. کنترلگر معادلِ قطعی مبتنی بر تاریخچه، در شرایط اطلاعات ناقص به معیار نزدیکتر عمل میکند. PPO یک سیاست تحلیلیِ منجمد را با هزینه اجرای تغییرکرده سازگار میکند، اما سود گزارششده فقط بخش کوچکی از فاصله تا مرجع جدید را میبندد. نتایج هم ارزش معیارهای تحلیلی و هم محدودیت روشهای RL آزمودهشده را نشان میدهند.
ایدههای کلیدی
- یک بازی معاملاتی با راهحل تحلیلی، معیار مستقیمی برای ارزیابی سیاست یادگیری تقویتی فراهم میکند.
- PPO در نبود جریان تصادفی سفارشهای معاملهگران ناآگاه به کنش مرجع نزدیک میشود، اما در وضعیت تصادفیِ آزمودهشده دقیق نیست.
- توان بازنمایی بازیگر، وقتی منتقد کنشها را بهطور غیرقابلاتکا رتبهبندی میکند، تصمیمهای دقیق را تضمین نمیکند.
- یک کنترلگر علّی مبتنی بر تاریخچه مشاهدهپذیر، در شرایط اطلاعات ناقص به مرجع نزدیکتر میماند.
- آغاز PPO از سیاست تحلیلی، تا حدی سازگاری با هزینههای اجرایی تغییریافته را ممکن میکند.
برچسبها
متن کامل
# When a Correct Reward Is Not Enough: Diagnosing and Guiding PPO in an Analytically Solved Broker-Trader Game # When a Correct Reward Is Not Enough: Diagnosing and Guiding PPO in an Analytically Solved Broker-Trader Game Reinforcement learning (RL) is increasingly used for financial optimal-control problems when complex dynamics make analytical strategies difficult to obtain. There are financial mathematics literactures which provides many solved models whose equations and controls could evaluate and guide learning; we ask whether RL can exploit these results. We place a proximal policy optimisation (PPO) agent in an analytically solved continuous-time broker--trader game. PPO replaces the broker and chooses its trading speed while interacting with an informed trader and stochastic uninformed order flow. We derive a finite-step reward from the broker's continuous-time payoff and verify its discrete implementation through grid refinement and an exact one-step identity. With zero uninformed flow, a validation-selected PPO--FFNN approaches the reference action. With stochastic uninformed flow, the tested PPO--FFNN and PPO--LSTM remain inaccurate, although supervised learning confirms that their actors can represent the action. Monte Carlo diagnostics show that their critics do not reliably rank nearby actions; potential-based reward shaping also gives no reliable improvement. Under partial information, a causal certainty-equivalent controller based on the broker's observable history remains close to the reference, while PPO has larger errors and lower payoffs. Finally, we freeze the analytical policy and train PPO to adjust it after the execution cost changes. Halving the cost yields a repeatable improvement that closes \(2.22\%\) of the gap to the changed-cost reference. The analytical solution therefore provides both a benchmark for diagnosing RL and a useful starting policy for adaptation.
با ذکر منبع و مطابق مجوز اثر، بهطور کامل نمایش داده میشود. مجوز: abstract CC0
این خلاصه را عامل پژوهشی Stratmill بر پایه متن اصلی نوشته است؛ نسخهای از اثر منبع نیست.