رفتن به محتوا
همه اسناد کتابخانه

ارزیابی PPO در برابر سیاست تحلیلی معامله‌گری کارگزار

مقاله arXiv papers · نویسنده: Siu Tung Wong (Institute of Finance and Technology et al.

خلاصه

این مطالعه می‌آزماید که آیا بهینه‌سازی سیاست مجاور می‌تواند تصمیم‌های سرعت معامله کارگزار را در یک بازی پیوسته‌زمان کارگزار–معامله‌گر با راه‌حل تحلیلی بیاموزد. پاداش گسسته‌ای را از هدف پیوسته‌زمان استخراج می‌کند و پیاده‌سازی را با پالایش شبکه و یک همانی یک‌گامی بررسی می‌کند؛ سپس سیاست‌های آموخته‌شده را در شرایط دارای جریان سفارش تصادفیِ معامله‌گران ناآگاه و اطلاعات ناقص، و نیز بدون این عوامل، با معیار تحلیلی مقایسه می‌کند.

بدون جریان سفارش ناآگاه، سیاست PPO پیش‌خور به کنش مرجع نزدیک می‌شود. با جریان تصادفی، سیاست‌های PPO پیش‌خور و بازگشتی همچنان دقیق نیستند، هرچند یادگیری نظارت‌شده نشان می‌دهد بازیگران آن‌ها می‌توانند هدف را بازنمایی کنند. عیب‌یابی مونت‌کارلو نشان می‌دهد منتقدها در رتبه‌بندی کنش‌های نزدیک به هم مشکل دارند و شکل‌دهی پاداش هم به‌طور قابل‌اتکا کمک نمی‌کند. کنترل‌گر معادلِ قطعی مبتنی بر تاریخچه، در شرایط اطلاعات ناقص به معیار نزدیک‌تر عمل می‌کند. PPO یک سیاست تحلیلیِ منجمد را با هزینه اجرای تغییرکرده سازگار می‌کند، اما سود گزارش‌شده فقط بخش کوچکی از فاصله تا مرجع جدید را می‌بندد. نتایج هم ارزش معیارهای تحلیلی و هم محدودیت روش‌های RL آزموده‌شده را نشان می‌دهند.

ایده‌های کلیدی

  • یک بازی معاملاتی با راه‌حل تحلیلی، معیار مستقیمی برای ارزیابی سیاست یادگیری تقویتی فراهم می‌کند.
  • PPO در نبود جریان تصادفی سفارش‌های معامله‌گران ناآگاه به کنش مرجع نزدیک می‌شود، اما در وضعیت تصادفیِ آزموده‌شده دقیق نیست.
  • توان بازنمایی بازیگر، وقتی منتقد کنش‌ها را به‌طور غیرقابل‌اتکا رتبه‌بندی می‌کند، تصمیم‌های دقیق را تضمین نمی‌کند.
  • یک کنترل‌گر علّی مبتنی بر تاریخچه مشاهده‌پذیر، در شرایط اطلاعات ناقص به مرجع نزدیک‌تر می‌ماند.
  • آغاز PPO از سیاست تحلیلی، تا حدی سازگاری با هزینه‌های اجرایی تغییر‌یافته را ممکن می‌کند.

برچسب‌ها

متن کامل
# When a Correct Reward Is Not Enough: Diagnosing and Guiding PPO in an Analytically Solved Broker-Trader Game


# When a Correct Reward Is Not Enough: Diagnosing and Guiding PPO in an Analytically Solved Broker-Trader Game









Reinforcement learning (RL) is increasingly used for financial optimal-control problems when complex dynamics make analytical strategies difficult to obtain. There are financial mathematics literactures which provides many solved models whose equations and controls could evaluate and guide learning; we ask whether RL can exploit these results. We place a proximal policy optimisation (PPO) agent in an analytically solved continuous-time broker--trader game. PPO replaces the broker and chooses its trading speed while interacting with an informed trader and stochastic uninformed order flow. We derive a finite-step reward from the broker's continuous-time payoff and verify its discrete implementation through grid refinement and an exact one-step identity. With zero uninformed flow, a validation-selected PPO--FFNN approaches the reference action. With stochastic uninformed flow, the tested PPO--FFNN and PPO--LSTM remain inaccurate, although supervised learning confirms that their actors can represent the action. Monte Carlo diagnostics show that their critics do not reliably rank nearby actions; potential-based reward shaping also gives no reliable improvement. Under partial information, a causal certainty-equivalent controller based on the broker's observable history remains close to the reference, while PPO has larger errors and lower payoffs. Finally, we freeze the analytical policy and train PPO to adjust it after the execution cost changes. Halving the cost yields a repeatable improvement that closes \(2.22\%\) of the gap to the changed-cost reference. The analytical solution therefore provides both a benchmark for diagnosing RL and a useful starting policy for adaptation.

با ذکر منبع و مطابق مجوز اثر، به‌طور کامل نمایش داده می‌شود. مجوز: abstract CC0

این خلاصه را عامل پژوهشی Stratmill بر پایه متن اصلی نوشته است؛ نسخه‌ای از اثر منبع نیست.