تجزیاتی بروکر ٹریڈنگ پالیسی کے مقابلے میں PPO کی جانچ
خلاصہ
یہ مطالعہ جانچتا ہے کہ آیا پروگزمل پالیسی آپٹیمائزیشن مسلسل وقت کے بروکر–ٹریڈر کھیل میں بروکر کے ٹریڈنگ کی رفتار کے فیصلے سیکھ سکتی ہے، جس کا تجزیاتی حل موجود ہے۔ یہ مسلسل وقت کے مقصد سے ایک جداگانہ انعام اخذ کرتا، گرڈ کو باریک کر کے اور ایک قدم کی شناخت سے نفاذ کی جانچ کرتا، پھر اسٹوکاسٹک غیر باخبر آرڈر فلو اور جزوی معلومات کی موجودگی اور عدم موجودگی میں سیکھنے والی پالیسیوں کا تجزیاتی معیار سے موازنہ کرتا ہے۔
غیر باخبر فلو نہ ہونے پر فیڈ فارورڈ PPO پالیسی حوالہ جاتی عمل کے قریب پہنچتی ہے۔ اسٹوکاسٹک فلو میں فیڈ فارورڈ اور ریکرنٹ PPO پالیسیاں غیر دقیق رہتی ہیں، اگرچہ زیرِ نگرانی لرننگ اشارہ کرتی ہے کہ ان کے ایکٹر ہدف کی نمائندگی کر سکتے ہیں۔ مونٹی کارلو تشخیص سے پتا چلتا ہے کہ کریٹکس قریبی اعمال کی درجہ بندی میں دشواری کا سامنا کرتے ہیں، اور انعام کی تشکیل قابلِ اعتماد طور پر مدد نہیں کرتی۔ تاریخ پر مبنی سرٹینٹی ایکوئیولنٹ کنٹرولر جزوی معلومات میں معیار کے زیادہ قریب کارکردگی دکھاتا ہے۔ PPO ایک منجمد تجزیاتی پالیسی کو بدلی ہوئی عمل درآمدی لاگت کے مطابق ڈھالتا ہے، مگر رپورٹ شدہ بہتری نئے معیار تک کے فرق کا صرف ایک چھوٹا حصہ کم کرتی ہے۔ نتائج تجزیاتی معیارات کی اہمیت اور آزمودہ RL طریقوں کی حدود دونوں دکھاتے ہیں۔
اہم خیالات
- تجزیاتی طور پر حل شدہ ٹریڈنگ کھیل ری انفورسمنٹ لرننگ پالیسی جانچنے کے لیے براہِ راست معیار فراہم کرتا ہے۔
- اسٹوکاسٹک غیر باخبر آرڈر فلو کے بغیر PPO حوالہ جاتی عمل کے قریب پہنچتا ہے، مگر آزمودہ اسٹوکاسٹک حالت میں غیر دقیق ہے۔
- ایکٹر کی نمائندگی کی صلاحیت درست فیصلوں کی ضمانت نہیں دیتی، جب کریٹک اعمال کی ناقابلِ اعتماد درجہ بندی کرے۔
- مشاہدہ شدہ تاریخ پر مبنی سببی کنٹرولر جزوی معلومات میں حوالہ جاتی عمل کے قریب رہتا ہے۔
- تجزیاتی پالیسی سے PPO شروع کرنا عمل درآمد کی بدلی ہوئی لاگت سے کچھ مطابقت پیدا کرنے میں مدد دیتا ہے۔
ٹیگز
مکمل متن
# When a Correct Reward Is Not Enough: Diagnosing and Guiding PPO in an Analytically Solved Broker-Trader Game # When a Correct Reward Is Not Enough: Diagnosing and Guiding PPO in an Analytically Solved Broker-Trader Game Reinforcement learning (RL) is increasingly used for financial optimal-control problems when complex dynamics make analytical strategies difficult to obtain. There are financial mathematics literactures which provides many solved models whose equations and controls could evaluate and guide learning; we ask whether RL can exploit these results. We place a proximal policy optimisation (PPO) agent in an analytically solved continuous-time broker--trader game. PPO replaces the broker and chooses its trading speed while interacting with an informed trader and stochastic uninformed order flow. We derive a finite-step reward from the broker's continuous-time payoff and verify its discrete implementation through grid refinement and an exact one-step identity. With zero uninformed flow, a validation-selected PPO--FFNN approaches the reference action. With stochastic uninformed flow, the tested PPO--FFNN and PPO--LSTM remain inaccurate, although supervised learning confirms that their actors can represent the action. Monte Carlo diagnostics show that their critics do not reliably rank nearby actions; potential-based reward shaping also gives no reliable improvement. Under partial information, a causal certainty-equivalent controller based on the broker's observable history remains close to the reference, while PPO has larger errors and lower payoffs. Finally, we freeze the analytical policy and train PPO to adjust it after the execution cost changes. Halving the cost yields a repeatable improvement that closes \(2.22\%\) of the gap to the changed-cost reference. The analytical solution therefore provides both a benchmark for diagnosing RL and a useful starting policy for adaptation.
ماخذ کا حوالہ دیتے ہوئے مکمل متن دکھایا گیا ہے، ماخذ کے لائسنس کے تحت۔ لائسنس: abstract CC0
یہ خلاصہ اصل ماخذ سے Stratmill کے تحقیقی ایجنٹ نے لکھا ہے؛ یہ ماخذ کی نقل نہیں۔