تشخيص PPO بمقارنته بسياسة تداول تحليلية للوسيط
الملخص
تختبر هذه الدراسة ما إذا كان تحسين السياسة القريب يستطيع تعلم قرارات وسيط بشأن سرعة التداول في لعبة مستمرة الزمن بين وسيط ومتداول لها حل تحليلي. وتشتق مكافأة متقطعة من الهدف المستمر زمنيًا، وتتحقق من التنفيذ عبر تنقيح الشبكة وهوية الخطوة الواحدة، ثم تقارن السياسات المتعلمة بالمعيار التحليلي في إعدادات تشمل تدفق أوامر عشوائيًا من متداولين غير مطلعين ومعلومات جزئية وأخرى لا تشملها.
في غياب التدفق غير المطلع، تقترب سياسة PPO ذات التغذية الأمامية من الفعل المرجعي. ومع وجود تدفق عشوائي، تظل سياسات PPO ذات التغذية الأمامية والمتكررة غير دقيقة، رغم أن التعلم الخاضع للإشراف يشير إلى قدرة الفاعلين على تمثيل الهدف. وتشير تشخيصات مونت كارلو إلى أن النقاد يواجهون صعوبة في ترتيب الأفعال المتقاربة، ولا يساعد تشكيل المكافأة بصورة موثوقة. ويؤدي متحكم مكافئ اليقين المستند إلى التاريخ أداءً أقرب إلى المعيار في ظل المعلومات الجزئية. ويكيّف PPO سياسة تحليلية مجمدة مع تكلفة تنفيذ متغيرة، لكن المكسب المُبلغ عنه لا يسد سوى جزء صغير من الفجوة إلى المرجع الجديد. وتوضح النتائج قيمة المعايير التحليلية وحدود أساليب RL المختبرة.
الأفكار الرئيسية
- توفر لعبة تداول محلولة تحليليًا معيارًا مباشرًا لتقييم سياسة تعلم معزز.
- يقترب PPO من الفعل المرجعي عند غياب تدفق أوامر عشوائي غير مطلع، لكنه غير دقيق في الإعداد العشوائي المختبر.
- لا تضمن القدرة التمثيلية للفاعل دقة القرارات عندما يرتب الناقد الأفعال على نحو غير موثوق.
- يبقى متحكم سببي يستند إلى التاريخ المرصود أقرب إلى المرجع في ظل معلومات جزئية.
- يدعم بدء PPO من السياسة التحليلية بعض التكيف مع تغير تكاليف التنفيذ.
الوسوم
النص الكامل
# When a Correct Reward Is Not Enough: Diagnosing and Guiding PPO in an Analytically Solved Broker-Trader Game # When a Correct Reward Is Not Enough: Diagnosing and Guiding PPO in an Analytically Solved Broker-Trader Game Reinforcement learning (RL) is increasingly used for financial optimal-control problems when complex dynamics make analytical strategies difficult to obtain. There are financial mathematics literactures which provides many solved models whose equations and controls could evaluate and guide learning; we ask whether RL can exploit these results. We place a proximal policy optimisation (PPO) agent in an analytically solved continuous-time broker--trader game. PPO replaces the broker and chooses its trading speed while interacting with an informed trader and stochastic uninformed order flow. We derive a finite-step reward from the broker's continuous-time payoff and verify its discrete implementation through grid refinement and an exact one-step identity. With zero uninformed flow, a validation-selected PPO--FFNN approaches the reference action. With stochastic uninformed flow, the tested PPO--FFNN and PPO--LSTM remain inaccurate, although supervised learning confirms that their actors can represent the action. Monte Carlo diagnostics show that their critics do not reliably rank nearby actions; potential-based reward shaping also gives no reliable improvement. Under partial information, a causal certainty-equivalent controller based on the broker's observable history remains close to the reference, while PPO has larger errors and lower payoffs. Finally, we freeze the analytical policy and train PPO to adjust it after the execution cost changes. Halving the cost yields a repeatable improvement that closes \(2.22\%\) of the gap to the changed-cost reference. The analytical solution therefore provides both a benchmark for diagnosing RL and a useful starting policy for adaptation.
يُعرض النص كاملًا مع نسبه إلى مصدره وفقًا لترخيصه. الترخيص: abstract CC0
أعدّ وكيل الأبحاث في Stratmill هذا الملخص استنادًا إلى المصدر الأصلي؛ وهو ليس نسخة منه.