コンテンツへスキップ
ライブラリの全資料

PPOと解析的なブローカー売買方策の比較

記事 arXiv papers · 著者: Siu Tung Wong (Institute of Finance and Technology et al.

サマリー

この研究は、連続時間のブローカー・トレーダーゲームにおいて、近接方策最適化がブローカーの売買速度を学習できるかを検証します。このゲームには解析解があります。連続時間の目的関数から離散報酬を導き、格子細分化と一段階の恒等式を用いて実装を確認したうえで、確率的な非情報注文フローの有無や部分情報の条件下で、学習済み方策を解析的ベンチマークと比較します。

非情報フローがない場合、フィードフォワード型のPPO方策は参照行動に近づきます。確率的なフローがある場合、教師あり学習ではアクターが目標を表現できると示されるものの、フィードフォワード型とリカレント型のPPO方策は正確さを保てません。モンテカルロ診断では、近接した行動の順位づけに苦戦するクリティックが示され、報酬整形も安定した改善につながりません。履歴に基づく確実性等価コントローラーは、部分情報の下でベンチマークにより近い性能を示します。PPOは固定された解析方策を、変化した執行コストに適応させますが、報告された改善で新たな基準との差が埋まるのはわずかです。結果は解析的ベンチマークの有用性と、検証したRL手法の限界の双方を示しています。

主なアイデア

  • 解析的に解ける売買ゲームが、強化学習方策を評価する直接的なベンチマークになります。
  • 確率的な非情報注文フローがない場合、PPOは参照行動に近づきますが、検証した確率的条件では正確さを欠きます。
  • アクターが目標を表現できても、クリティックが行動の順位を不正確に評価すれば、正確な判断は保証されません。
  • 観測可能な履歴に基づく因果的コントローラーは、部分情報の下で基準方策により近い動作を保ちます。
  • 解析方策を初期値としてPPOを開始すると、執行コストの変化にある程度適応できます。

タグ

全文
# When a Correct Reward Is Not Enough: Diagnosing and Guiding PPO in an Analytically Solved Broker-Trader Game


# When a Correct Reward Is Not Enough: Diagnosing and Guiding PPO in an Analytically Solved Broker-Trader Game









Reinforcement learning (RL) is increasingly used for financial optimal-control problems when complex dynamics make analytical strategies difficult to obtain. There are financial mathematics literactures which provides many solved models whose equations and controls could evaluate and guide learning; we ask whether RL can exploit these results. We place a proximal policy optimisation (PPO) agent in an analytically solved continuous-time broker--trader game. PPO replaces the broker and chooses its trading speed while interacting with an informed trader and stochastic uninformed order flow. We derive a finite-step reward from the broker's continuous-time payoff and verify its discrete implementation through grid refinement and an exact one-step identity. With zero uninformed flow, a validation-selected PPO--FFNN approaches the reference action. With stochastic uninformed flow, the tested PPO--FFNN and PPO--LSTM remain inaccurate, although supervised learning confirms that their actors can represent the action. Monte Carlo diagnostics show that their critics do not reliably rank nearby actions; potential-based reward shaping also gives no reliable improvement. Under partial information, a causal certainty-equivalent controller based on the broker's observable history remains close to the reference, while PPO has larger errors and lower payoffs. Finally, we freeze the analytical policy and train PPO to adjust it after the execution cost changes. Halving the cost yields a repeatable improvement that closes \(2.22\%\) of the gap to the changed-cost reference. The analytical solution therefore provides both a benchmark for diagnosing RL and a useful starting policy for adaptation.

出典を明記したうえで、ライセンスに従って全文を掲載しています。 ライセンス: abstract CC0

この要約は原文をもとにStratmillのリサーチエージェントが作成したもので、出典の複製ではありません。