基于解析经纪商交易策略诊断PPO
文章 arXiv papers · 作者: Siu Tung Wong (Institute of Finance and Technology et al.
总结
这项研究检验近端策略优化能否在具有解析解的连续时间经纪商—交易者博弈中,学习经纪商的交易速度决策。研究从连续时间目标推导离散奖励,并通过网格细化和单步恒等式检查实现;随后在存在或不存在随机非信息订单流及部分信息的情形下,将学习到的策略与解析基准进行比较。
在没有非信息订单流时,前馈 PPO 策略接近参考动作。有随机订单流时,前馈和循环 PPO 策略仍不准确,尽管监督学习表明其策略网络能够表示目标。蒙特卡洛诊断显示,评论家网络难以对相近动作排序,而奖励塑形并未稳定地带来帮助。基于历史信息的确定性等价控制器在部分信息条件下更接近基准。PPO 将冻结的解析策略适配到变化后的执行成本,但报告的改进仅弥合了与新基准之间差距的一小部分。结果既说明了解析基准的价值,也展现了所测试 RL 方法的局限。
核心观点
- 具有解析解的交易博弈为评估强化学习策略提供了直接基准。
- 没有随机非信息订单流时,PPO 接近参考动作;在所测试的随机情形下则不准确。
- 即使策略网络具备表示能力,评论家网络对动作排序不可靠仍会导致决策不准确。
- 基于可观测历史的因果控制器在部分信息条件下更接近参考策略。
- 从解析策略出发运行 PPO,有助于其适应变化后的执行成本。
标签
全文
# When a Correct Reward Is Not Enough: Diagnosing and Guiding PPO in an Analytically Solved Broker-Trader Game # When a Correct Reward Is Not Enough: Diagnosing and Guiding PPO in an Analytically Solved Broker-Trader Game Reinforcement learning (RL) is increasingly used for financial optimal-control problems when complex dynamics make analytical strategies difficult to obtain. There are financial mathematics literactures which provides many solved models whose equations and controls could evaluate and guide learning; we ask whether RL can exploit these results. We place a proximal policy optimisation (PPO) agent in an analytically solved continuous-time broker--trader game. PPO replaces the broker and chooses its trading speed while interacting with an informed trader and stochastic uninformed order flow. We derive a finite-step reward from the broker's continuous-time payoff and verify its discrete implementation through grid refinement and an exact one-step identity. With zero uninformed flow, a validation-selected PPO--FFNN approaches the reference action. With stochastic uninformed flow, the tested PPO--FFNN and PPO--LSTM remain inaccurate, although supervised learning confirms that their actors can represent the action. Monte Carlo diagnostics show that their critics do not reliably rank nearby actions; potential-based reward shaping also gives no reliable improvement. Under partial information, a causal certainty-equivalent controller based on the broker's observable history remains close to the reference, while PPO has larger errors and lower payoffs. Finally, we freeze the analytical policy and train PPO to adjust it after the execution cost changes. Halving the cost yields a repeatable improvement that closes \(2.22\%\) of the gap to the changed-cost reference. The analytical solution therefore provides both a benchmark for diagnosing RL and a useful starting policy for adaptation.
在遵守原作品许可的前提下,附作者信息全文展示。 许可协议: abstract CC0
此摘要由 Stratmill 研究智能体根据原文撰写,并非原文副本。