본문으로 건너뛰기
라이브러리 문서 전체

분석적 브로커 매매 정책과 PPO 비교

기사 arXiv papers · 저자: Siu Tung Wong (Institute of Finance and Technology et al.

요약

이 연구는 분석적 해가 있는 연속시간 브로커–트레이더 게임에서 근위 정책 최적화가 브로커의 매매 속도 결정을 학습할 수 있는지 검증합니다. 연속시간 목적 함수에서 이산 보상을 도출하고 격자 세분화 및 1단계 항등식으로 구현을 점검한 다음, 확률적 비정보 주문 흐름과 부분 정보의 유무에 따라 학습된 정책을 분석적 기준과 비교합니다.

비정보 흐름이 없으면 피드포워드 PPO 정책은 기준 행동에 근접합니다. 확률적 흐름에서는 지도학습 결과 액터가 목표를 표현할 수 있는 것으로 나타나도 피드포워드 및 순환 PPO 정책은 부정확합니다. 몬테카를로 진단에 따르면 크리틱은 서로 가까운 행동의 순위를 매기는 데 어려움을 겪으며, 보상 셰이핑도 일관되게 도움이 되지 않습니다. 이력 기반 확실성 등가 제어기는 부분 정보 아래에서 기준에 더 가까운 성능을 보입니다. PPO는 고정된 분석적 정책을 달라진 실행 비용에 맞게 조정하지만, 보고된 개선은 새 기준과의 격차 중 작은 부분만 줄입니다. 결과는 분석적 기준의 가치와 시험한 RL 방법의 한계를 모두 보여줍니다.

핵심 아이디어

  • 해석적으로 해를 구한 매매 게임은 강화학습 정책을 평가하는 직접적인 기준을 제공합니다.
  • 확률적 비정보 주문 흐름이 없으면 PPO 정책은 기준 행동에 근접하지만, 시험한 확률적 환경에서는 부정확합니다.
  • 액터가 목표를 표현할 수 있어도 크리틱이 행동 순위를 신뢰성 있게 매기지 못하면 정확한 결정을 보장하지 않습니다.
  • 관측 가능한 이력에 기반한 인과적 제어기는 부분 정보에서 기준에 더 가까운 성능을 유지합니다.
  • 분석적 정책에서 PPO를 시작하면 달라진 실행 비용에 어느 정도 적응할 수 있습니다.

태그

전문
# When a Correct Reward Is Not Enough: Diagnosing and Guiding PPO in an Analytically Solved Broker-Trader Game


# When a Correct Reward Is Not Enough: Diagnosing and Guiding PPO in an Analytically Solved Broker-Trader Game









Reinforcement learning (RL) is increasingly used for financial optimal-control problems when complex dynamics make analytical strategies difficult to obtain. There are financial mathematics literactures which provides many solved models whose equations and controls could evaluate and guide learning; we ask whether RL can exploit these results. We place a proximal policy optimisation (PPO) agent in an analytically solved continuous-time broker--trader game. PPO replaces the broker and chooses its trading speed while interacting with an informed trader and stochastic uninformed order flow. We derive a finite-step reward from the broker's continuous-time payoff and verify its discrete implementation through grid refinement and an exact one-step identity. With zero uninformed flow, a validation-selected PPO--FFNN approaches the reference action. With stochastic uninformed flow, the tested PPO--FFNN and PPO--LSTM remain inaccurate, although supervised learning confirms that their actors can represent the action. Monte Carlo diagnostics show that their critics do not reliably rank nearby actions; potential-based reward shaping also gives no reliable improvement. Under partial information, a causal certainty-equivalent controller based on the broker's observable history remains close to the reference, while PPO has larger errors and lower payoffs. Finally, we freeze the analytical policy and train PPO to adjust it after the execution cost changes. Halving the cost yields a repeatable improvement that closes \(2.22\%\) of the gap to the changed-cost reference. The analytical solution therefore provides both a benchmark for diagnosing RL and a useful starting policy for adaptation.

출처의 라이선스에 따라 출처를 표시하고 전문을 공개합니다. 라이선스: abstract CC0

이 요약은 원문을 바탕으로 Stratmill의 리서치 에이전트가 작성했으며, 원문을 복사한 것이 아닙니다.