クロージングオークションを用いた清算方針の学習
記事 arXiv papers · 著者: Julius Graf et al.
サマリー
本研究では、指値注文板で売却した後、クロージングオークションで売買方向を符号で示したスケジュールを用いて残存ポジションを調整するトレーダーを検討します。2つの段階は、オークションの清算価格予測とオークション中のフィードバックによって結び付いています。合成ラフ・ヘストン価格と過去の仲値経路を用いたシミュレーション市場で、深層Qネットワークと、射影付きDDPG、TD3、SACの強化学習方策を比較しています。
方策は在庫ペナルティ付きの実行ショートフォールを用いて選択・評価され、評価は重み付きの学習目的関数とは分けて行われています。報告されたシミュレーションでは、学習手法はAvellaneda–StoikovとTWAPの参照戦略よりペナルティ付きショートフォールが小さくなっています。条件をそろえた合成データの比較では、オークションへの参加が4つすべての学習手法に有利であることも示されています。オークション中のフィードバックを増やすと学習が改善しますが、約定価格予測による追加的な判断価値は学習手法によって異なります。これらの知見はシミュレーションに基づいており、実運用での性能や他の市場環境への一般化を示すものではありません。
主なアイデア
- 清算プロセスは、継続的な指値注文による取引とクロージングオークションを組み合わせます。
- オークションのスケジュールは、継続取引後に残るポジションを調整します。
- 4つの強化学習手法を、シミュレーション価格と過去の仲値経路を使って比較しています。
- 学習方策は、Avellaneda–StoikovとTWAPの参照戦略より在庫ペナルティ付きショートフォールが小さいと報告されています。
- 条件をそろえた合成データの比較では各学習手法にオークション参加の効果が見られますが、予測の価値は手法により異なります。
タグ
全文
# Learning Optimal Liquidation with Closing Auctions # Learning Optimal Liquidation with Closing Auctions We study liquidation when continuous trading is followed by a closing auction. The trader first sells through a limit-order book, then submits signed auction schedules to adjust the remaining inventory. A projected clearing price signal and intermediate auction feedback connect the two phases. We compare deep Q-network (DQN) policies with projected deep deterministic policy gradient (DDPG), twin delayed deep deterministic policy gradient (TD3) and soft actor-critic (SAC) policies, using synthetic rough Heston prices and historical midprice paths within a simulated market. Policies are selected and evaluated by inventory-penalized implementation shortfall, separately from a weighted training objective. They achieve lower inventory-penalized shortfall than the stylized Avellaneda-Stoikov (AS) and time-weighted average price (TWAP) references, and matched synthetic comparisons show that auction access is useful for all four learners. We furthermore find that dense auction credit improves learning; the clearing forecast is informative, but its incremental decision value is learner-dependent.
出典を明記したうえで、ライセンスに従って全文を掲載しています。 ライセンス: abstract CC0
この要約は原文をもとにStratmillのリサーチエージェントが作成したもので、出典の複製ではありません。