学习带收盘集合竞价的市场清算策略
文章 arXiv papers · 作者: Julius Graf et al.
总结
本研究考虑一名交易者先通过限价订单簿卖出,然后在收盘集合竞价中使用带方向的计划调整剩余库存。两个阶段通过对竞价成交价的预测以及竞价期间的反馈相连接。作者在模拟市场中比较深度 Q 网络与投影 DDPG、TD3 和 SAC 强化学习策略,使用合成粗糙 Heston 价格和历史中间价路径。
策略选择和评估依据是纳入库存惩罚的执行偏差,且评估过程与加权训练目标分开。报告模拟结果显示,学习方法的惩罚后执行偏差低于 Avellaneda–Stoikov 和 TWAP 基准策略。匹配的合成数据比较还表明,集合竞价机会有助于全部四种学习方法。更密集的集合竞价反馈能改善学习,而成交价预测带来的额外决策价值因学习方法而异。这些发现基于模拟;说明并未证明它们在实盘交易中的表现,或能推广至其他市场状况。
核心观点
- 平仓过程结合了连续限价订单交易和收盘集合竞价。
- 集合竞价计划用于调整连续交易后的剩余持仓。
- 研究使用模拟价格和历史中间价路径比较四种强化学习方法。
- 据报告,学习策略的库存惩罚后执行偏差低于 Avellaneda–Stoikov 和 TWAP 基准策略。
- 在匹配的合成数据比较中,集合竞价机会有助于每种学习方法,但预测的价值取决于学习方法。
标签
全文
# Learning Optimal Liquidation with Closing Auctions # Learning Optimal Liquidation with Closing Auctions We study liquidation when continuous trading is followed by a closing auction. The trader first sells through a limit-order book, then submits signed auction schedules to adjust the remaining inventory. A projected clearing price signal and intermediate auction feedback connect the two phases. We compare deep Q-network (DQN) policies with projected deep deterministic policy gradient (DDPG), twin delayed deep deterministic policy gradient (TD3) and soft actor-critic (SAC) policies, using synthetic rough Heston prices and historical midprice paths within a simulated market. Policies are selected and evaluated by inventory-penalized implementation shortfall, separately from a weighted training objective. They achieve lower inventory-penalized shortfall than the stylized Avellaneda-Stoikov (AS) and time-weighted average price (TWAP) references, and matched synthetic comparisons show that auction access is useful for all four learners. We furthermore find that dense auction credit improves learning; the clearing forecast is informative, but its incremental decision value is learner-dependent.
在遵守原作品许可的前提下,附作者信息全文展示。 许可协议: abstract CC0
此摘要由 Stratmill 研究智能体根据原文撰写,并非原文副本。