Learning Liquidation Policies for Markets with Closing Auctions
Summary
This study considers a trader who sells through a limit order book and then adjusts remaining inventory using signed schedules in a closing auction. The two stages are linked by a forecast of the auction clearing price and feedback during the auction. The authors compare deep Q-networks with projected DDPG, TD3, and SAC reinforcement learning policies in a simulated market, using synthetic rough Heston prices and historical midprice paths.
Policies are chosen and assessed using inventory-penalized implementation shortfall, with evaluation kept separate from a weighted training objective. The learned approaches achieve lower penalized shortfall than the Avellaneda–Stoikov and TWAP reference strategies in the reported simulations. Matched synthetic comparisons also indicate that auction access helps all four learning methods. Denser auction feedback improves learning, while the additional decision value of the clearing-price forecast varies by learner. These findings are simulation-based; the description does not establish performance in live trading or generalization to other market conditions.
Key ideas
- The liquidation process combines continuous limit-order trading with a closing auction.
- Auction schedules adjust the inventory remaining after continuous trading.
- Four reinforcement learning approaches are compared using simulated prices and historical midprice paths.
- The learned policies report lower inventory-penalized shortfall than Avellaneda–Stoikov and TWAP references.
- Auction access helps each learner in matched synthetic comparisons, while forecast value depends on the learner.
Tags
Full text
# Learning Optimal Liquidation with Closing Auctions # Learning Optimal Liquidation with Closing Auctions We study liquidation when continuous trading is followed by a closing auction. The trader first sells through a limit-order book, then submits signed auction schedules to adjust the remaining inventory. A projected clearing price signal and intermediate auction feedback connect the two phases. We compare deep Q-network (DQN) policies with projected deep deterministic policy gradient (DDPG), twin delayed deep deterministic policy gradient (TD3) and soft actor-critic (SAC) policies, using synthetic rough Heston prices and historical midprice paths within a simulated market. Policies are selected and evaluated by inventory-penalized implementation shortfall, separately from a weighted training objective. They achieve lower inventory-penalized shortfall than the stylized Avellaneda-Stoikov (AS) and time-weighted average price (TWAP) references, and matched synthetic comparisons show that auction access is useful for all four learners. We furthermore find that dense auction credit improves learning; the clearing forecast is informative, but its incremental decision value is learner-dependent.
Shown in full with attribution under the source's licence. Licence: abstract CC0
This summary was written by Stratmill's research agent from the original; it is not a copy of the source.