종가 경매 시장의 청산 정책 학습
기사 arXiv papers · 저자: Julius Graf et al.
요약
이 연구는 지정가 주문장으로 매도한 뒤 종가 경매에서 부호가 있는 일정에 따라 잔여 재고를 조정하는 트레이더를 살펴봅니다. 두 단계는 경매 청산 가격 예측과 경매 중 피드백으로 연결됩니다. 저자들은 시뮬레이션 시장에서 심층 Q 네트워크와 투영된 DDPG, TD3, SAC 강화학습 정책을 비교하며, 합성 rough Heston 가격과 과거 중간가격 경로를 사용합니다.
정책은 재고 페널티를 반영한 실행 부족분을 기준으로 선택하고 평가하며, 평가 과정은 가중 훈련 목적함수와 분리합니다. 보고된 시뮬레이션에서 학습된 접근법은 Avellaneda–Stoikov 및 TWAP 기준 전략보다 페널티 반영 부족분이 낮습니다. 조건을 맞춘 합성 비교에서도 경매 참여가 네 가지 학습 방법 모두에 도움이 되는 것으로 나타납니다. 경매 피드백을 더 촘촘히 제공하면 학습이 개선되며, 청산 가격 예측의 추가적인 의사결정 가치는 학습 방법에 따라 다릅니다. 이 결과는 시뮬레이션에 기반하며, 실거래 성과나 다른 시장 상황으로의 일반화를 입증하지 않습니다.
핵심 아이디어
- 청산 과정은 연속 지정가 주문 거래와 종가 경매를 결합합니다.
- 경매 일정은 연속 거래 후 남은 재고를 조정합니다.
- 시뮬레이션 가격과 과거 중간가격 경로를 사용해 네 가지 강화학습 접근법을 비교합니다.
- 학습된 정책은 Avellaneda–Stoikov 및 TWAP 기준 전략보다 재고 페널티 반영 부족분이 낮다고 보고합니다.
- 조건을 맞춘 합성 비교에서 경매 참여는 각 학습 방법에 도움이 되며, 예측의 가치는 학습 방법에 따라 다릅니다.
태그
전문
# Learning Optimal Liquidation with Closing Auctions # Learning Optimal Liquidation with Closing Auctions We study liquidation when continuous trading is followed by a closing auction. The trader first sells through a limit-order book, then submits signed auction schedules to adjust the remaining inventory. A projected clearing price signal and intermediate auction feedback connect the two phases. We compare deep Q-network (DQN) policies with projected deep deterministic policy gradient (DDPG), twin delayed deep deterministic policy gradient (TD3) and soft actor-critic (SAC) policies, using synthetic rough Heston prices and historical midprice paths within a simulated market. Policies are selected and evaluated by inventory-penalized implementation shortfall, separately from a weighted training objective. They achieve lower inventory-penalized shortfall than the stylized Avellaneda-Stoikov (AS) and time-weighted average price (TWAP) references, and matched synthetic comparisons show that auction access is useful for all four learners. We furthermore find that dense auction credit improves learning; the clearing forecast is informative, but its incremental decision value is learner-dependent.
출처의 라이선스에 따라 출처를 표시하고 전문을 공개합니다. 라이선스: abstract CC0
이 요약은 원문을 바탕으로 Stratmill의 리서치 에이전트가 작성했으며, 원문을 복사한 것이 아닙니다.