가상자산 주문 실행의 전문가 혼합 모형: 안정성과 꼬리 위험
기사 arXiv papers · 저자: Alexander Ardaiz et al.
요약
이 연구는 BTC/USDT 주문 실행을 대상으로 이중 심층 Q 학습과 K-평균으로 구획한 전문가 혼합 모형, 매개변수 수를 맞춘 밀집 신경망을 비교합니다. 평가는 5분 평균 집계 바이낸스 지정가 주문 장부 데이터를 사용하며, 집행 부족분, 학습 시드별 변동, 정책 붕괴, 정책 내 꼬리 위험에 초점을 맞춥니다. 보상과 페널티 설계 때문에 조기 청산 비용이 거의 없는 재생 환경에서 시간가중 평균 가격 및 즉시 청산 기준과 결과를 비교합니다.
학습된 구성 중 기준 학습기보다 평균 부족분을 유의하게 개선한 것은 없으며, 이 사양에서는 모두 두 단순 기준보다 평균 부족분이 높습니다. 반복 시드 및 절제 분석은 붕괴 방지가 내재적인 전문가 혼합 효과라기보다 점감 탐색과 관련 있음을 보여줍니다. 탐색과 보상을 함께 변경하면 붕괴가 다시 나타날 수 있습니다. 가장 큰 전문가 구성은 시드 간 분산이 가장 낮지만 다중 비교 보정에는 강건하지 않으며, 전문가 수가 늘수록 정책 내 꼬리 위험은 악화됩니다. 이 결과는 시험한 환경과 사양에 한정되며, 시드 수에 따라 원인 해석이 달라지는 점은 평가의 불확실성을 보여줍니다.
핵심 아이디어
- 시험한 전문가 혼합 구성은 기본 이중 심층 Q 학습보다 평균 집행 부족분을 유의하게 낮추지 못합니다.
- 지정된 재생 환경에서는 시간가중 평균 가격과 즉시 청산의 평균 집행 부족분이 학습된 구성보다 낮습니다.
- 점감 탐색은 전문가 구획 없이 관측된 정책 붕괴를 억제합니다.
- 전문가 수가 늘 때 시드 간 분산과 정책 내 꼬리 위험은 서로 다르게 변합니다.
- 학습 시드 수에 따라 실패 원인 해석이 달라질 수 있으므로 반복 시드 평가가 중요합니다.
태그
전문
# Mixture-of-Experts for Cryptocurrency Order Execution: Training Stability, Tail Risk, and Failure Modes
# Mixture-of-Experts for Cryptocurrency Order Execution: Training Stability, Tail Risk, and Failure Modes
Deep reinforcement-learning policies for order execution can vary substantially across training seeds, so apparent architectural gains may reflect favourable training realisations rather than reproducible properties of the architecture. We evaluate vanilla Double Deep Q-Learning (DDQL), K-means-partitioned mixtures of DDQL experts at $K \in \{2, 4, 8\}$, and dense networks parameter-matched to the $K{=}4$ and $K{=}8$ expert budgets on 5-minute mean-aggregated BTC/USDT limit order book data from Binance. No learned configuration significantly improves mean implementation shortfall over DDQL. Under the reported specification, all have higher mean shortfall than TWAP (0.39 bps) and immediate liquidation (0.21 bps) in an environment whose frictionless replay and terminal-urgency penalty make early liquidation nearly costless; 11/100 vanilla-DDQL runs, versus none in either MoE $K{\geq}4$ arm, converge to a policy that waits until forced liquidation. We then decompose this specification on a device-matched baseline. Annealed exploration alone eliminates observed collapses (12/100 to 0/100; exact McNemar $p{=}4.9{\times}10^{-4}$), matching the elimination under expert partitioning. Combining annealed exploration with the aligned reward restores collapse in 19/30 runs; with all three specification changes, it rises to 48/100. In this environment, expert partitioning is unnecessary to suppress collapse and appears to mask a training-specification failure rather than confer an intrinsic performance benefit. No MoE $K{=}8$ run collapses under any of the six specifications tested. Across-seed dispersion is lowest at $K{=}8$ but non-monotone and not robust to family-wise adjustment, while within-policy tail risk worsens monotonically with $K$. The apparent attribution of the failure mode reverses between 30 and 100 seeds, illustrating the importance of repeated-seed evaluation.출처의 라이선스에 따라 출처를 표시하고 전문을 공개합니다. 라이선스: abstract CC0
이 요약은 원문을 바탕으로 Stratmill의 리서치 에이전트가 작성했으며, 원문을 복사한 것이 아닙니다.