الانتقال إلى المحتوى
جميع مستندات المكتبة

مزيج الخبراء لتنفيذ أوامر العملات المشفرة: الاستقرار ومخاطر الذيل

مقال arXiv papers · المؤلف: Alexander Ardaiz et al.

الملخص

تقارن هذه الدراسة التعلم المزدوج العميق Q بمزيج خبراء مقسم بخوارزمية K-means وبشبكات كثيفة متطابقة المعلمات لتنفيذ أوامر BTC/USDT. ويستخدم التقييم بيانات دفتر أوامر محددة من بينانس مجمعة بمتوسطات لخمس دقائق، ويركز على عجز التنفيذ وتباين بذور التدريب وانهيار السياسة ومخاطر الذيل داخل السياسة. وتُقارن النتائج بمعياري متوسط السعر المرجح بالوقت والتصفية الفورية في بيئة إعادة تشغيل تكون فيها التصفية المبكرة شبه مجانية بسبب تصميم المكافأة والعقوبة المذكور.

لا يحسن أي إعداد متعلم متوسط عجز التنفيذ بدرجة دالة مقارنة بالمتعلم الأساسي، ولكل منها متوسط عجز أعلى من كلا المعيارين البسيطين في هذه المواصفات. وتعزو تحليلات البذور المتكررة والاستئصال منع الانهيار إلى الاستكشاف الملدّن، لا إلى فائدة جوهرية من مزيج الخبراء؛ وقد تؤدي التغييرات المجمعة في الاستكشاف والمكافأة إلى عودة الانهيار. ويظهر أكبر إعداد للخبراء أقل تشتت بين البذور، لكن ذلك لا يصمد أمام تصحيح المقارنات المتعددة، بينما تتفاقم مخاطر الذيل داخل السياسة مع زيادة عدد الخبراء. وتخص هذه النتائج البيئة والمواصفات المختبرة، كما يبرز تغير إسناد الأسباب باختلاف عدد البذور عدم اليقين في التقييم.

الأفكار الرئيسية

  • لا تخفض إعدادات مزيج الخبراء المختبرة متوسط عجز التنفيذ بدرجة دالة مقارنة بالتعلم المزدوج العميق Q الأساسي.
  • في بيئة إعادة التشغيل المحددة، يتفوق معيار متوسط السعر المرجح بالوقت والتصفية الفورية على الإعدادات المتعلمة في متوسط عجز التنفيذ.
  • يحد الاستكشاف الملدّن من انهيارات السياسة المرصودة من دون الحاجة إلى تقسيم الخبراء.
  • يتغير التشتت بين البذور ومخاطر الذيل داخل السياسة على نحو مختلف مع زيادة عدد الخبراء.
  • قد يتغير إسناد أسباب الإخفاق باختلاف عدد بذور التدريب، لذا يهم التقييم ببذور متكررة.

الوسوم

النص الكامل
# Mixture-of-Experts for Cryptocurrency Order Execution: Training Stability, Tail Risk, and Failure Modes


# Mixture-of-Experts for Cryptocurrency Order Execution: Training Stability, Tail Risk, and Failure Modes









Deep reinforcement-learning policies for order execution can vary substantially across training seeds, so apparent architectural gains may reflect favourable training realisations rather than reproducible properties of the architecture. We evaluate vanilla Double Deep Q-Learning (DDQL), K-means-partitioned mixtures of DDQL experts at $K \in \{2, 4, 8\}$, and dense networks parameter-matched to the $K{=}4$ and $K{=}8$ expert budgets on 5-minute mean-aggregated BTC/USDT limit order book data from Binance. No learned configuration significantly improves mean implementation shortfall over DDQL. Under the reported specification, all have higher mean shortfall than TWAP (0.39 bps) and immediate liquidation (0.21 bps) in an environment whose frictionless replay and terminal-urgency penalty make early liquidation nearly costless; 11/100 vanilla-DDQL runs, versus none in either MoE $K{\geq}4$ arm, converge to a policy that waits until forced liquidation. We then decompose this specification on a device-matched baseline. Annealed exploration alone eliminates observed collapses (12/100 to 0/100; exact McNemar $p{=}4.9{\times}10^{-4}$), matching the elimination under expert partitioning. Combining annealed exploration with the aligned reward restores collapse in 19/30 runs; with all three specification changes, it rises to 48/100. In this environment, expert partitioning is unnecessary to suppress collapse and appears to mask a training-specification failure rather than confer an intrinsic performance benefit. No MoE $K{=}8$ run collapses under any of the six specifications tested. Across-seed dispersion is lowest at $K{=}8$ but non-monotone and not robust to family-wise adjustment, while within-policy tail risk worsens monotonically with $K$. The apparent attribution of the failure mode reverses between 30 and 100 seeds, illustrating the importance of repeated-seed evaluation.

يُعرض النص كاملًا مع نسبه إلى مصدره وفقًا لترخيصه. الترخيص: abstract CC0

أعدّ وكيل الأبحاث في Stratmill هذا الملخص استنادًا إلى المصدر الأصلي؛ وهو ليس نسخة منه.