الانتقال إلى المحتوى
جميع مستندات المكتبة

التعلم المعزز العميق لتداول أزواج العملات الرقمية ديناميكيًا

مقال arXiv papers · المؤلف: Damian Lebiedź et al.

الملخص

تختبر هذه الدراسة التعلم المعزز العميق بوصفه طبقة تنفيذ إضافية لتداول الأزواج في أسواق العملات الرقمية المتقلبة. يرشح إطارها أولًا الأزواج المرشحة ويرتبها، ثم يستخدم وكيلًا للتحسين بسياسة القرب مع طبقة للذاكرة طويلة وقصيرة الأمد لاتخاذ قرارات التنفيذ. ويقيّد الوكيل تصميمٌ بمخاطر ثابتة ومتوسط متكيف وضوابط مخاطر حتمية، بهدف الحد من مخاطر التباعد. ويستخدم التقييم بيانات العقود الآجلة باينانس USD-M بالساعة، ويقارن السياسة المتعلمة بخط أساس استدلالي.

ترجح النتائج المعلنة خارج العينة سياسة التعلم المعزز، ويشير اختبار إعادة المعاينة الكتلية الدائرية الساكنة إلى تفوق ذي دلالة إحصائية ومعدل حسب المخاطر عند مستوى 10 بالمئة. ولا تبلغ النتيجة عتبة الدلالة الأشد البالغة 5 بالمئة، وتعزو الدراسة ذلك إلى ارتفاع التباين الخاص بالأصول الرقمية. تقتصر الأدلة على البيانات والإعداد المختبرين؛ ولا يقدم الوصف مقادير الأداء أو افتراضات تكاليف المعاملات أو اختبارات عبر منصات وفترات أخرى. ويوفر التصميم الهجين وسيلة لتقييد قرارات التنفيذ المتعلمة، لكن متانة النهج على نطاق أوسع تظل سؤالًا مفتوحًا.

الأفكار الرئيسية

  • يجمع النظام المقترح بين الاختيار الإحصائي للأزواج وطبقة تنفيذ متعلمة.
  • يتخذ وكيل PPO مزود بـLSTM قرارات التنفيذ ضمن حدود مخاطر حتمية.
  • تقيّم الدراسة النهج باستخدام بيانات العقود الآجلة باينانس USD-M بالساعة ومقارنته بخط أساس استدلالي.
  • التفوق المعدل حسب المخاطر المعلن ذو دلالة عند مستوى 10 بالمئة، لكنه لا يبلغ مستوى 5 بالمئة.
  • تقتصر النتائج على السوق والبيانات وإعداد الاستراتيجية المختبرة.

الوسوم

النص الكامل
# Dynamic Multi-Pair Trading Strategy in Cryptocurrency Markets with Deep Reinforcement Learning


# Dynamic Multi-Pair Trading Strategy in Cryptocurrency Markets with Deep Reinforcement Learning









This study aims to determine whether the application of Deep Reinforcement Learning (DRL) as a specialized execution overlay can enhance pair trading in highly volatile cryptocurrency markets. Although classical implementations of the strategy have proven successful in traditional equities, they frequently exhibit rigidity and suffer from severe divergence risks when applied to high-variance environments. To address this need, this research introduces novel concepts. To construct a robust system, we developed a hierarchical "Filter-then-Rank" pair selection methodology and a proprietary "Fixed Risk, Adaptive Mean" execution model. The system employs a Proximal Policy Optimization (PPO) agent with a Long Short-Term Memory (LSTM) layer to govern execution decisions within strict deterministic risk management boundaries. Evaluated on 1-hour interval data from the Binance USD-M Futures market, the optimized RL policy achieved an out-of-sample performance that substantially outperformed the heuristic baseline. A stationary circular block bootstrap robustness check confirms that the agent's risk-adjusted outperformance is statistically significant at the 10 percent level. Although falling marginally short of the stricter 5 percent threshold, this result highlights the extreme idiosyncratic variance characteristic of digital assets. Ultimately, this thesis contributes to the quantitative finance literature by introducing a hybrid architecture that combines statistical arbitrage with DRL execution policies. Furthermore, it delivers a novel framework for safe reinforcement learning via deterministic shielding, proving that anchoring a neural policy to statistically robust boundaries successfully mitigates severe divergence risks.

يُعرض النص كاملًا مع نسبه إلى مصدره وفقًا لترخيصه. الترخيص: abstract CC0

أعدّ وكيل الأبحاث في Stratmill هذا الملخص استنادًا إلى المصدر الأصلي؛ وهو ليس نسخة منه.