الانتقال إلى المحتوى
جميع مستندات المكتبة

TD3 ذو الأفعال المستمرة لتداول الأسهم وبيتكوين

مقال arXiv papers · المؤلف: Naseh Majidi et al.

الملخص

تطبق هذه الدراسة أسلوب Twin-Delayed Deep Deterministic Policy Gradient (TD3)، وهو أسلوب تعلم معزز ذو أفعال مستمرة، على أسعار الإغلاق اليومية لسهم أمازون وبيتكوين. ويمثل فعل التداول كلاً من المركز الواجب اتخاذه وعدد الأسهم الواجب تداولها، موسعاً بذلك الأساليب التي تستخدم أفعالاً متقطعة. وتقارن الورقة الاستراتيجيات الناتجة بأساليب التحليل الفني وأساليب أخرى للتعلم المعزز واستراتيجيات عشوائية وحتمية.

يُقيّم الأداء باستخدام العائد ونسبة شارب. وتشير النتائج المبلّغ عنها إلى أن اختيار المركز وحجم التداول معاً قد يحسن الأداء وفق هذين المقياسين. ولا تقدم الوثيقة تفاصيل التقييم أو النتائج المحددة أو أدلة على المتانة عبر أصول أو فترات سوق أخرى. ولذلك تصف استنتاجاتها الأسواق وأساليب المقارنة المدروسة، ولا تثبت أن TD3 ذا الأفعال المستمرة سيحقق أداء أفضل في سياقات أخرى.

الأفكار الرئيسية

  • يُستخدم TD3 لتوليد أفعال تداول من أسعار الإغلاق اليومية لأمازون وبيتكوين.
  • تمثل مساحة الأفعال كلاً من مركز التداول وعدد الأسهم المتداولة.
  • تقارن الدراسة TD3 بأساليب التحليل الفني والتعلم المعزز والاستراتيجيات العشوائية والحتمية.
  • العائد ونسبة شارب هما مقياسا الأداء المذكوران.
  • ترجح المقارنة المبلّغ عنها الجمع بين اختيار المركز وتحديد حجم التداول ضمن الإعدادات التي اختبرتها الدراسة.

الوسوم

النص الكامل
# Algorithmic Trading Using Continuous Action Space Deep Reinforcement Learning


# Algorithmic Trading Using Continuous Action Space Deep Reinforcement Learning









Price movement prediction has always been one of the traders' concerns in financial market trading. In order to increase their profit, they can analyze the historical data and predict the price movement. The large size of the data and complex relations between them lead us to use algorithmic trading and artificial intelligence. This paper aims to offer an approach using Twin-Delayed DDPG (TD3) and the daily close price in order to achieve a trading strategy in the stock and cryptocurrency markets. Unlike previous studies using a discrete action space reinforcement learning algorithm, the TD3 is continuous, offering both position and the number of trading shares. Both the stock (Amazon) and cryptocurrency (Bitcoin) markets are addressed in this research to evaluate the performance of the proposed algorithm. The achieved strategy using the TD3 is compared with some algorithms using technical analysis, reinforcement learning, stochastic, and deterministic strategies through two standard metrics, Return and Sharpe ratio. The results indicate that employing both position and the number of trading shares can improve the performance of a trading system based on the mentioned metrics.

يُعرض النص كاملًا مع نسبه إلى مصدره وفقًا لترخيصه. الترخيص: abstract CC0

أعدّ وكيل الأبحاث في Stratmill هذا الملخص استنادًا إلى المصدر الأصلي؛ وهو ليس نسخة منه.