الانتقال إلى المحتوى
جميع مستندات المكتبة

تداول البيتكوين بنماذج الأسعار LSTM وPPO

مقال arXiv papers · المؤلف: Fengrui Liu et al.

الملخص

تقترح الدراسة إطارًا آليًا لتداول البيتكوين عالي التردد يجمع بين نموذج أسعار يستخدم الذاكرة طويلة وقصيرة الأمد (LSTM) والتحسين بسياسة المنطقة القريبة (PPO). ويمثل الأسعار حالات الوكيل، والتداولات أفعالًا، والعوائد مكافآت. وبالنسبة إلى مكوّن التنبؤ بالأسعار، تقارن الدراسة آلات متجهات الدعم، والمدركات متعددة الطبقات، وLSTMs، والشبكات الالتفافية الزمنية، والمحولات؛ وترجح التجارب المذكورة LSTM.

يقيّم المؤلفون استراتيجية PPO الناتجة في بيئة محاكاة باستخدام بيانات متزامنة، ويقارنونها بمعايير شائعة للاستراتيجيات على أصل واحد. ويذكرون عوائد أعلى من أقوى معيار، ومكاسب خلال فترات التقلب والارتفاع. تقتصر هذه النتائج على حالة البيتكوين والمحاكاة في الدراسة؛ ولا يحدد المقتطف تكاليف المعاملات أو أثر السوق أو تصميم العينة أو المتانة خارج العينة. لذلك لا تثبت الادعاءات أن النهج سينتقل إلى التداول الفعلي أو إلى منتجات أخرى.

الأفكار الرئيسية

  • يستخدم الإطار PPO لإنشاء تداولات البيتكوين انطلاقًا من حالات الأسعار ومكافآت العوائد.
  • يوفر LSTM أساس نموذج الأسعار للسياسة بعد مقارنته بعدة أساليب تنبؤ أخرى.
  • تقيّم الدراسة الاستراتيجية في بيئة محاكاة ببيانات متزامنة.
  • ترجح مقارنات المعايير المذكورة الطريقة المقترحة ضمن إعداد الاختبار لأصل واحد.
  • لا يصف المقتطف تكاليف التداول أو أدلة المتانة في الأسواق الفعلية.

الوسوم

النص الكامل
# Bitcoin Transaction Strategy Construction Based on Deep Reinforcement Learning


# Bitcoin Transaction Strategy Construction Based on Deep Reinforcement Learning









The emerging cryptocurrency market has lately received great attention for asset allocation due to its decentralization uniqueness. However, its volatility and brand new trading mode have made it challenging to devising an acceptable automatically-generating strategy. This study proposes a framework for automatic high-frequency bitcoin transactions based on a deep reinforcement learning algorithm-proximal policy optimization (PPO). The framework creatively regards the transaction process as actions, returns as awards and prices as states to align with the idea of reinforcement learning. It compares advanced machine learning-based models for static price predictions including support vector machine (SVM), multi-layer perceptron (MLP), long short-term memory (LSTM), temporal convolutional network (TCN), and Transformer by applying them to the real-time bitcoin price and the experimental results demonstrate that LSTM outperforms. Then an automatically-generating transaction strategy is constructed building on PPO with LSTM as the basis to construct the policy. Extensive empirical studies validate that the proposed method performs superiorly to various common trading strategy benchmarks for a single financial product. The approach is able to trade bitcoins in a simulated environment with synchronous data and obtains a 31.67% more return than that of the best benchmark, improving the benchmark by 12.75%. The proposed framework can earn excess returns through both the period of volatility and surge, which opens the door to research on building a single cryptocurrency trading strategy based on deep learning. Visualizations of trading the process show how the model handles high-frequency transactions to provide inspiration and demonstrate that it can be expanded to other financial products.

يُعرض النص كاملًا مع نسبه إلى مصدره وفقًا لترخيصه. الترخيص: abstract CC0

أعدّ وكيل الأبحاث في Stratmill هذا الملخص استنادًا إلى المصدر الأصلي؛ وهو ليس نسخة منه.