LSTM قیمت ماڈلز اور PPO کے ساتھ بٹ کوائن ٹریڈنگ
خلاصہ
مطالعہ بٹ کوائن کی خودکار ہائی فریکوئنسی ٹریڈنگ کا ایک خاکہ پیش کرتا ہے، جو لانگ شارٹ ٹرم میموری (LSTM) قیمت ماڈل کو پروکسمل پالیسی آپٹیمائزیشن (PPO) کے ساتھ ملاتا ہے۔ اس میں قیمتوں کو ایجنٹ کی حالتیں، ٹریڈز کو اعمال اور منافع کو انعامات کے طور پر ظاہر کیا گیا ہے۔ قیمت کی پیش گوئی کے جزو کے لیے یہ سپورٹ ویکٹر مشینز، ملٹی لیئر پرسیپٹرونز، LSTMs، ٹیمپورل کنولوشنل نیٹ ورکس اور ٹرانسفارمرز کا موازنہ کرتا ہے؛ بیان کردہ تجربات میں LSTM کو بہتر پایا گیا۔
مصنفین نتیجے کی PPO اسٹریٹیجی کو ہم آہنگ ڈیٹا کے ساتھ فرضی ماحول میں جانچتے اور ایک ہی اثاثے کے عام اسٹریٹیجی معیارات سے اس کا موازنہ کرتے ہیں۔ وہ سب سے مضبوط معیار سے زیادہ منافع اور اتار چڑھاؤ اور چڑھتی قیمتوں دونوں کے ادوار میں فائدے کی اطلاع دیتے ہیں۔ یہ نتائج مطالعے کے بٹ کوائن اور فرضی ماحول تک محدود ہیں؛ اقتباس میں لین دین کی لاگت، مارکیٹ اثر، نمونے کے ڈیزائن یا نمونے سے باہر مضبوطی کی تفصیل نہیں۔ لہٰذا دعوے یہ ثابت نہیں کرتے کہ یہ طریقہ لائیو ٹریڈنگ یا دوسری مصنوعات میں بھی کارآمد ہوگا۔
اہم خیالات
- یہ خاکہ قیمت کی حالتوں اور منافع کے انعامات سے بٹ کوائن ٹریڈز بنانے کے لیے PPO استعمال کرتا ہے۔
- کئی دوسرے پیش گوئی طریقوں سے موازنے کے بعد LSTM پالیسی کے قیمت ماڈل کی بنیاد فراہم کرتا ہے۔
- مطالعہ ہم آہنگ ڈیٹا کے ساتھ فرضی ماحول میں اسٹریٹیجی کا جائزہ لیتا ہے۔
- بیان کردہ معیاری موازنے آزمودہ واحد اثاثے کی صورت میں مجوزہ طریقے کے حق میں ہیں۔
- اقتباس میں ٹریڈنگ لاگت یا لائیو مارکیٹ میں مضبوطی کے شواہد بیان نہیں کیے گئے۔
ٹیگز
مکمل متن
# Bitcoin Transaction Strategy Construction Based on Deep Reinforcement Learning # Bitcoin Transaction Strategy Construction Based on Deep Reinforcement Learning The emerging cryptocurrency market has lately received great attention for asset allocation due to its decentralization uniqueness. However, its volatility and brand new trading mode have made it challenging to devising an acceptable automatically-generating strategy. This study proposes a framework for automatic high-frequency bitcoin transactions based on a deep reinforcement learning algorithm-proximal policy optimization (PPO). The framework creatively regards the transaction process as actions, returns as awards and prices as states to align with the idea of reinforcement learning. It compares advanced machine learning-based models for static price predictions including support vector machine (SVM), multi-layer perceptron (MLP), long short-term memory (LSTM), temporal convolutional network (TCN), and Transformer by applying them to the real-time bitcoin price and the experimental results demonstrate that LSTM outperforms. Then an automatically-generating transaction strategy is constructed building on PPO with LSTM as the basis to construct the policy. Extensive empirical studies validate that the proposed method performs superiorly to various common trading strategy benchmarks for a single financial product. The approach is able to trade bitcoins in a simulated environment with synchronous data and obtains a 31.67% more return than that of the best benchmark, improving the benchmark by 12.75%. The proposed framework can earn excess returns through both the period of volatility and surge, which opens the door to research on building a single cryptocurrency trading strategy based on deep learning. Visualizations of trading the process show how the model handles high-frequency transactions to provide inspiration and demonstrate that it can be expanded to other financial products.
ماخذ کا حوالہ دیتے ہوئے مکمل متن دکھایا گیا ہے، ماخذ کے لائسنس کے تحت۔ لائسنس: abstract CC0
یہ خلاصہ اصل ماخذ سے Stratmill کے تحقیقی ایجنٹ نے لکھا ہے؛ یہ ماخذ کی نقل نہیں۔