Giao dịch Bitcoin với mô hình giá LSTM và PPO
Tóm tắt
Nghiên cứu đề xuất một khung giao dịch Bitcoin tần suất cao tự động, kết hợp mô hình giá bộ nhớ dài-ngắn hạn (LSTM) với phương pháp tối ưu hóa chính sách lân cận (PPO). Mô hình biểu diễn giá là trạng thái của tác nhân, giao dịch là hành động và lợi nhuận là phần thưởng. Với thành phần dự báo giá, nghiên cứu so sánh máy vectơ hỗ trợ, perceptron nhiều lớp, LSTM, mạng tích chập thời gian và Transformer; các thí nghiệm được báo cáo nghiêng về LSTM.
Các tác giả đánh giá chiến lược PPO thu được trong môi trường mô phỏng với dữ liệu đồng bộ và so sánh với các chuẩn chiến lược phổ biến cho một tài sản. Họ báo cáo lợi nhuận cao hơn chuẩn mạnh nhất và mô tả mức tăng trong cả giai đoạn biến động lẫn đi lên. Kết quả chỉ giới hạn ở bối cảnh Bitcoin và mô phỏng của nghiên cứu; trích đoạn không nêu chi phí giao dịch, tác động thị trường, thiết kế mẫu hay độ vững ngoài mẫu. Vì vậy, các tuyên bố này không chứng minh phương pháp có thể áp dụng cho giao dịch thực hoặc sản phẩm khác.
Ý chính
- Khung này dùng PPO để tạo giao dịch Bitcoin từ trạng thái giá và phần thưởng lợi nhuận.
- Một LSTM làm cơ sở mô hình giá cho chính sách sau khi được so sánh với một số phương pháp dự báo khác.
- Nghiên cứu đánh giá chiến lược trong môi trường mô phỏng với dữ liệu đồng bộ.
- Các so sánh chuẩn được báo cáo nghiêng về phương pháp đề xuất trong bối cảnh thử nghiệm một tài sản.
- Trích đoạn không mô tả chi phí giao dịch hoặc bằng chứng về độ vững trong thị trường thực.
Thẻ
Toàn văn
# Bitcoin Transaction Strategy Construction Based on Deep Reinforcement Learning # Bitcoin Transaction Strategy Construction Based on Deep Reinforcement Learning The emerging cryptocurrency market has lately received great attention for asset allocation due to its decentralization uniqueness. However, its volatility and brand new trading mode have made it challenging to devising an acceptable automatically-generating strategy. This study proposes a framework for automatic high-frequency bitcoin transactions based on a deep reinforcement learning algorithm-proximal policy optimization (PPO). The framework creatively regards the transaction process as actions, returns as awards and prices as states to align with the idea of reinforcement learning. It compares advanced machine learning-based models for static price predictions including support vector machine (SVM), multi-layer perceptron (MLP), long short-term memory (LSTM), temporal convolutional network (TCN), and Transformer by applying them to the real-time bitcoin price and the experimental results demonstrate that LSTM outperforms. Then an automatically-generating transaction strategy is constructed building on PPO with LSTM as the basis to construct the policy. Extensive empirical studies validate that the proposed method performs superiorly to various common trading strategy benchmarks for a single financial product. The approach is able to trade bitcoins in a simulated environment with synchronous data and obtains a 31.67% more return than that of the best benchmark, improving the benchmark by 12.75%. The proposed framework can earn excess returns through both the period of volatility and surge, which opens the door to research on building a single cryptocurrency trading strategy based on deep learning. Visualizations of trading the process show how the model handles high-frequency transactions to provide inspiration and demonstrate that it can be expanded to other financial products.
Hiển thị toàn văn kèm ghi nguồn theo giấy phép của tài liệu gốc. Giấy phép: abstract CC0
Bản tóm tắt này do tác nhân nghiên cứu của Stratmill biên soạn từ tài liệu gốc; đây không phải bản sao của tài liệu.