본문으로 건너뛰기
라이브러리 문서 전체

LSTM 가격 모형과 PPO을 활용한 비트코인 트레이딩

기사 arXiv papers · 저자: Fengrui Liu et al.

요약

이 연구는 장단기 메모리(LSTM) 가격 모형과 근접 정책 최적화(PPO)를 결합한 자동화 고빈도 비트코인 트레이딩 프레임워크를 제안합니다. 가격은 에이전트의 상태, 거래는 행동, 수익률은 보상으로 나타냅니다. 가격 예측 요소로 서포트 벡터 머신, 다층 퍼셉트론, LSTM, 시간 합성곱 네트워크, 트랜스포머를 비교했으며, 보고된 실험에서는 LSTM이 더 나은 결과를 보였습니다.

저자들은 동기화된 데이터를 사용해 시뮬레이션 환경에서 결과로 나온 PPO 전략을 평가하고, 단일 자산의 일반적인 전략 벤치마크와 비교합니다. 가장 강력한 벤치마크보다 높은 수익률을 보고하며, 변동성이 큰 기간과 상승 기간 모두에서 이익이 났다고 설명합니다. 이 결과는 연구의 비트코인 설정과 시뮬레이션에 한정됩니다. 발췌문은 거래 비용, 시장 충격, 표본 설계 또는 표본 외 견고성을 명시하지 않습니다. 따라서 이러한 주장은 해당 접근법이 실거래나 다른 상품에도 적용된다는 점을 입증하지 않습니다.

핵심 아이디어

  • 이 프레임워크는 PPO을 사용해 가격 상태와 수익 보상으로부터 비트코인 거래를 생성합니다.
  • 여러 예측 방법과 비교한 뒤 LSTM을 정책의 가격 모형 기반으로 사용합니다.
  • 이 연구는 동기화된 데이터를 사용하는 시뮬레이션 환경에서 전략을 평가합니다.
  • 보고된 벤치마크 비교에서는 테스트한 단일 자산 설정에서 제안된 방법이 더 나은 결과를 보였습니다.
  • 발췌문에는 트레이딩 비용이나 실거래 환경의 견고성을 보여주는 근거가 없습니다.

태그

전문
# Bitcoin Transaction Strategy Construction Based on Deep Reinforcement Learning


# Bitcoin Transaction Strategy Construction Based on Deep Reinforcement Learning









The emerging cryptocurrency market has lately received great attention for asset allocation due to its decentralization uniqueness. However, its volatility and brand new trading mode have made it challenging to devising an acceptable automatically-generating strategy. This study proposes a framework for automatic high-frequency bitcoin transactions based on a deep reinforcement learning algorithm-proximal policy optimization (PPO). The framework creatively regards the transaction process as actions, returns as awards and prices as states to align with the idea of reinforcement learning. It compares advanced machine learning-based models for static price predictions including support vector machine (SVM), multi-layer perceptron (MLP), long short-term memory (LSTM), temporal convolutional network (TCN), and Transformer by applying them to the real-time bitcoin price and the experimental results demonstrate that LSTM outperforms. Then an automatically-generating transaction strategy is constructed building on PPO with LSTM as the basis to construct the policy. Extensive empirical studies validate that the proposed method performs superiorly to various common trading strategy benchmarks for a single financial product. The approach is able to trade bitcoins in a simulated environment with synchronous data and obtains a 31.67% more return than that of the best benchmark, improving the benchmark by 12.75%. The proposed framework can earn excess returns through both the period of volatility and surge, which opens the door to research on building a single cryptocurrency trading strategy based on deep learning. Visualizations of trading the process show how the model handles high-frequency transactions to provide inspiration and demonstrate that it can be expanded to other financial products.

출처의 라이선스에 따라 출처를 표시하고 전문을 공개합니다. 라이선스: abstract CC0

이 요약은 원문을 바탕으로 Stratmill의 리서치 에이전트가 작성했으며, 원문을 복사한 것이 아닙니다.