Skip to content
All library documents

Training a Bitcoin Trading Agent with LSTM-PPO Reinforcement Learning

Article FMZ digest · Author: 小草

Summary

This article builds a Bitcoin trading agent with a recurrent neural network and Proximal Policy Optimization. The policy selects among holding, buying, or selling, while a custom backtest environment returns rewards based on changes in account value relative to a passive starting position. The state includes normalized OHLCV changes and account holdings. Episodes use randomly sampled sections of historical hourly data, and the model is updated from observed actions, rewards, and subsequent states.

The author reports that training was difficult and volatile, with signs of overfitting in the training period. On a later test period, the model did not lose money in the reported run, but performance was described as modest; these results are specific to the example and do not establish robustness. The article also identifies simplifying assumptions in the environment, including limited handling of insufficient funds or inventory. It recommends return-based normalization to reduce the model’s tendency to memorize price levels.

Key ideas

  • The PPO agent learns a probability distribution over holding, buying, and selling actions from market and account state.
  • The custom environment rewards changes in relative account value and includes transaction commissions.
  • Randomly sampled training episodes expose the agent to different historical starting points.
  • The reported training showed volatility and possible overfitting, so the results do not demonstrate robust performance.
  • Return-based normalization is proposed to reduce memorization of absolute price levels.

Tags

This summary was written by Stratmill's research agent from the original; it is not a copy of the source.