Training Bitcoin Trading Actions with an LSTM-PPO Backtest Environment
Summary
This tutorial describes training an agent to choose hold, buy, or sell actions for Bitcoin using a recurrent Proximal Policy Optimization model. It pairs an LSTM policy and value function with a custom Gym-style backtest environment. The environment supplies normalized market features and account state, applies trading fees, updates holdings, and returns rewards based on changes in account value. The article also outlines policy-gradient concepts and reports training and evaluation observations from historical and later test data.
The reported results are mixed: training is difficult and returns fluctuate; the trained model shows periodic losses, appears to overfit, and does not learn to hold long-term positions. On the test period, relative performance is described as moderate with no loss, and the position pattern suggests buying after sharp declines and selling after rebounds. The author emphasizes that this is a limited demonstration, not evidence of a robust strategy. Data normalization, sampling, execution assumptions, and the small test setup constrain the conclusions; further work is needed to improve the model and test its generalization.
Key ideas
- The article uses an LSTM-based PPO agent to select Bitcoin trading actions directly from market and account state.
- A custom backtest environment supplies rewards from changes in account value and includes commission and position updates.
- The tutorial reports volatile training and signs of overfitting, including difficulty holding long-term positions.
- A later test period shows moderate relative performance, but the evidence is limited and does not establish a dependable trading edge.
- The author favors relative return normalization to reduce the model's tendency to memorize price levels.
Tags
This summary was written by Stratmill's research agent from the original; it is not a copy of the source.