Training a Bitcoin Trading Agent with Reinforcement Learning
Summary
This tutorial describes a Bitcoin trading environment for reinforcement learning using OpenAI Gym and a PPO agent from Stable Baselines. The environment exposes recent OHLCV data alongside account history, permits buy, sell, or hold actions at fractional sizes, applies trading commissions, and samples random segments of historical data for training. It also outlines a visualization of price, volume, net worth, and trades.
For evaluation, the author separates earlier data for training from a later time-series segment for testing, noting that random k-fold validation can leak future information. Initial training results looked unusually strong, but an environment bug changed the reward chart; after further adjustments, some agents still failed while others performed better. The author reports that training and testing omitted commissions, making real-money profitability doubtful, and that a trained agent initially went bankrupt on unseen data. Switching algorithms and rewarding changes in net worth improved the reported test outcome, but the evidence is limited to one dataset and an unfinished experiment. Hyperparameter optimization and more robust validation remain future work.
Key ideas
- The Gym environment uses recent market history and account state as observations for a reinforcement learning agent.
- Random training slices expose the agent to varied historical periods, while evaluation uses a later sequential segment.
- Time-series validation must avoid training on information from the future.
- Early results were affected by an environment bug, and performance varied substantially across agents.
- The reported experiments excluded commissions and do not establish reliable live profitability.
Tags
This summary was written by Stratmill's research agent from the original; it is not a copy of the source.