Skip to content
All library documents

Building and Testing a Bitcoin Trading Agent with Reinforcement Learning

Article FMZ digest · Author: 发明者量化-小小梦

Summary

This tutorial builds a Bitcoin trading environment for reinforcement learning, using historical OHLCV data, a rolling observation window, and actions that buy, sell, or hold at varying position sizes. It describes random data slices for training, serial traversal for evaluation, and scaling observations that include both market history and account information. The article also emphasizes limiting observations to information available at the time to reduce look-ahead bias, and it discusses transaction commissions and episode boundaries.

The reported experiments show why results need careful validation. Early training appeared extremely profitable, but a bug was found; after correction, some agents performed well while others failed. The initially strong agents then went bankrupt on unseen test data. Changing the algorithm and rewarding changes in net worth improved the reported test outcome, but profitable models were not consistently established, and some earlier results omitted commissions. The tutorial is a useful environment-design example, not evidence that the learned policy is robust or suitable for live trading.

Key ideas

  • The environment represents buy, sell, and hold decisions alongside a selectable trade size.
  • Observations combine recent market bars with account state and trade history.
  • Random training slices expose the agent to varied sequences, while serial evaluation tests unseen data chronologically.
  • A discovered environment bug and poor unseen-data results show the danger of trusting training rewards.
  • Commissions, reward design, and out-of-sample testing materially affect the reported performance.

Tags

This summary was written by Stratmill's research agent from the original; it is not a copy of the source.