Skip to content
All library documents

Designing Reinforcement Learning Agents for Trading

Article Freqtrade docs

Summary

The document explains how FreqAI trains trading agents through reinforcement learning. An agent processes historical candles and chooses among actions such as entering or exiting long and short positions. A custom reward function scores its decisions, while state information such as current profit, position, and trade duration can inform training and live or dry operation. The guide contrasts this approach with supervised classifiers and regressors, which use labeled targets but may offer more robust predictions.

It describes how to connect agent actions to strategy signals, choose an environment with three, four, or five available actions, and configure training and model parameters. Users can customize the reward function and environment, with logging available to track events and metrics. The examples explain the framework's interfaces rather than demonstrate trading performance. The document warns that the supplied reward function is illustrative, agents can exploit poorly designed rewards, and the simplified training environment omits real strategy features such as custom exits, stop losses, and leverage controls. Training behavior therefore may not carry over directly to live trading or backtests.

Key ideas

  • A reinforcement learning agent chooses trading actions and learns from rewards assigned by its environment.
  • The reward function shapes agent behavior and should be designed carefully to avoid exploiting unintended shortcuts.
  • FreqAI environments offer different action sets, which determine how agent decisions map to strategy entries and exits.
  • Training occurs in a simplified environment that does not reproduce all live strategy rules or controls.
  • The provided reward example demonstrates functionality and is not presented as production-ready.

Tags

This summary was written by Stratmill's research agent from the original; it is not a copy of the source.