Skip to content
All library documents

Training and Backtesting Deep Reinforcement Learning Stock Agents

Article FinRL

Summary

The tutorial outlines an end-to-end FinRL workflow for training and comparing deep reinforcement learning agents on Dow 30 equities. It describes downloading and preprocessing market data, adding technical indicators plus VIX and a turbulence measure, dividing the history into training and trading periods, and fitting A2C, DDPG, PPO, TD3, and SAC agents. It also identifies training controls such as timesteps, learning rate, batch size, and replay-buffer capacity, explaining their intended roles and tradeoffs.

For evaluation, the tutorial runs trained agents on a later trading period and compares their performance with mean-variance optimization and the DJIA. It provides a reproducible sequence of steps, but no reported performance results or details about transaction costs, slippage, risk-adjusted metrics, or robustness checks. The stated trading window ends in March 2026, so this example is limited to that sample and should not be read as evidence that any agent is profitable or will generalize.

Key ideas

  • The workflow separates data preparation, agent training, and backtesting into distinct stages.
  • Five reinforcement learning algorithms are trained using the same stock trading setup.
  • Technical indicators, VIX, and a turbulence index are included in the prepared market data.
  • The backtest compares agents with mean-variance optimization and the DJIA, but gives no performance findings.

Tags

Full text
# FinRL Stock Trading 2026 Tutorial


## FinRL Stock Trading 2026 Tutorial

### Step 1: Clone the Repository

```bash
git clone https://github.com/AI4Finance-Foundation/FinRL.git
cd FinRL
```

### Step 2: Create and Activate Virtual Environment

```bash
python3 -m venv venv
source venv/bin/activate
```

### Step 3: Install FinRL

```bash
pip install -e .
```

### Step 4: Run the Scripts

**1. Data Download & Preprocessing**

```bash
python examples/FinRL_StockTrading_2026_1_data.py
```

This script downloads DOW 30 stock data from Yahoo Finance, adds technical indicators (MACD, RSI, etc.), VIX, and turbulence index, then splits the data into training set (2014–2025) and trading set (2026-01-01 to 2026-03-20), saving them as `train_data.csv` and `trade_data.csv`.

**2. Train DRL Agents**

```bash
python examples/FinRL_StockTrading_2026_2_train.py
```

This script trains 5 DRL agents (A2C, DDPG, PPO, TD3, SAC) using Stable Baselines 3 on the training data. Trained models are saved to the `trained_models/` directory.

**Key Hyperparameters:**

| Parameter | Description | Default in Script |
|-----------|-------------|-------------------|
| `total_timesteps` | Total number of environment interactions for training. **This is the most important parameter** — higher values give the agent more experience to learn from, leading to better trading performance. Start with a small value (e.g., 1,000) for a quick test, then increase (e.g., 20,000–200,000) for serious training. | 20,000 |
| `learning_rate` | Controls how much the model weights are updated at each step. Too high causes instability; too low causes slow learning. | 0.001 |
| `batch_size` | Number of samples used per gradient update. Larger batches give more stable updates but require more memory. | 100 |
| `buffer_size` | Size of the replay buffer (for off-policy algorithms like DDPG, TD3, SAC). Stores past experiences for the agent to learn from. Larger buffers retain more diverse experiences. | 1,000,000 |

**3. Backtest**

```bash
python examples/FinRL_StockTrading_2026_3_Backtest.py
```

This script loads the trained agents, runs them on the trading data, and compares their performance against two baselines: Mean Variance Optimization (MVO) and the DJIA index. Results are printed to the console and a plot is saved as `backtest_result.png`.

Shown in full with attribution under the source's licence. Licence: MIT

This summary was written by Stratmill's research agent from the original; it is not a copy of the source.