Using Monte Carlo Sampling and OOB Error to Select RL Trading Agents
Summary
The article extends a reinforcement learning trading system based on random decision forests. It addresses overfitting by selecting price increment features and tuning regularization against out-of-bag classification error, then uses Monte Carlo sampling in the optimizer to generate multiple agents with different settings. The best candidate is saved for later use, and the system tracks model errors and trading rewards.
The reported example shows training losses substantially lower than out-of-bag losses, and the author describes test performance as unstable despite profitability over a limited period at a chosen signal threshold. All agents use the same closing-price data, so the experiment does not establish the value of adding different inputs. The article recommends considering both trade count and out-of-bag error when selecting a model, with similar training and test errors as a desired sign. These results are an illustrative, limited backtest rather than evidence of robust live performance.
Key ideas
- The system creates multiple reinforcement learning agents by randomly sampling model settings during optimization.
- Recursive feature elimination compares price increments as candidate predictors using out-of-bag error.
- Regularization and out-of-bag model selection are used to reduce overfitting.
- The example reports a gap between training and out-of-bag errors and unstable test behavior.
- Model choice should consider trade count and classification error alongside the test results.
Tags
This summary was written by Stratmill's research agent from the original; it is not a copy of the source.