Skip to content
All library documents

Why Model-Free Reinforcement Learning Can Still Overfit in Trading

Article Quant Q&A · Author: Ali H. Askar

Summary

The document considers whether Soft Actor-Critic reinforcement learning can optimize the risk-aversion parameter in an Avellaneda–Stoikov market-making strategy and then be used in live trading. The answers explain that being model-free does not make an agent immune to overfitting. Generalization depends on how well its training environment represents the live market, including the effects of the agent's own orders and trades.

The responses identify market making as especially difficult to simulate because trading actions can influence market behavior. A poorly specified environment or state space, and the use of an off-the-shelf algorithm, can all undermine results. A realistic simulator is presented as a major challenge; live fine-tuning is mentioned as a possible later step, while beginning live training is described as costly. The discussion offers caution rather than empirical evidence or a tested deployment recipe, and does not establish that any particular agent will succeed out of sample.

Key ideas

  • Model-free reinforcement learning can still overfit to its training data or environment.
  • Live performance depends on whether the agent generalizes to market conditions it did not encounter during training.
  • Market-making simulations are difficult because the agent's actions can affect the market it is trying to model.
  • The environment and state space need to represent trading interactions realistically.
  • Live training may provide direct market experience, but the response notes that it can be costly.

Tags

Full text
# can Soft Actor-Critic reinforcement learning algorithms be used in real-time trading?


# can Soft Actor-Critic reinforcement learning algorithms be used in real-time trading?












I am scratching my head with an optimization problem for Avellaneda and Stoikov market-making algorithm (optimizing the risk aversion parameter), and I've come across https://github.com/im1235/ISAC

which is using SACs to optimize the gamma parameter.

since SAC is a model-free reinforcement learning, does this mean it is not prone to overfitting?

or in other words, can it be applied to live to trade?

## Answer by autoencoder (score 1)

https://quant.stackexchange.com/a/71485

Overfitting depends on whether your agent can generalize to the real world of trading, not on whether it is model free or not. When you train your agent with historical market data, or simulated data, you need to make sure the interactions with the market is as realistic as possible, which is hard especially in the case of market making, where your actions affect the market a lot. Personally, I think it very hard to make it work unless you have a really good market simulator. However, if it does work out-of-sample, you could start to run it alive and finetune it with real market data. I would prefer to train it directly by running the agent live in the very beginning, though it would be costly.

## Answer by IDontKnowCode (score 0)

https://quant.stackexchange.com/a/71514

I agree with autoencoder; your environment must really be precisely defined. Markets react to trades and creating accurate simulator for your environment is a challenging problem. I think overfitting will depend entirely on the definition of your environment and state space and can happen if you use out of the box algorithms. However, if such an accurately defined environment and state space does exist, it is not entirely impossible to create a successful RL application.

Shown in full with attribution under the source's licence. Licence: CC BY-SA 4.0 (Stack Exchange)

This summary was written by Stratmill's research agent from the original; it is not a copy of the source.