When Reinforcement Learning Helps with Historical Portfolio Optimization
Summary
The discussion questions whether reinforcement learning is useful when a portfolio model is trained on fixed historical prices. In that setup, market states advance independently of the chosen action, and next-period rewards can be deterministic given the state and action. If actions affect only immediate reward, this weakens the usual reason to use reinforcement learning: optimizing a sequence of decisions for long-term outcomes rather than choosing greedily.
The reply argues that the result depends on how the state and reward are defined. Different action sequences can produce different terminal wealth or Sharpe outcomes, even with fixed prices, so a sequential optimization objective may still justify reinforcement learning. It also suggests bootstrapping historical data to add randomness. These points are conceptual rather than supported by an empirical comparison; the cited papers are not described. The exchange does not establish when such methods outperform simpler optimization, and notes that ignoring trading costs can make the problem closer to greedy decision-making.
Key ideas
- With fixed historical prices, actions may not affect the next market state.
- Reinforcement learning is most motivated when decisions influence future rewards or portfolio outcomes.
- Different sequences of portfolio actions can still produce different terminal performance on the same price history.
- Bootstrapping historical data is suggested as a way to add randomness to training.
- The exchange provides no direct evidence that reinforcement learning outperforms simpler methods.
Tags
Full text
# Why use reinforcement learning for portfolio optimization with historical market data?
# Why use reinforcement learning for portfolio optimization with historical market data?
One of the main advantages of (deep) reinforcement learning approaches (compared to more widely known supervised deep learning approaches) is the fact that it enables us to automatically take sequentiality into account. It's clear that the optimal action at time $t$ doesn't necessarily have to be the one that maximizes the expectation of immediate reward (greedy is not necessarily optimal in the long run). Therefore the framework of (D)RL seems appropriate for portfolio optimization where we are interested in maximizing a certain objective (say Sharpe's ratio) over a longer period.
However, many papers that deal with (D)RL applications in portfolio optimization use historical market data to build a deterministic MDP to train the model on. Under such approaches the state at time $t$ is a often list of historical returns for a chosen set of assets. Such an MDP deterministic in the sense that 1) the state at $t+1$ ($s_{t+1}$) will not depend on action at $t$ ($a_t$) since it consists of (already fixed) historical data 2) the reward at $t+1$ ($r_{t+1}$) will be a deterministic function of the previous action and state ($a_t, s_t$)
Therefore, the optimal action at $t$ will be greedy (since whatever we do it will not affect the next state) and the advantages of the DRL approach seem to be gone while its disadvantages (sample inefficiency, instability, etc.) are still there.
My question is the following: Why should one even try to use (deep) reinforcement learning for portfolio optimization when given historical market data (i.e. deterministic MDP to train on)?
Addendum: It's clear to me that we can artificially introduce sequentiality by say including current portfolio weights in the state vector (for the sake of say taking into account trading costs), but it still seems to me that a) under small trading costs the optimal action will be close to greedy, thereby still not achieving full sequentiality as desirable for RL based approaches b) many researchers who use DRL approaches totally ignore trading costs and don't seem to be bothered by what I've outlined above
## Answer by Aaron (score 1)
https://quant.stackexchange.com/a/54475
Whether you will end up with a deterministic or stochastic MDP fully depends on what you treat as the state and/or reward. Even working with fixed historical price data, different actions (sequence) will lead to totally different terminal wealth/Sharpe ratio, making it legitimate, or even favorable, to apply RL. You can even adopt bootstrapping type of training with fixed real data to enhance the randomness in the RL problem. Two good reference papers can be found here:
https://arxiv.org/abs/1904.11392
https://arxiv.org/abs/1907.11718
Hope this helps!Shown in full with attribution under the source's licence. Licence: CC BY-SA 4.0 (Stack Exchange)
This summary was written by Stratmill's research agent from the original; it is not a copy of the source.