Reinforcement Learning Methods for Portfolio Allocation
Summary
The responses survey reinforcement learning (RL) applications in quantitative finance, with portfolio allocation as the main example. They describe critic-only methods, which choose actions using learned value estimates; actor-only methods, which optimize actions more directly; and actor-critic methods, which combine an action policy with a learned evaluator. The discussion also distinguishes value-based Q-learning, which derives a policy from an optimized value function, from policy-based methods that directly optimize cumulative rewards.
The material mentions applications to option pricing, investment management, inverse reinforcement learning for trading, and portfolio optimization, and points readers to books, papers, and implementation resources. It offers no original experiments or performance evidence. It cautions indirectly that portfolio strategies may appear to outperform classical allocation approaches when transaction costs are omitted and reallocation is unrealistically frequent. The references are suggestions for further study, not a detailed implementation guide.
Key ideas
- Reinforcement learning in finance is commonly applied to portfolio allocation and investment management.
- Critic-only methods select actions from learned value estimates, while actor-only methods optimize actions more directly.
- Actor-critic methods pair an action policy with a critic that evaluates its choices.
- Q-learning optimizes a value function to derive a policy, while policy-based approaches optimize the policy directly.
- Reported portfolio advantages can be misleading when transaction costs and trading frequency are not modeled.
Tags
Full text
# Reinforcement learning in finance
# Reinforcement learning in finance
In brief, what are some mainstream and recent applications of reinforcement learning in finance that fall outside of the usual scope of agent-based modeling?
## Answer by Dhruv Mahajan (score 5)
https://quant.stackexchange.com/a/59208
Really recommend this book for RL in finance :
- Dixon et al (2020) Machine Learning in Finance: From Theory to Practice
He talks about QLBS, q-learning setup for black scholes, RL for investment management and inverse RL for trading.
## Answer by Igor Rivin (score 3)
https://quant.stackexchange.com/a/59196
See Alexandr Honchar's post on portfolio optimization with RL: https://medium.com/swlh/ai-for-portfolio-management-from-markowitz-to-reinforcement-learning-cffedcbba566
## Answer by Giogre (score 2)
https://quant.stackexchange.com/a/79655
I second the previous answer by Igor Rivin. In quantitative finance contexts Reinforcement Learning (RL) is chiefly employed to derive automated Portfolio Allocation Strategies showing superior performances (especially when the analyses are theoretical, and do not consider transaction costs, freely reallocating way too frequently) with respect to classical techniques such as Markowitz' MVO, Black-Litterman, Minimum Variance.
See this review paper from 2018, in which RL algorithms are classified in three distinct camps:
- Critic-only approach: select a value function to maximize that basically cares as much for exploration than exploitation when called upon to decide for action. Choice of action happens after having detected the current state of the system/environment (the market). The decision will be based on past rewards;
- Actor-only approach: more greedy, prefers exploitation;
- Actor-critic approach: "The key idea is to simultaneously use an actor, which determines the agent’s action given the current state of the environment, and a critic, which judges the selected action. Simply speaking, the actor learns to choose the action which is considered best by the critic and the critic learns to improve its judgment."
On the topic of Q-learning, it pertains to what kind of object function the RL agent intends to optimize. It is implemented with Neural Networks ("Deep RL"); the whole concept of it is linked to the Bellman equations, for which I would refer readers to this clear and informative video lecture. Q-learning consists in a family of algorithm that emphasize the optimization of the value function first and foremost, and from this then they derive an optimal policy. The opposite trend is that of policy-based methods, which directly optimize the objective function for the policy (usually cumulative rewards) using numerical techniques such as Gradient Descent-based algorithms.
Here is an available report, good to get an idea for a Python and a C++ implementation of a RL-driven portfolio allocation strategy.Shown in full with attribution under the source's licence. Licence: CC BY-SA 4.0 (Stack Exchange)
This summary was written by Stratmill's research agent from the original; it is not a copy of the source.