Reinforcement Learning: From Dynamic Programming to Policy Optimization
Summary
The document gives a brief historical and conceptual overview of reinforcement learning. It starts with Markov decision processes and dynamic programming, then describes temporal-difference learning and Monte Carlo methods as ways to estimate value without knowing system dynamics in advance. It presents Q-learning as a value-iteration-like method that also avoids requiring those dynamics, and contrasts it with SARSA, which updates using actions selected under the current policy and is characterized as more conservative.
The overview then notes the combination of deep learning with Q-learning and the emergence of policy-based optimization. It distinguishes value-based approaches, which estimate action or state values, from policy-based methods, which evaluate and optimize a policy as a whole. This is a general introduction rather than a trading application: it gives no market examples, implementation details, empirical comparisons, or guidance on risks such as exploration, reward design, and out-of-sample performance.
Key ideas
- Reinforcement learning is commonly framed as learning in a Markov decision process.
- Dynamic programming can solve problems when system dynamics are known, while temporal-difference and Monte Carlo methods can learn without that prior knowledge.
- Q-learning learns action values without requiring a model of system dynamics.
- SARSA is on-policy because its update uses an action selected under the current policy.
- Deep Q-learning extends value-based learning with deep networks, while policy-based methods optimize policies directly.
Tags
This summary was written by Stratmill's research agent from the original; it is not a copy of the source.