Multi-Step and Off-Policy Methods in Reinforcement Learning
Summary
This page summarizes a lecture on multi-step and off-policy algorithms in reinforcement learning. It identifies the subject as methods that learn from sequences of rewards across multiple steps while handling data generated by a policy different from the one being evaluated or improved. The lecture is attributed to research scientist Hado van Hasselt.
The summary also notes that the lecture discusses techniques for reducing variance, a central challenge when learning from sampled returns. These ideas can inform reinforcement-learning approaches to trading, where an agent learns decisions from sequential market interaction, but the page does not describe a particular trading application. It provides no equations, algorithm comparisons, experiments, or results; the linked lecture or accompanying material is required to assess the methods in detail.
Key ideas
- The lecture covers multi-step reinforcement-learning algorithms.
- It also addresses learning off-policy from data generated by another policy.
- Variance-reduction techniques are part of the discussion.
- The page gives no trading-specific example or empirical results.
Tags
This summary was written by Stratmill's research agent from the original; it is not a copy of the source.