Skip to content
All library documents

Exploration and Exploitation in Reinforcement Learning

Article BigQuant

Summary

This page introduces the exploration–exploitation problem in reinforcement learning: a learning agent must decide when to test actions that may reveal useful information and when to choose actions based on what it has learned so far. The lecture is attributed to researcher Hado van Hasselt and is described as part of a DeepMind and UCL course. The page provides links to a video and a PDF, but their contents are not included in the supplied text.

The excerpt gives no specific algorithm, trading application, empirical comparison, or performance evidence. Its value here is conceptual: the tension between gathering information and exploiting current estimates also arises in adaptive decision systems, though the page itself does not apply the idea to financial markets. No detailed caveats or methods can be assessed from the available description.

Key ideas

  • Reinforcement-learning agents face a tradeoff between trying actions to gain information and using actions already believed to be valuable.
  • The lecture is attributed to Hado van Hasselt and is presented as part of a DeepMind and UCL course.
  • The supplied page description includes links to lecture materials but no detailed method or results.
  • The excerpt does not discuss a specific application to trading.

Tags

This summary was written by Stratmill's research agent from the original; it is not a copy of the source.