Skip to content
All library documents

Six ICLR 2022 Deep Learning Studies on Forecasting and Reinforcement Learning

Article BigQuant

Summary

This overview surveys six ICLR 2022 studies: periodic time-series forecasting, model-based policy optimization, disentangled representations, Stein-based sampling, deployment-efficient reinforcement learning, and training with privileged observations. The forecasting study, DEPTS, models global periodic states and progressively explains observed and predicted components. The policy-optimization work argues that learned model gradients matter alongside prediction accuracy and proposes separate models for prediction and gradient estimation. Other studies extend disentanglement into composed feature spaces, adapt Stein variational sampling through mirror maps, derive deployment-complexity bounds, and use a variational latent variable to learn from information available during training but hidden at decision time.

The article reports experiments on synthetic and real time series, continuous-control benchmarks, representation-learning tasks, and several reinforcement-learning settings, including games and mahjong. It also summarizes theoretical results for sampling convergence and deployment complexity. These are brief secondary descriptions rather than full paper evaluations; reported improvements depend on the studied tasks, and the article does not establish direct trading performance or applicability to financial markets.

Key ideas

  • DEPTS separates global periodic structure from local signal components and expands them through successive prediction layers.
  • Model-based policy optimization can depend on model-gradient accuracy as well as prediction accuracy.
  • RecurD propagates disentanglement constraints through composed feature representations.
  • Mirror-based Stein methods adapt variational sampling to constrained or geometrically difficult distributions.
  • Deployment-efficient reinforcement learning studies how to reduce the number of policy deployments needed for learning.
  • VLOG uses information available during training to guide learning while making decisions from observable information.

Tags

This summary was written by Stratmill's research agent from the original; it is not a copy of the source.