Fully Parameterized Quantile Functions for Distributional Reinforcement Learning
Summary
This article explains Fully Parameterized Quantile Functions (FQF), a distributional reinforcement-learning method for estimating reward distributions in Q-learning. Earlier approaches use fixed reward categories or predetermined quantiles, which require choices about the distribution and can trade accuracy against model size and training time. Implicit Quantile Networks address this by sampling quantiles, but random samples may be inefficient and add uncertainty.
FQF replaces random quantile generation with a learned network that produces a state-dependent quantile distribution. A softmax output is accumulated to form ordered quantiles, while another network estimates the quantile function. The article summarizes results reported by the original method on 55 Atari games and notes the computational cost of adding the quantile-generation model; the source authors recommend 32 quantiles. It then discusses an MQL5 implementation that packages the components in a neural-network layer. This is a technical account of reinforcement learning, not a trading strategy evaluation, and the article cautions that its example programs are not intended for live trading without thorough testing.
Key ideas
- Fixed quantile methods require assumptions about the reward distribution and choices that affect model cost and accuracy.
- IQN samples quantiles to approximate a wider range of reward distributions, but sampling may be inefficient.
- FQF learns a state-dependent quantile distribution using a neural network and cumulative softmax outputs.
- The method adds computation for quantile generation, and the article reports original-study tests on 55 Atari games.
- The MQL5 example implements the method for study and testing, not as evidence of live trading performance.
Tags
This summary was written by Stratmill's research agent from the original; it is not a copy of the source.