Nuclear Norm Curiosity Rewards for Reinforcement Learning in Trading
Summary
The article presents nuclear norm maximization as an intrinsic reward for reinforcement learning agents exploring environments with sparse external rewards. It forms a matrix from encoded representations of a visited state and its nearest neighboring states, using the nuclear norm as a continuous proxy for state diversity. The proposed reward normalizes this measure by the matrix’s Frobenius norm to limit the influence of low entropy and adapt the reward scale to matrix dimensions. The approach is integrated into RE3, which already uses an encoder and nearest states, with a random convolutional encoder and decomposed trading rewards.
The article cites prior research reporting better exploration performance than other methods, including under added noise. In its own trading implementation, testing reportedly produced more varied agent behavior than plain RE3, but also more chaotic trading; the reported profit factor was 1.02. The author says balancing exploration against exploitation remains unresolved. These results are specific to the described experiment and do not establish that the method generalizes or is profitable in live markets.
Key ideas
- The method uses nuclear norm as a continuous proxy for diversity among encoded neighboring states.
- The intrinsic reward is normalized by the Frobenius norm to reduce the influence of low entropy and control its scale.
- The article integrates the reward into RE3 with a random convolutional encoder and decomposed trading rewards.
- The reported trading experiment produced more varied but more chaotic behavior than plain RE3, with a profit factor of 1.02.
- The balance between exploration and exploitation remains an open issue in the implementation.
Tags
This summary was written by Stratmill's research agent from the original; it is not a copy of the source.