Skip to content
All library documents

Exploratory Reinforcement Learning for Speculative Trading and Pairs Trading

Article arXiv papers · Author: Yun Zhao et al.

Summary

This study formulates speculative trading as a sequential optimal stopping problem, choosing entry and exit times under general utility functions and price processes. It first relaxes the stopping problem by representing entry and exit events as jumps of Cox processes, whose bounded intensity controls are selected by the agent.

In the exploratory reinforcement learning formulation, randomized policies place probability distributions over these intensities, and Shannon differential entropy regularizes the objective. The resulting exploratory Hamilton-Jacobi-Bellman equations yield optimal policies in Gibbs form. The authors establish error estimates and convergence of the learning objective to the original problem's value function, then demonstrate an algorithm in a pairs-trading application. The supplied description gives no data or performance results for that application, and the stated convergence concerns the modeled objective rather than a demonstrated live-trading outcome.

Key ideas

  • Speculative trading is modeled as an optimal stopping problem over entry and exit times.
  • A relaxed formulation represents stopping events with Cox process jumps controlled through bounded intensities.
  • Randomized intensity policies are regularized using Shannon differential entropy.
  • The framework derives exploratory HJB equations and Gibbs-form optimal policies.
  • The work establishes error and convergence results and demonstrates an algorithm for pairs trading.

Tags

Full text
# Reinforcement Learning for Speculative Trading under Exploratory Framework


# Reinforcement Learning for Speculative Trading under Exploratory Framework









We study a speculative trading problem within the exploratory reinforcement learning (RL) framework of Wang et al. [2020]. The problem is formulated as a sequential optimal stopping problem over entry and exit times under general utility function and price process. We first consider a relaxed version of the problem in which the stopping times are modeled by the jump times of Cox processes driven by bounded, non-randomized intensity controls. Under the exploratory formulation, the agent's randomized control is characterized via the probability measure over the jump intensities, and their objective function is regularized by Shannon's differential entropy. This yields a system of the exploratory HJB equations and Gibbs distributions in closed-form as the optimal policy. Error estimates and convergence of the RL objective to the value function of the original problem are established. Finally, an RL algorithm is designed, and its implementation is showcased in a pairs-trading application.

Shown in full with attribution under the source's licence. Licence: abstract CC0

This summary was written by Stratmill's research agent from the original; it is not a copy of the source.