Skip to content
All library documents

Tsallis Entropy Exploration for Control with Latent Factors

Article arXiv papers · Author: Ryan Donnelly et al.

Summary

This work studies optimal control when a model contains latent factors and the decision maker chooses a distribution over actions rather than selecting a single action directly. It introduces Tsallis entropy as a reward for exploration of the state space, in both discrete-time and continuous-time settings. The authors derive an optimal state distribution with a q-Gaussian form and characterize its location through forward-backward equations for the respective settings.

The paper also examines how the resulting exploration solutions relate to standard dynamic optimal control. For a model-agnostic approach, it develops an optimal policy following the general idea of soft Q-learning. The method is presented as potentially useful for building more robust statistical arbitrage strategies, rather than as evidence that a particular trading system has been tested or improved. The supplied description gives theoretical results and an application direction, but no market data, trading performance, or practical implementation details. Its relevance to trading therefore depends on how the control framework is translated into a strategy and evaluated empirically.

Key ideas

  • The framework controls a distribution over actions in models with latent factors.
  • Tsallis entropy rewards exploration of the state space in discrete and continuous time.
  • The derived optimal state distribution has a q-Gaussian form.
  • The analysis relates exploration-aware solutions to standard dynamic optimal control.
  • A model-agnostic policy is developed along the lines of soft Q-learning and may inform statistical arbitrage research.

Tags

Full text
# Exploratory Control with Tsallis Entropy for Latent Factor Models


# Exploratory Control with Tsallis Entropy for Latent Factor Models









We study optimal control in models with latent factors where the agent controls the distribution over actions, rather than actions themselves, in both discrete and continuous time. To encourage exploration of the state space, we reward exploration with Tsallis Entropy and derive the optimal distribution over states - which we prove is $q$-Gaussian distributed with location characterized through the solution of an FBS$Δ$E and FBSDE in discrete and continuous time, respectively. We discuss the relation between the solutions of the optimal exploration problems and the standard dynamic optimal control solution. Finally, we develop the optimal policy in a model-agnostic setting along the lines of soft $Q$-learning. The approach may be applied in, e.g., developing more robust statistical arbitrage trading strategies.

Shown in full with attribution under the source's licence. Licence: abstract CC0

This summary was written by Stratmill's research agent from the original; it is not a copy of the source.