跳至正文
返回文库全部文档

潜在因子控制中的 Tsallis 熵探索

文章 arXiv papers · 作者: Ryan Donnelly et al.

总结

本文研究含潜在因子的模型中的最优控制问题,其中决策者选择的是行动分布,而非直接选择单一行动。研究在离散时间和连续时间设定中引入 Tsallis 熵作为探索状态空间的奖励。作者推导出具有 q-Gaussian 形式的最优状态分布,并通过相应设定下的前向–后向方程刻画其位置。

本文还考察由此得到的探索解与标准动态最优控制之间的关系。针对模型无关的方法,作者依据软 Q 学习的一般思路构建了最优策略。该方法被提出作为构建更稳健统计套利策略的潜在参考,并非某个具体交易系统已经经过测试或改进的证据。所提供的描述包含理论结果和应用方向,但没有市场数据、交易绩效或实际实施细节。因此,这项研究与交易的相关性取决于如何将控制框架转化为策略并进行实证评估。

核心观点

  • 该框架控制含潜在因子的模型中的行动分布。
  • Tsallis 熵在离散和连续时间中为探索状态空间提供奖励。
  • 推导出的最优状态分布具有 q-Gaussian 形式。
  • 分析将考虑探索的解与标准动态最优控制联系起来。
  • 研究沿用软 Q 学习思路开发了模型无关策略,或可为统计套利研究提供参考。

标签

全文
# Exploratory Control with Tsallis Entropy for Latent Factor Models


# Exploratory Control with Tsallis Entropy for Latent Factor Models









We study optimal control in models with latent factors where the agent controls the distribution over actions, rather than actions themselves, in both discrete and continuous time. To encourage exploration of the state space, we reward exploration with Tsallis Entropy and derive the optimal distribution over states - which we prove is $q$-Gaussian distributed with location characterized through the solution of an FBS$Δ$E and FBSDE in discrete and continuous time, respectively. We discuss the relation between the solutions of the optimal exploration problems and the standard dynamic optimal control solution. Finally, we develop the optimal policy in a model-agnostic setting along the lines of soft $Q$-learning. The approach may be applied in, e.g., developing more robust statistical arbitrage trading strategies.

在遵守原作品许可的前提下,附作者信息全文展示。 许可协议: abstract CC0

此摘要由 Stratmill 研究智能体根据原文撰写,并非原文副本。