コンテンツへスキップ
ライブラリの全資料

潜在因子を含む制御におけるツァリスエントロピー探索

記事 arXiv papers · 著者: Ryan Donnelly et al.

サマリー

この研究は、モデルに潜在因子が含まれ、意思決定者が単一の行動を直接選ぶのではなく、行動の分布を選択する最適制御を調べています。離散時間と連続時間の両方で、状態空間の探索に対する報酬としてツァリスエントロピーを導入しています。著者らはqガウス型の最適状態分布を導き、それぞれの設定における前進後退方程式によってその位置を特徴づけています。

この論文は、得られた探索解と標準的な動的最適制御との関係も検討しています。モデルに依存しない方法として、ソフトQ学習の一般的な考え方に沿った最適方策を開発しています。この手法は、より頑健な統計的裁定戦略の構築に役立つ可能性があるものとして提示されていますが、特定の取引システムが検証または改善された証拠ではありません。提示された説明には理論結果と応用の方向性はありますが、市場データ、取引成績、実装の詳細はありません。取引との関連性は、この制御枠組みを戦略に変換し、実証的に評価する方法に左右されます。

主なアイデア

  • 潜在因子を含むモデルで、行動の分布を制御する枠組みです。
  • ツァリスエントロピーは離散時間と連続時間の状態空間探索に報酬を与えます。
  • 導かれる最適状態分布はqガウス型です。
  • 探索を考慮した解と標準的な動的最適制御との関係を分析しています。
  • ソフトQ学習に沿ったモデル非依存方策を開発し、統計的裁定の研究に役立つ可能性を示しています。

タグ

全文
# Exploratory Control with Tsallis Entropy for Latent Factor Models


# Exploratory Control with Tsallis Entropy for Latent Factor Models









We study optimal control in models with latent factors where the agent controls the distribution over actions, rather than actions themselves, in both discrete and continuous time. To encourage exploration of the state space, we reward exploration with Tsallis Entropy and derive the optimal distribution over states - which we prove is $q$-Gaussian distributed with location characterized through the solution of an FBS$Δ$E and FBSDE in discrete and continuous time, respectively. We discuss the relation between the solutions of the optimal exploration problems and the standard dynamic optimal control solution. Finally, we develop the optimal policy in a model-agnostic setting along the lines of soft $Q$-learning. The approach may be applied in, e.g., developing more robust statistical arbitrage trading strategies.

出典を明記したうえで、ライセンスに従って全文を掲載しています。 ライセンス: abstract CC0

この要約は原文をもとにStratmillのリサーチエージェントが作成したもので、出典の複製ではありません。