잠재 요인 제어를 위한 차리스 엔트로피 탐색
기사 arXiv papers · 저자: Ryan Donnelly et al.
요약
이 연구는 모델에 잠재 요인이 있고 의사결정자가 단일 행동을 직접 선택하는 대신 행동에 대한 분포를 고르는 상황에서 최적 제어를 살펴봅니다. 이산시간과 연속시간 설정 모두에서 상태 공간 탐색에 대한 보상으로 차리스 엔트로피를 도입합니다. 각 설정의 순방향·역방향 방정식을 통해 q-가우시안 형태의 최적 상태 분포를 도출하고 그 위치를 규정합니다.
논문은 이렇게 얻은 탐색 해법과 표준 동적 최적 제어의 관계도 살펴봅니다. 모델에 구애받지 않는 접근법으로 소프트 Q 학습의 일반적인 아이디어를 따르는 최적 정책을 개발합니다. 이 방법은 특정 트레이딩 시스템을 테스트하거나 개선했다는 증거가 아니라 더 견고한 통계적 차익거래 전략 구축에 잠재적으로 유용한 것으로 제시됩니다. 제공된 설명에는 이론적 결과와 적용 방향은 있지만 시장 데이터, 트레이딩 성과, 실제 구현 세부 사항은 없습니다. 따라서 트레이딩 관련성은 제어 틀을 전략으로 어떻게 옮기고 실증 평가하는지에 달려 있습니다.
핵심 아이디어
- 잠재 요인이 있는 모델에서 행동 분포를 제어하는 틀입니다.
- 차리스 엔트로피는 이산시간 및 연속시간에서 상태 공간 탐색을 보상합니다.
- 도출된 최적 상태 분포는 q-가우시안 형태를 띱니다.
- 탐색을 고려한 해법과 표준 동적 최적 제어의 관계를 분석합니다.
- 소프트 Q 학습 방식의 모델 비의존 정책을 개발하며, 통계적 차익거래 연구에 참고가 될 수 있습니다.
태그
전문
# Exploratory Control with Tsallis Entropy for Latent Factor Models # Exploratory Control with Tsallis Entropy for Latent Factor Models We study optimal control in models with latent factors where the agent controls the distribution over actions, rather than actions themselves, in both discrete and continuous time. To encourage exploration of the state space, we reward exploration with Tsallis Entropy and derive the optimal distribution over states - which we prove is $q$-Gaussian distributed with location characterized through the solution of an FBS$Δ$E and FBSDE in discrete and continuous time, respectively. We discuss the relation between the solutions of the optimal exploration problems and the standard dynamic optimal control solution. Finally, we develop the optimal policy in a model-agnostic setting along the lines of soft $Q$-learning. The approach may be applied in, e.g., developing more robust statistical arbitrage trading strategies.
출처의 라이선스에 따라 출처를 표시하고 전문을 공개합니다. 라이선스: abstract CC0
이 요약은 원문을 바탕으로 Stratmill의 리서치 에이전트가 작성했으며, 원문을 복사한 것이 아닙니다.