跳至正文
返回文库全部文档

用于投机交易与配对交易的探索式强化学习

文章 arXiv papers · 作者: Yun Zhao et al.

总结

本研究将投机交易表述为序贯最优停止问题,在一般效用函数和价格过程中选择入场与离场时点。首先,研究通过将入场和离场事件表示为考克斯过程的跳跃来放宽停止问题;其有界强度控制由智能体选择。

在探索式强化学习表述中,随机化策略为这些强度赋予概率分布,并使用香农微分熵对目标进行正则化。由此得到的探索式哈密顿—雅可比—贝尔曼方程给出吉布斯形式的最优策略。作者建立了误差估计,并证明学习目标收敛到原问题的值函数,随后在配对交易应用中展示了一种算法。所提供的描述未给出该应用的数据或表现结果,且所述收敛针对建模目标,并非已展示的实盘交易结果。

核心观点

  • 投机交易被建模为围绕入场与离场时点的最优停止问题。
  • 一种放宽后的表述用考克斯过程跳跃表示停止事件,并通过有界强度进行控制。
  • 随机化强度策略使用香农微分熵进行正则化。
  • 该框架推导出探索式HJB方程和吉布斯形式的最优策略。
  • 研究建立了误差与收敛结果,并展示了用于配对交易的算法。

标签

全文
# Reinforcement Learning for Speculative Trading under Exploratory Framework


# Reinforcement Learning for Speculative Trading under Exploratory Framework









We study a speculative trading problem within the exploratory reinforcement learning (RL) framework of Wang et al. [2020]. The problem is formulated as a sequential optimal stopping problem over entry and exit times under general utility function and price process. We first consider a relaxed version of the problem in which the stopping times are modeled by the jump times of Cox processes driven by bounded, non-randomized intensity controls. Under the exploratory formulation, the agent's randomized control is characterized via the probability measure over the jump intensities, and their objective function is regularized by Shannon's differential entropy. This yields a system of the exploratory HJB equations and Gibbs distributions in closed-form as the optimal policy. Error estimates and convergence of the RL objective to the value function of the original problem are established. Finally, an RL algorithm is designed, and its implementation is showcased in a pairs-trading application.

在遵守原作品许可的前提下,附作者信息全文展示。 许可协议: abstract CC0

此摘要由 Stratmill 研究智能体根据原文撰写,并非原文副本。