跳至正文
返回文库全部文档

用于主动股票高频交易的PPO强化学习

文章 arXiv papers · 作者: Antonio Briola et al.

总结

本文介绍一种端到端深度强化学习方法,用于主动股票高频交易。智能体使用近端策略优化交易一单位英特尔股票,状态由不同组合的限价订单簿特征构成。为提高信噪比,训练数据选取价格变动较大的样本;连续的一个月留作验证,之后另取一个月用于测试。超参数通过序贯模型优化进行调整。

作者报告称,尽管市场环境嘈杂且不断变化,智能体仍能在订单簿中发现重复模式,并在测试期取得稳定的正回报。这是针对单只股票、跨越数月的短期实验,其范围较窄。描述未提供回报统计、交易成本、基准比较,也未证明学习到的行为能在其他时期、资产或实盘执行中持续有效。根据价格变动选择训练样本也可能影响智能体的学习内容,因此不应将报告结果视为广泛的盈利能力证明。

核心观点

  • 智能体使用近端策略优化交易单只股票的一单位。
  • 其状态表示因纳入的限价订单簿特征不同而有所区别。
  • 训练侧重价格大幅波动的样本,并设有独立的验证期和测试期。
  • 超参数通过序贯模型优化进行调整。
  • 报告称测试回报为正,但尚未证明其具备更广泛的泛化能力,也未说明执行成本。

标签

全文
# Deep Reinforcement Learning for Active High Frequency Trading


# Deep Reinforcement Learning for Active High Frequency Trading









We introduce the first end-to-end Deep Reinforcement Learning (DRL) based framework for active high frequency trading in the stock market. We train DRL agents to trade one unit of Intel Corporation stock by employing the Proximal Policy Optimization algorithm. The training is performed on three contiguous months of high frequency Limit Order Book data, of which the last month constitutes the validation data. In order to maximise the signal to noise ratio in the training data, we compose the latter by only selecting training samples with largest price changes. The test is then carried out on the following month of data. Hyperparameters are tuned using the Sequential Model Based Optimization technique. We consider three different state characterizations, which differ in their LOB-based meta-features. Analysing the agents' performances on test data, we argue that the agents are able to create a dynamic representation of the underlying environment. They identify occasional regularities present in the data and exploit them to create long-term profitable trading strategies. Indeed, agents learn trading strategies able to produce stable positive returns in spite of the highly stochastic and non-stationary environment.

在遵守原作品许可的前提下,附作者信息全文展示。 许可协议: abstract CC0

此摘要由 Stratmill 研究智能体根据原文撰写,并非原文副本。