利用预测信号进行限价单执行的深度强化学习
文章 arXiv papers · 作者: Peer Nagy et al.
总结
本研究使用深度强化学习,将短期交易预测转化为限价单挂单和限价订单簿中的库存决策。研究人员在一个由历史订单簿消息构建的模拟NASDAQ股票交易环境中,使用ABIDES模拟器训练智能体。智能体观察当前订单簿、近期市场历史和方向信号,然后利用深度双重决斗Q学习以及优先异步经验回放来学习执行策略。
为了单独评估执行效果,不依赖某种特定的预测方法,研究人员测试了不同噪声水平下的合成阿尔法信号。他们报告称,训练出的智能体能够有效管理库存和挂单,在使用相同信号时优于启发式策略。这一比较支持该执行方法在所述模拟器中的表现,但文中没有提供绩效数据或实盘结果证据。因此,合成信号和模拟环境的质量与真实性限制了对其可部署性的判断。
核心观点
- 智能体学习将方向预测转化为限价单和库存决策。
- 训练使用基于历史消息的NASDAQ限价订单簿模拟器。
- 智能体观察订单簿状态、近期历史数据和短期预测。
- 研究人员使用不同噪声水平的合成信号,以便专注评估执行质量而非预测质量。
- 在模拟中,训练出的策略优于使用相同信号的启发式策略,但其实际市场表现尚未得到证明。
标签
全文
# Asynchronous Deep Double Duelling Q-Learning for Trading-Signal Execution in Limit Order Book Markets # Asynchronous Deep Double Duelling Q-Learning for Trading-Signal Execution in Limit Order Book Markets We employ deep reinforcement learning (RL) to train an agent to successfully translate a high-frequency trading signal into a trading strategy that places individual limit orders. Based on the ABIDES limit order book simulator, we build a reinforcement learning OpenAI gym environment and utilise it to simulate a realistic trading environment for NASDAQ equities based on historic order book messages. To train a trading agent that learns to maximise its trading return in this environment, we use Deep Duelling Double Q-learning with the APEX (asynchronous prioritised experience replay) architecture. The agent observes the current limit order book state, its recent history, and a short-term directional forecast. To investigate the performance of RL for adaptive trading independently from a concrete forecasting algorithm, we study the performance of our approach utilising synthetic alpha signals obtained by perturbing forward-looking returns with varying levels of noise. Here, we find that the RL agent learns an effective trading strategy for inventory management and order placing that outperforms a heuristic benchmark trading strategy having access to the same signal.
在遵守原作品许可的前提下,附作者信息全文展示。 许可协议: abstract CC0
此摘要由 Stratmill 研究智能体根据原文撰写,并非原文副本。