用于限价订单簿消息流的自回归深度状态空间模型
文章 arXiv papers · 作者: Peer Nagy et al.
总结
本文开发了一种自回归模型,用于生成经过词元化的限价订单簿消息,再由模拟器据此更新订单簿状态。其架构采用简化的结构化状态空间层,处理长消息序列和状态序列。作者还为LOBSTER数据中的NASDAQ订单消息创建了词元器,将连续数字分组成词元。
在样本外评估中,模型以较低的困惑度逼近观测到的消息分布。根据生成的订单流计算出的中间价收益与数据显著相关,论文将此视为具备条件预测能力的证据。作者提出,细致的合成订单流可支持预测以外的研究,包括为高频强化学习提供模拟环境。这些结果仅涉及所报告的数据集和指标;摘要未提供更广泛的市场验证、执行结果,或模拟订单流能够重现交易应用所需全部特征的证据。
核心观点
- 模型以自回归方式生成词元化的限价订单簿消息序列。
- 结构化状态空间层用于处理较长的消息序列和订单簿状态序列。
- 研究将自定义词元器应用于LOBSTER数据中的NASDAQ股票订单消息。
- 样本外评估报告称,困惑度较低,生成的中间价收益与观测值显著相关。
- 合成订单流可用作高频研究的建模环境,但描述未提及更广泛的验证。
标签
全文
# 2309.00638 # Generative AI for End-to-End Limit Order Book Modelling: A Token-Level Autoregressive Generative Model of Message Flow Using a Deep State Space Network Developing a generative model of realistic order flow in financial markets is a challenging open problem, with numerous applications for market participants. Addressing this, we propose the first end-to-end autoregressive generative model that generates tokenized limit order book (LOB) messages. These messages are interpreted by a Jax-LOB simulator, which updates the LOB state. To handle long sequences efficiently, the model employs simplified structured state-space layers to process sequences of order book states and tokenized messages. Using LOBSTER data of NASDAQ equity LOBs, we develop a custom tokenizer for message data, converting groups of successive digits to tokens, similar to tokenization in large language models. Out-of-sample results show promising performance in approximating the data distribution, as evidenced by low model perplexity. Furthermore, the mid-price returns calculated from the generated order flow exhibit a significant correlation with the data, indicating impressive conditional forecast performance. Due to the granularity of generated data, and the accuracy of the model, it offers new application areas for future work beyond forecasting, e.g. acting as a world model in high-frequency financial reinforcement learning applications. Overall, our results invite the use and extension of the model in the direction of autoregressive large financial models for the generation of high-frequency financial data and we commit to open-sourcing our code to facilitate future research.
在遵守原作品许可的前提下,附作者信息全文展示。 许可协议: abstract CC0
此摘要由 Stratmill 研究智能体根据原文撰写,并非原文副本。