본문으로 건너뛰기
라이브러리 문서 전체

지정가 주문장 메시지 흐름을 위한 자기회귀 심층 상태공간 모형

기사 arXiv papers · 저자: Peer Nagy et al.

요약

이 논문은 토큰화된 지정가 주문장 메시지를 생성하고 시뮬레이터가 이를 사용해 주문장 상태를 갱신하는 자기회귀 모형을 개발합니다. 이 구조는 간소화된 구조적 상태공간 계층을 사용해 긴 메시지 및 상태 시퀀스를 처리합니다. 저자들은 LOBSTER 데이터의 NASDAQ 주문 메시지용 토크나이저도 만들어 연속된 숫자를 토큰으로 묶습니다.

표본 외 평가에서 이 모형은 낮은 퍼플렉서티로 관측된 메시지 분포를 근사합니다. 생성된 주문 흐름에서 계산한 중간 가격 수익률은 데이터와 유의한 상관관계를 보이며, 논문은 이를 조건부 예측 능력의 근거로 제시합니다. 저자들은 상세한 합성 주문 흐름이 예측을 넘어 고빈도 강화학습을 위한 시뮬레이션 환경 등에도 활용될 수 있다고 제안합니다. 이 결과는 보고된 데이터셋과 측정치에 관한 것입니다. 초록에는 더 넓은 시장에서의 검증, 실행 결과, 또는 시뮬레이션 흐름이 트레이딩 응용에 필요한 모든 특성을 재현한다는 증거가 없습니다.

핵심 아이디어

  • 이 모형은 토큰화된 지정가 주문장 메시지 시퀀스를 자기회귀 방식으로 생성합니다.
  • 구조적 상태공간 계층으로 긴 메시지와 주문장 상태 시퀀스를 처리합니다.
  • 맞춤형 토크나이저를 LOBSTER 데이터의 NASDAQ 주식 주문 메시지에 적용합니다.
  • 표본 외 평가에서는 퍼플렉서티가 낮고 생성 및 관측 중간 가격 수익률 간 상관관계가 유의하다고 보고합니다.
  • 합성 주문 흐름은 고빈도 연구의 모델링 환경으로 쓰일 수 있지만, 더 폭넓은 검증은 설명되어 있지 않습니다.

태그

전문
# 2309.00638


# Generative AI for End-to-End Limit Order Book Modelling: A Token-Level Autoregressive Generative Model of Message Flow Using a Deep State Space Network









Developing a generative model of realistic order flow in financial markets is a challenging open problem, with numerous applications for market participants. Addressing this, we propose the first end-to-end autoregressive generative model that generates tokenized limit order book (LOB) messages. These messages are interpreted by a Jax-LOB simulator, which updates the LOB state. To handle long sequences efficiently, the model employs simplified structured state-space layers to process sequences of order book states and tokenized messages. Using LOBSTER data of NASDAQ equity LOBs, we develop a custom tokenizer for message data, converting groups of successive digits to tokens, similar to tokenization in large language models. Out-of-sample results show promising performance in approximating the data distribution, as evidenced by low model perplexity. Furthermore, the mid-price returns calculated from the generated order flow exhibit a significant correlation with the data, indicating impressive conditional forecast performance. Due to the granularity of generated data, and the accuracy of the model, it offers new application areas for future work beyond forecasting, e.g. acting as a world model in high-frequency financial reinforcement learning applications. Overall, our results invite the use and extension of the model in the direction of autoregressive large financial models for the generation of high-frequency financial data and we commit to open-sourcing our code to facilitate future research.

출처의 라이선스에 따라 출처를 표시하고 전문을 공개합니다. 라이선스: abstract CC0

이 요약은 원문을 바탕으로 Stratmill의 리서치 에이전트가 작성했으며, 원문을 복사한 것이 아닙니다.