본문으로 건너뛰기
라이브러리 문서 전체

예측 신호를 활용한 지정가 주문 실행의 심층 강화학습

기사 arXiv papers · 저자: Peer Nagy et al.

요약

이 연구는 심층 강화학습으로 단기 트레이딩 예측을 지정가 주문 배치와 지정가 주문장 내 재고 결정으로 전환합니다. 과거 주문장 메시지로 구축한 모의 NASDAQ 주식 환경에서 ABIDES 시뮬레이터를 사용해 에이전트를 학습합니다. 에이전트는 현재 주문장, 최근 시장 기록, 방향성 신호를 관찰한 뒤 심층 이중 결투 Q 학습과 우선순위 기반 비동기 경험 재생으로 실행 정책을 학습합니다.

특정 예측 방법과 분리해 실행을 평가하기 위해 연구자들은 잡음 수준이 다른 합성 알파 신호를 시험합니다. 학습된 에이전트가 재고를 관리하고 주문을 효과적으로 배치해 같은 신호를 받은 휴리스틱 전략보다 나은 성과를 냈다고 보고합니다. 이 비교는 설명된 시뮬레이터 안에서 실행 방법을 뒷받침하지만, 문서에는 성과 수치나 실거래 결과의 근거가 없습니다. 따라서 합성 신호와 모의 환경의 품질 및 현실성이 실제 적용에 관한 결론을 제한합니다.

핵심 아이디어

  • 에이전트는 방향성 예측을 지정가 주문 및 재고 결정으로 전환하는 법을 학습합니다.
  • 과거 메시지 기반 NASDAQ 지정가 주문장 시뮬레이터로 학습합니다.
  • 에이전트는 주문장 상태, 최근 기록, 단기 예측을 관찰합니다.
  • 잡음 수준이 다른 합성 신호를 사용해 예측 품질과 분리하여 실행을 평가합니다.
  • 시뮬레이션에서는 학습된 정책이 같은 신호를 쓰는 휴리스틱보다 나았지만, 실거래 성과는 입증되지 않았습니다.

태그

전문
# Asynchronous Deep Double Duelling Q-Learning for Trading-Signal Execution in Limit Order Book Markets


# Asynchronous Deep Double Duelling Q-Learning for Trading-Signal Execution in Limit Order Book Markets









We employ deep reinforcement learning (RL) to train an agent to successfully translate a high-frequency trading signal into a trading strategy that places individual limit orders. Based on the ABIDES limit order book simulator, we build a reinforcement learning OpenAI gym environment and utilise it to simulate a realistic trading environment for NASDAQ equities based on historic order book messages. To train a trading agent that learns to maximise its trading return in this environment, we use Deep Duelling Double Q-learning with the APEX (asynchronous prioritised experience replay) architecture. The agent observes the current limit order book state, its recent history, and a short-term directional forecast. To investigate the performance of RL for adaptive trading independently from a concrete forecasting algorithm, we study the performance of our approach utilising synthetic alpha signals obtained by perturbing forward-looking returns with varying levels of noise. Here, we find that the RL agent learns an effective trading strategy for inventory management and order placing that outperforms a heuristic benchmark trading strategy having access to the same signal.

출처의 라이선스에 따라 출처를 표시하고 전문을 공개합니다. 라이선스: abstract CC0

이 요약은 원문을 바탕으로 Stratmill의 리서치 에이전트가 작성했으며, 원문을 복사한 것이 아닙니다.