활성 고빈도 주식 트레이딩을 위한 PPO 강화학습
기사 arXiv papers · 저자: Antonio Briola et al.
요약
이 문서는 능동적 고빈도 주식 트레이딩을 위한 종단 간 심층 강화학습 접근법을 설명합니다. 에이전트는 서로 다른 지정가 주문장 특성 집합으로 구성한 상태를 바탕으로 근접 정책 최적화를 사용해 Intel 주식 1주를 거래합니다. 신호 대 잡음비를 높이기 위해 큰 가격 변동이 발생한 학습 데이터를 선택합니다. 연속된 한 달은 검증에, 그다음 별도의 한 달은 테스트에 사용합니다. 순차적 모형 기반 최적화로 하이퍼파라미터를 조정합니다.
저자들은 에이전트가 주문장에서 반복되는 패턴을 찾아내고, 잡음이 많고 변하는 상황에서도 테스트 기간에 안정적인 양의 수익률을 낸다고 보고합니다. 이는 짧은 몇 달 동안 단일 종목을 대상으로 한 제한적인 실험의 근거입니다. 설명에는 수익률 통계, 거래 비용, 벤치마크 비교 또는 학습된 행동이 다른 기간과 자산, 실거래에서도 유지된다는 근거가 없습니다. 가격 움직임에 따라 학습 표본을 선택하는 방식도 에이전트가 학습하는 내용에 영향을 줄 수 있으므로, 보고된 결과를 수익성의 일반적인 증거로 받아들여서는 안 됩니다.
핵심 아이디어
- 에이전트는 단일 종목 1주를 거래할 때 근접 정책 최적화를 사용합니다.
- 상태 표현은 포함하는 지정가 주문장 특성에 따라 달라집니다.
- 학습은 큰 가격 변동이 있는 표본을 중시하며 검증 기간과 테스트 기간을 따로 둡니다.
- 순차적 모형 기반 최적화로 하이퍼파라미터를 조정합니다.
- 양의 테스트 수익률이 보고되었지만, 더 폭넓은 일반화와 체결 비용은 입증되지 않았습니다.
태그
전문
# Deep Reinforcement Learning for Active High Frequency Trading # Deep Reinforcement Learning for Active High Frequency Trading We introduce the first end-to-end Deep Reinforcement Learning (DRL) based framework for active high frequency trading in the stock market. We train DRL agents to trade one unit of Intel Corporation stock by employing the Proximal Policy Optimization algorithm. The training is performed on three contiguous months of high frequency Limit Order Book data, of which the last month constitutes the validation data. In order to maximise the signal to noise ratio in the training data, we compose the latter by only selecting training samples with largest price changes. The test is then carried out on the following month of data. Hyperparameters are tuned using the Sequential Model Based Optimization technique. We consider three different state characterizations, which differ in their LOB-based meta-features. Analysing the agents' performances on test data, we argue that the agents are able to create a dynamic representation of the underlying environment. They identify occasional regularities present in the data and exploit them to create long-term profitable trading strategies. Indeed, agents learn trading strategies able to produce stable positive returns in spite of the highly stochastic and non-stationary environment.
출처의 라이선스에 따라 출처를 표시하고 전문을 공개합니다. 라이선스: abstract CC0
이 요약은 원문을 바탕으로 Stratmill의 리서치 에이전트가 작성했으며, 원문을 복사한 것이 아닙니다.