PPO Reinforcement Learning for Active High-Frequency Stock Trading
Summary
The document describes an end-to-end deep reinforcement learning approach to active high-frequency stock trading. Agents use Proximal Policy Optimization to trade one unit of Intel stock, with states built from different sets of limit order book features. Training data are selected for large price changes to increase the signal-to-noise ratio; a contiguous month is reserved for validation, followed by a separate month for testing. Hyperparameters are tuned with sequential model-based optimization.
The authors report that agents find recurring patterns in the order book and produce stable positive returns on the test period despite noisy, changing conditions. This is evidence from a narrowly scoped experiment in one stock over a short sequence of months. The description gives no return statistics, transaction costs, benchmark comparisons, or evidence that the learned behavior survives other periods, assets, and live execution. Selecting training samples by price movement may also shape what the agents learn, so the reported result should not be taken as broad proof of profitability.
Key ideas
- The agents use Proximal Policy Optimization to trade one unit of a single stock.
- Their state representations vary in the limit order book features they include.
- Training emphasizes samples with large price moves, with separate validation and test periods.
- Hyperparameters are tuned through sequential model-based optimization.
- Positive test returns are reported, but broader generalization and execution costs are not established.
Tags
Full text
# Deep Reinforcement Learning for Active High Frequency Trading # Deep Reinforcement Learning for Active High Frequency Trading We introduce the first end-to-end Deep Reinforcement Learning (DRL) based framework for active high frequency trading in the stock market. We train DRL agents to trade one unit of Intel Corporation stock by employing the Proximal Policy Optimization algorithm. The training is performed on three contiguous months of high frequency Limit Order Book data, of which the last month constitutes the validation data. In order to maximise the signal to noise ratio in the training data, we compose the latter by only selecting training samples with largest price changes. The test is then carried out on the following month of data. Hyperparameters are tuned using the Sequential Model Based Optimization technique. We consider three different state characterizations, which differ in their LOB-based meta-features. Analysing the agents' performances on test data, we argue that the agents are able to create a dynamic representation of the underlying environment. They identify occasional regularities present in the data and exploit them to create long-term profitable trading strategies. Indeed, agents learn trading strategies able to produce stable positive returns in spite of the highly stochastic and non-stationary environment.
Shown in full with attribution under the source's licence. Licence: abstract CC0
This summary was written by Stratmill's research agent from the original; it is not a copy of the source.