Deep Reinforcement Learning for Limit-Order Execution with Forecast Signals
Summary
The study uses deep reinforcement learning to turn a short-term trading forecast into limit-order placement and inventory decisions in a limit order book. It trains an agent in a simulated NASDAQ equity environment built from historical order-book messages, using the ABIDES simulator. The agent observes the current book, recent market history, and a directional signal, then learns an execution policy with Deep Double Duelling Q-learning and prioritized asynchronous experience replay.
To assess execution separately from any particular forecasting method, the researchers test synthetic alpha signals with different levels of noise. They report that the learned agent manages inventory and places orders effectively, outperforming a heuristic strategy given the same signal. This comparison supports the execution method within the described simulator, but the document does not provide performance figures or evidence of live-market results. The quality and realism of the synthetic signals and simulated environment therefore limit conclusions about deployment.
Key ideas
- The agent learns to convert a directional forecast into limit-order and inventory decisions.
- Training uses a historical-message-based NASDAQ limit order book simulator.
- The agent observes book state, recent history, and a short-term forecast.
- Synthetic signals with varying noise let the researchers focus on execution rather than forecasting quality.
- In simulation, the learned policy outperforms a heuristic using the same signal, though live performance is not established.
Tags
Full text
# Asynchronous Deep Double Duelling Q-Learning for Trading-Signal Execution in Limit Order Book Markets # Asynchronous Deep Double Duelling Q-Learning for Trading-Signal Execution in Limit Order Book Markets We employ deep reinforcement learning (RL) to train an agent to successfully translate a high-frequency trading signal into a trading strategy that places individual limit orders. Based on the ABIDES limit order book simulator, we build a reinforcement learning OpenAI gym environment and utilise it to simulate a realistic trading environment for NASDAQ equities based on historic order book messages. To train a trading agent that learns to maximise its trading return in this environment, we use Deep Duelling Double Q-learning with the APEX (asynchronous prioritised experience replay) architecture. The agent observes the current limit order book state, its recent history, and a short-term directional forecast. To investigate the performance of RL for adaptive trading independently from a concrete forecasting algorithm, we study the performance of our approach utilising synthetic alpha signals obtained by perturbing forward-looking returns with varying levels of noise. Here, we find that the RL agent learns an effective trading strategy for inventory management and order placing that outperforms a heuristic benchmark trading strategy having access to the same signal.
Shown in full with attribution under the source's licence. Licence: abstract CC0
This summary was written by Stratmill's research agent from the original; it is not a copy of the source.