Clustering Limit Order Events to Build Order Flow Signals
Summary
ClusterLOB groups individual market events from market-by-order data using six time-dependent features and K-means++ clustering. The resulting groups are interpreted as directional, opportunistic, and market-making participants, providing a way to study different behaviors within a limit order book.
The study uses one year of NASDAQ data spanning small-, medium-, and large-tick stocks. It measures order flow imbalance by cluster in 30-minute intervals and evaluates signals through trading strategies, including comparisons with non-clustered benchmarks on training and test data. It also examines imbalance signals for order additions, cancellations, and trades. The reported test performance is stronger for the strategy selected by Sharpe ratio on training data, but the document gives no detailed performance figures or further evidence about generalization beyond this dataset.
Key ideas
- Market-by-order events are represented with six time-dependent features before clustering.
- K-means++ assigns events to three groups interpreted as directional, opportunistic, and market-making behavior.
- Cluster-level order flow imbalances are computed in 30-minute intervals as trading signals.
- The study compares strategies using these signals with benchmarks and separates training from test data.
- Separate imbalance analyses cover order additions, cancellations, and trades.
Tags
Full text
# ClusterLOB: Enhancing Trading Strategies by Clustering Orders in Limit Order Books # ClusterLOB: Enhancing Trading Strategies by Clustering Orders in Limit Order Books In the rapidly evolving world of financial markets, understanding the dynamics of limit order book (LOB) is crucial for unraveling market microstructure and participant behavior. We introduce ClusterLOB as a method to cluster individual market events in a stream of market-by-order (MBO) data into different groups. To do so, each market event is augmented with six time-dependent features. By applying the K-means++ clustering algorithm to the resulting order features, we are then able to assign each new order to one of three distinct clusters, which we identify as directional, opportunistic, and market-making participants, each capturing unique trading behaviors. Our experimental results are performed on one year of MBO data containing small-tick, medium-tick, and large-tick stocks from NASDAQ. To validate the usefulness of our clustering, we compute order flow imbalances across each cluster within 30-minute buckets during the trading day. We treat each cluster's imbalance as a signal that provides insights into trading strategies and participants' responses to varying market conditions. To assess the effectiveness of these signals, we identify the trading strategy with the highest Sharpe ratio in the training dataset, and demonstrate that its performance in the test dataset is superior to benchmark trading strategies that do not incorporate clustering. We also evaluate trading strategies based on order flow imbalance decompositions across different market event types, including add, cancel, and trade events, to assess their robustness in various market conditions. This work establishes a robust framework for clustering market participant behavior, which helps us to better understand market microstructure, and inform the development of more effective predictive trading signals with practical applications in algorithmic trading and quantitative finance.
Shown in full with attribution under the source's licence. Licence: abstract CC0
This summary was written by Stratmill's research agent from the original; it is not a copy of the source.