Separating Order-Book Events from Tick Data by One-Second Bins
Summary
The document describes an exploratory workflow for inspecting full-depth equity market data and separating order-related events by type. It recommends loading a sample file into a data analysis environment, parsing the date and timestamp fields, and checking the available event labels. These include order additions, deletions, cancellations, fills, executions, trades, and crosses, with bid and ask sides identified for many event types.
The example then groups records into one-second time intervals and filters a selected interval for a particular event, such as bid additions. This demonstrates how to isolate event categories and timestamps for analysis. It does not provide a complete classification of which labels represent market orders versus limit orders, nor does it explain how to handle ambiguous events, aggregation, or validation. The method therefore offers a starting point for examining event data rather than a finished order-classification procedure.
Key ideas
- Parse the date and timestamp fields before grouping full-depth records by time.
- Inspect the dataset’s event labels to understand which order-book actions it records.
- Group records into one-second intervals to examine activity at that resolution.
- Filter each interval by event type and side to isolate selected order-book events.
- The example does not fully map event labels to market-order and limit-order categories.
Tags
Full text
# Separate market and limit orders from market depth/tick data
# Separate market and limit orders from market depth/tick data
From the website https://www.algoseek.com/equities/, we can get a sample of the full depth market/tick data. From the paper https://arxiv.org/pdf/1710.03870.pdf page 8, I would like to extract the market orders and limit orders separately with timestamp of 1 second . Is it possible to do such a thing? If so, how?
## Answer by Alexey Golyshev (score 3, accepted)
https://quant.stackexchange.com/a/41173
- download the data
- open Jupyter Notebook
```
import pandas as pd
data = pd.read_csv('IBM.FullDepth.20140128.csv', parse_dates=[['Date', 'Timestamp']])
data['EventType'].unique()
```
> array(['ADD BID', 'ADD ASK', 'DELETE ASK', 'DELETE BID', 'TRADE ASK', 'EXECUTE BID', 'FILL BID', 'TRADE BID', 'FILL ASK', 'EXECUTE ASK', 'CROSS', 'CANCEL ASK', 'CANCEL BID'], dtype=object)
```
grouped = data.groupby(pd.Grouper(key='Date_Timestamp', freq='1s'))
groups = grouped.groups
keys = list(groups.keys())
df=grouped.get_group(keys[0])
df[df.EventType=='ADD BID']
```Shown in full with attribution under the source's licence. Licence: CC BY-SA 4.0 (Stack Exchange)
This summary was written by Stratmill's research agent from the original; it is not a copy of the source.