오더북 재구성을 위한 NASDAQ ITCH 메시지 파싱
코드 Machine Learning for Trading
요약
이 노트북은 바이너리 형식의 NASDAQ TotalView-ITCH 주문별 메시지 데이터를 구조화된 레코드로 디코딩하는 방법을 설명합니다. 메시지 프레이밍과 유형별 구조를 살펴보고 필드를 언패킹하는 방법을 보여줍니다. 자정부터 나노초 단위로 표현된 타임스탬프와 소수점 위치가 암묵적으로 정해진 정수 형식의 가격도 변환합니다. 주요 워크플로는 메시지를 순차적으로 읽고 지원되는 유형을 디코딩해 버퍼에 담은 뒤 메시지 유형별로 분할된 Parquet 파일에 기록합니다. 따라서 전체 세션을 메모리에 올리지 않고 처리할 수 있습니다.
문서는 파싱한 이벤트를 후속 오더북 재구성에 연결하고 개별 주문을 수정하는 메시지와 거래소 수준 이벤트를 구분합니다. 같은 출력 스키마를 생성하는 Python 구현과 Rust 파서도 비교하며, 며칠 분량을 처리하거나 실제 운영에 사용할 때는 후자를 권합니다. 이 노트북은 샘플 워크플로를 설명할 뿐 트레이딩 성과 근거를 제시하지 않습니다. 범위는 한 거래소에 한정되고 지원되지 않는 메시지 유형은 건너뜁니다. 인용된 Rust 속도 비교는 다른 컴퓨터에서 얻은 결과입니다.
핵심 아이디어
- ITCH 프레임에는 길이, 메시지 유형, 해당 유형에 따라 필드가 달라지는 페이로드가 있습니다.
- 개별 주문 단위의 추가·체결·취소·삭제·교체 메시지는 오더북을 재구성하는 데 필요한 이벤트를 제공합니다.
- 분석하기 전에 타임스탬프와 가격을 인코딩된 형식에서 변환해야 합니다.
- 디코딩한 메시지를 버퍼에 모아 일괄 기록하면 큰 세션을 처리할 때 메모리 사용량을 줄일 수 있습니다.
- Python은 읽기와 디버깅에 유리하고, Rust는 더 큰 파싱 작업에 적합합니다.
태그
전문
# 01_itch_parser.py
```py
# ---
# jupyter:
# jupytext:
# formats: ipynb,py:percent
# text_representation:
# extension: .py
# format_name: percent
# format_version: '1.3'
# jupytext_version: 1.19.3
# kernelspec:
# display_name: Python 3 (ipykernel)
# language: python
# name: python3
# ---
# %% [markdown]
# # NASDAQ TotalView-ITCH: Order Book Data Parsing
#
# **Chapter 3: Market Microstructure**
#
# **Docker image**: `ml4t`
#
# ## Purpose
#
# This notebook demonstrates how to parse NASDAQ's TotalView-ITCH binary protocol.
# Understanding MBO (message-by-order) data is foundational for microstructure-based ML features.
#
# ## Learning Objectives
#
# After completing this notebook, you will be able to:
# - Read one ITCH v5.0 message out of a binary file: find where it starts, how long it is,
# and which of the twenty-odd message types it is.
# - Unpack a message's fields with Python's `struct` module, and convert the two encodings
# ITCH uses - a nanosecond offset from midnight, and a price held as an integer with four
# implied decimal places.
# - Run the parse over a full trading day, writing each message type to its own Parquet
# partition so that a day that does not fit in memory still lands on disk.
# - Load an already-parsed day and check that every message type reads back.
# - Say when the Python parser is the right tool and when the Rust one is.
#
# ## Book reference
#
# Section §3.3, *From raw messages to the limit order book* - the binary-parsing-at-scale
# subsection and Table 3.1. §3.2 places ITCH among the other feeds.
#
# ## Cross-References
#
# - **Downstream**: `02_itch_lob_reconstruction` (builds order book from these messages)
# - **Related**: `09_databento_mbo_analysis` (alternative MBO data source)
#
# ## Data Requirements
#
# ITCH sample data can be downloaded using:
# ```bash
# python data/equities/market/microstructure/nasdaq_itch_download.py --list # List available files
# python data/equities/market/microstructure/nasdaq_itch_download.py # Download default sample
# ```
#
# Files are ~5GB compressed from: https://emi.nasdaq.com/ITCH/
# %% [markdown]
# ## ITCH Message Types
#
# The [ITCH v5.0 specification](https://www.nasdaqtrader.com/content/technicalsupport/specifications/dataproducts/NQTVITCHSpecification.pdf) defines 20+ message types:
#
# | Type | Name | Description |
# |------|------|-------------|
# | **S** | System Event | Market open/close events |
# | **R** | Stock Directory | Ticker information and characteristics |
# | **H** | Trading Action | Trading halts, pauses, and resumptions |
# | **Y** | Reg SHO Restriction | Short sale price test restrictions |
# | **L** | Market Participant | Market maker positions |
# | **V** | MWCB Decline Level | Market-wide circuit breaker levels |
# | **W** | MWCB Status | Circuit breaker breach status |
# | **A** | Add Order | New limit order enters the book |
# | **F** | Add Order (MPID) | Same as A, with market participant ID |
# | **E** | Order Executed | Partial/full execution against standing order |
# | **C** | Order Executed w/Price | Execution at different price (hidden orders) |
# | **X** | Order Cancel | Partial cancellation |
# | **D** | Order Delete | Full removal from book |
# | **U** | Order Replace | Modify price/size (cancel + add) |
# | **P** | Trade | Non-displayed execution |
# | **Q** | Cross Trade | Opening/closing cross |
# | **B** | Broken Trade | Trade cancellation |
# | **I** | NOII | Net Order Imbalance Indicator (auction) |
# | **J** | LULD Auction Collar | Limit up-limit down price bands |
# | **K** | IPO Quoting Period | IPO quotation timing |
#
# By combining these messages chronologically, we can reconstruct the order book at any point in time.
# %%
"""NASDAQ TotalView-ITCH: Order Book Data Parsing — parse ITCH binary protocol into structured messages."""
import gzip
import os
import shutil
import struct
from collections import Counter, defaultdict
from datetime import date, datetime
from pathlib import Path
from time import time
import polars as pl
from itch_message_specs import (
FMT_DICT,
MESSAGE_SPECS,
NT_DICT,
flush_to_parquet,
parse_price4,
parse_timestamp,
print_message_formats,
)
from tqdm.auto import tqdm
from data import load_nasdaq_itch
from utils.paths import display_path
# %% tags=["parameters"]
SKIP_PARSING = False
# Messages to read before stopping. None parses the whole session, which is what the
# shipped output is built from; a small number is how you try the parser out first.
MAX_MESSAGES = None
# %% [markdown]
# The parse reads one directory and writes another. `load_nasdaq_itch(get_base_path=True)`
# resolves the parsed-message root from the repository's data configuration rather than a
# path typed here, so the same notebook runs against a local checkout and inside the Docker
# image. The raw binary sits in a `raw/` directory beside it, which is where the download
# script puts it.
#
# `must_exist=False` is what makes a clean start work. For every other notebook in this
# chapter an absent `messages/` directory is the download instruction and the loader is
# right to refuse; this notebook is the one that creates it, and it has to be told where
# before there is anything there. The directory itself is created in Section 4, where the
# parse is about to write: creating it here would leave an empty one behind on a machine
# with no raw feed, and an empty directory is exactly what stops the loader refusing.
# %%
MESSAGE_DIR = load_nasdaq_itch(get_base_path=True, must_exist=False)
ITCH_RAW_DIR = MESSAGE_DIR.parent / "raw"
print(f"Raw ITCH binary (input): {display_path(ITCH_RAW_DIR)}")
print(f"Parsed messages (output): {display_path(MESSAGE_DIR)}")
raw_present = ITCH_RAW_DIR.exists() and next(ITCH_RAW_DIR.iterdir(), None) is not None
print(f"Raw binary present: {raw_present}")
if not raw_present:
print(
" Fetch it first: "
"uv run python data/equities/market/microstructure/nasdaq_itch_download.py"
)
def message_type_dirs() -> list[Path]:
"""Parsed message-type subdirectories, one uppercase letter each.
Empty before Section 4 has run: `messages/` itself does not exist on a machine that
has never parsed, so every cell that lists message types goes through here.
"""
if not MESSAGE_DIR.exists():
return []
return [
d
for d in sorted(MESSAGE_DIR.iterdir())
if d.is_dir() and len(d.name) == 1 and d.name.isupper()
]
# %% [markdown]
# ## 1. Message Specifications
#
# Each ITCH message has a fixed binary structure. The format is defined in `itch_message_specs.py`
# using Python's `struct` module. Format codes:
# - `H` = unsigned short (2 bytes)
# - `I` = unsigned int (4 bytes)
# - `Q` = unsigned long long (8 bytes)
# - `s` = char (1 byte), `Ns` = N chars
# - `>` = big-endian byte order
# %%
# Show message formats (loaded from utils/itch_message_specs.py)
print_message_formats()
# %% [markdown]
# ## 2. Binary Parsing Example
#
# Let's demonstrate how binary parsing works by creating and parsing a sample Add Order message.
# %% [markdown]
# Packing a message by hand is the quickest way to see the layout. An Add Order carries the
# format `>HH6sQsI8sI`: big-endian, two unsigned shorts, a six-byte timestamp, an eight-byte
# order reference, a one-character side, the share count, an eight-character padded ticker
# and the price. Unpacking it below returns exactly these fields.
# %%
sample_add_order = struct.pack(
">HH6sQsI8sI",
1234, # stock_locate
5678, # tracking_number
b"\x00\x00\x00\x00\x00\x01", # timestamp (1 nanosecond)
9876543210, # order_reference_number
b"B", # buy_sell_indicator
100, # shares
b"AAPL ", # stock (padded to 8 chars)
1500000, # price: four implied decimals, so this is $150
)
print(f"Raw Add Order message ({len(sample_add_order)} bytes):")
print(f" Hex: {sample_add_order.hex()}")
# %%
# Parse the binary data using our struct format
parsed = struct.unpack(FMT_DICT["A"], sample_add_order)
add_order = NT_DICT["A"]._make(parsed)
print("Parsed Add Order Message:")
print("-" * 40)
for field, value in add_order._asdict().items():
if isinstance(value, bytes):
value = value.decode("ascii").strip()
print(f" {field:25}: {value}")
# %%
# Apply conversions using helper functions from utils.itch_message_specs
ts_ns = parse_timestamp(add_order.timestamp)
price = parse_price4(add_order.price)
print(f"Timestamp: {ts_ns:,} nanoseconds = {ts_ns / 1e9:.9f} seconds after midnight")
print(f"Price: ${price:.4f}")
# %% [markdown]
# ## 3. What the Pipeline Holds Right Now
#
# Before parsing anything, take stock: the raw binary is the input, the per-message-type
# Parquet directories are the output, and either can be absent. On a first run the parsed
# side is empty, and it stays empty until Section 4 writes it. The parsed messages are
# loaded in Section 5, after there is something to load.
# %%
# Check what data is available locally
print("ITCH Data Pipeline Status:")
print("-" * 50)
# Step 1: Raw binary from download
raw_files = []
if ITCH_RAW_DIR.exists():
raw_files = list(ITCH_RAW_DIR.glob("*.gz")) + list(ITCH_RAW_DIR.glob("*.bin"))
print(f"Raw binary files: {len(raw_files)}")
for f in raw_files:
print(f" {f.name} ({f.stat().st_size / 1e9:.2f} GB)")
# Step 2: Parsed messages (single uppercase letter = message type)
parsed_types = message_type_dirs()
parsed_with_data = [d for d in parsed_types if list(d.glob("*.parquet"))]
print(f"Parsed message types: {len(parsed_with_data)}")
for msg_dir in parsed_with_data:
name = MESSAGE_SPECS.get(msg_dir.name, {}).get("name", "Unknown")
n_files = len(list(msg_dir.glob("*.parquet")))
print(f" {msg_dir.name} ({name}): {n_files} files")
# %% [markdown]
# ## 4. Full Parser Implementation
#
# The parser below is written to be read. It processes one message at a time in Python, which
# is what makes each step visible and also what makes a full trading day take about twenty
# minutes. Section 6 covers the Rust parser, which emits the same Parquet schema and is what
# you would run over many days.
# %% [markdown]
# ### Parser Helpers
#
# The parse splits into three functions. `read_frame` takes the next message off the file
# and hands back its type and its raw bytes; `decode_message` turns those bytes into a
# dictionary of Python values; and `parse_itch_file` runs the loop, buffering decoded
# messages and writing them out in batches.
# %%
def read_frame(f, pbar) -> tuple[str, bytes] | None:
"""Read one ITCH message frame: 2-byte length + 1-byte type + payload.
Returns (msg_type, payload) on success, or None on EOF/truncation.
"""
# 2-byte big-endian length prefix (message size including type byte)
length_bytes = f.read(2)
if len(length_bytes) < 2:
return None
pbar.update(2)
msg_size = int.from_bytes(length_bytes, "big")
# 1-byte message type
msg_type_byte = f.read(1)
if len(msg_type_byte) < 1:
print(f"\nWarning: Truncated message at byte {f.tell()}, expected type byte")
return None
pbar.update(1)
msg_type = msg_type_byte.decode("ascii")
# Payload (msg_size includes type byte, so payload is msg_size - 1)
payload = f.read(msg_size - 1)
if len(payload) < msg_size - 1:
print(f"\nWarning: Truncated payload for message type {msg_type}")
return None
pbar.update(msg_size - 1)
return msg_type, payload
# %% [markdown]
# Decode binary payload into a Python dict, converting raw timestamp bytes
# to nanosecond integers and byte strings to stripped ASCII.
# %%
def decode_message(msg_type: str, payload: bytes) -> dict | None:
"""Unpack binary payload into a dict, converting timestamps and strings.
Returns parsed message dict, or None on struct error.
"""
try:
parsed = struct.unpack(FMT_DICT[msg_type], payload)
msg = NT_DICT[msg_type]._make(parsed)._asdict()
except struct.error:
return None
# Convert timestamp: nanoseconds since midnight
if "timestamp" in msg:
msg["timestamp"] = int.from_bytes(msg["timestamp"], "big")
# Decode string fields
for field, value in msg.items():
if isinstance(value, bytes):
msg[field] = value.decode("ascii").strip()
return msg
# %% [markdown]
# The main parser reads the binary file sequentially, buffering decoded messages
# and flushing to Parquet periodically to bound memory usage.
# %%
def parse_itch_file(
itch_file: Path,
trading_day: date,
output_dir: Path,
max_buffered_messages: int = 10_000_000,
max_messages: int | None = None,
) -> dict[str, int]:
"""Parse ITCH binary file and store messages as Parquet.
Args:
itch_file: Path to binary ITCH file (.bin, not .gz).
trading_day: Trading date for timestamp construction.
output_dir: Directory for Parquet output (one subdir per message type).
max_buffered_messages: Flush threshold (total buffered messages).
max_messages: Optional limit for testing.
Returns:
Dictionary with message type counts.
"""
# Midnight timestamp for the trading day (ITCH timestamps are nanoseconds offset)
base_ts = datetime(trading_day.year, trading_day.month, trading_day.day)
file_counters: dict[str, int] = defaultdict(int)
buffers = defaultdict(list)
counts = Counter()
file_size = itch_file.stat().st_size
start_time = time()
with (
itch_file.open("rb") as f,
tqdm(total=file_size, desc="Parsing ITCH", unit="B", unit_scale=True) as pbar,
):
while True:
if max_messages and sum(counts.values()) >= max_messages:
print(f"\nLimit reached: {max_messages:,} messages")
break
frame = read_frame(f, pbar)
if frame is None:
break
msg_type, payload = frame
counts[msg_type] += 1
if msg_type not in FMT_DICT:
continue
msg = decode_message(msg_type, payload)
if msg is None:
continue
# Check for end of messages
if msg_type == "S" and msg.get("event_code") == "C":
print("\nEnd of Messages")
flush_to_parquet(buffers, output_dir, base_ts, file_counters)
break
buffers[msg_type].append(msg)
# Periodic flush
if sum(len(v) for v in buffers.values()) >= max_buffered_messages:
flush_to_parquet(buffers, output_dir, base_ts, file_counters)
# Final flush
if any(buffers.values()):
flush_to_parquet(buffers, output_dir, base_ts, file_counters)
elapsed = time() - start_time
total = sum(counts.values())
print(f"Parsed {total:,} messages in {elapsed:.1f}s ({total / elapsed:,.0f} msg/s)")
return dict(counts)
# %%
# Locate ITCH data file (skip if SKIP_PARSING is set)
if SKIP_PARSING:
print("SKIP_PARSING=True: skipping ITCH binary parsing (uses pre-parsed data)")
itch_file = None
counts = {}
else:
# Clear any existing parsed data to avoid schema conflicts
# (Different parser versions may produce different schemas)
# Set ITCH_KEEP_EXISTING=1 to skip cleanup and use existing data
clear_existing = os.environ.get("ITCH_KEEP_EXISTING", "0") != "1"
if clear_existing and MESSAGE_DIR.exists() and list(MESSAGE_DIR.glob("*/part-*.parquet")):
print(f"Clearing existing parsed data in {MESSAGE_DIR}")
print(" (Set ITCH_KEEP_EXISTING=1 to keep existing data)")
shutil.rmtree(MESSAGE_DIR)
MESSAGE_DIR.mkdir(parents=True, exist_ok=True)
# Find ITCH file (compressed or uncompressed)
gz_files = list(ITCH_RAW_DIR.glob("*.gz")) if ITCH_RAW_DIR.exists() else []
bin_files = list(ITCH_RAW_DIR.glob("*.bin")) if ITCH_RAW_DIR.exists() else []
if not gz_files and not bin_files:
raise FileNotFoundError(
f"No raw ITCH binary found at {ITCH_RAW_DIR}.\n"
"Download first:\n"
" uv run python data/equities/market/microstructure/nasdaq_itch_download.py"
)
# %%
# Decompress if needed and extract trading date
if not SKIP_PARSING and (gz_files or bin_files):
# Prefer uncompressed, otherwise decompress
if bin_files:
itch_file = bin_files[0]
print(f"Found uncompressed: {itch_file.name}")
else:
gz_file = gz_files[0]
itch_file = gz_file.with_suffix(".bin")
if not itch_file.exists():
print(f"Decompressing {gz_file.name}...")
with gzip.open(gz_file, "rb") as f_in, open(itch_file, "wb") as f_out:
shutil.copyfileobj(f_in, f_out)
print(f"Created: {itch_file.name} ({itch_file.stat().st_size / 1e9:.1f} GB)")
else:
print(f"Found: {itch_file.name}")
# Extract trading date from filename (format: MMDDYYYY.NASDAQ_ITCH50.bin)
date_str = itch_file.stem.split(".")[0]
if len(date_str) == 8 and date_str.isdigit():
trading_day = date(int(date_str[4:8]), int(date_str[:2]), int(date_str[2:4]))
else:
trading_day = date(2020, 1, 30) # Fallback
print(f"Trading day: {trading_day}")
# Parse ITCH file (full-day parse: ~22 min on the reference machine)
if itch_file and itch_file.exists():
# Created here, not where the path is resolved: an empty messages/ directory
# left behind by a run that found no raw binary would stop load_nasdaq_itch
# raising for every later notebook in the chapter.
MESSAGE_DIR.mkdir(parents=True, exist_ok=True)
counts = parse_itch_file(
itch_file=itch_file,
trading_day=trading_day,
output_dir=MESSAGE_DIR,
max_messages=MAX_MESSAGES,
)
print("\nMessage counts:")
for msg_type, count in sorted(counts.items(), key=lambda x: -x[1]):
name = MESSAGE_SPECS.get(msg_type, {}).get("name", "Unknown")
print(f" {msg_type} ({name:25}): {count:>10,}")
# %% [markdown]
# ## 5. Message Type Analysis
#
# The parse has written one Parquet directory per message type, so the store can now be
# read back. The cell below asserts on trade messages (`P`) because that is the type it
# goes on to load. A full session always contains them, so an empty `P` directory means
# the parse did not finish rather than that this day happened to have no trades. That
# reasoning holds only for a full parse: a run with `MAX_MESSAGES` set can stop before
# the first trade prints, so there the absence is reported and the remaining cells work
# with whichever types the trial did write.
# %%
trade_dir = MESSAGE_DIR / "P"
trade_files = sorted(trade_dir.glob("*.parquet")) if trade_dir.exists() else []
if not trade_files and MAX_MESSAGES is not None:
print(
f"No trade messages under {display_path(trade_dir)}: the parse stopped after "
f"{MAX_MESSAGES:,} messages and a `P` need not appear that early. "
"Set MAX_MESSAGES to None to read the whole session."
)
else:
assert trade_files, (
f"No parsed trade messages at {trade_dir}.\n"
"Section 4 writes them from the raw binary. If SKIP_PARSING is True, it was skipped\n"
"and the store has to have been written by an earlier run or by the Rust parser of\n"
"Section 6. If the raw binary is missing, fetch it first:\n"
" uv run python data/equities/market/microstructure/nasdaq_itch_download.py"
)
trades = pl.read_parquet(trade_dir / "*.parquet")
print(f"Loaded {len(trades):,} trade messages")
print(f"Columns: {trades.columns}")
# %%
# Message type distribution — use lazy scan to count without loading all data
written = [d for d in message_type_dirs() if list(d.glob("*.parquet"))]
print("Parsed Message Types:")
print("-" * 50)
if not written:
print(" Nothing parsed yet - Section 4 writes these from the raw binary.")
for msg_dir in written:
count = pl.scan_parquet(msg_dir / "*.parquet").select(pl.len()).collect().item()
name = MESSAGE_SPECS.get(msg_dir.name, {}).get("name", "Unknown")
print(f" {msg_dir.name} ({name:25}): {count:>12,} messages")
# %%
# Schema compatibility check — verify we can read each message type
written = [d for d in message_type_dirs() if list(d.glob("*.parquet"))]
print("Schema Compatibility Check:")
print("-" * 50)
if not written:
print(" Nothing parsed yet - Section 4 writes these from the raw binary.")
for msg_dir in written:
try:
sample = pl.scan_parquet(msg_dir / "*.parquet").head(5).collect()
name = MESSAGE_SPECS.get(msg_dir.name, {}).get("name", "Unknown")
print(f" [OK] {msg_dir.name} ({name:25}): cols={list(sample.columns)[:4]}...")
except Exception as e:
print(f" [FAIL] {msg_dir.name}: {e}")
# %% [markdown]
# ## 6. Production Parsing with Rust
#
# The Python parser above decodes one message per loop iteration, and a trading day holds a
# few hundred million of them. The same protocol parsed in Rust reads the file through a
# memory map and unpacks each message without copying it first, so it neither pays the
# per-message interpreter overhead nor holds the decoded messages in memory.
#
# **Repository**: [github.com/ml4t/itch-parser](https://github.com/ml4t/itch-parser)
#
# The figures in the book's Table 3.1, for the same 13 GB file on one machine, are about
# twenty-three minutes and roughly 8 GB of memory in Python against under five minutes and
# under 500 MB in Rust. Wall-clock timings move with disk and CPU, so read them as an order
# of magnitude rather than a ratio: the gap is large enough that it decides which parser you
# reach for, and not stable enough to quote to a decimal place. The throughput this notebook
# printed above is the Python side of the same comparison, measured on the machine that ran
# it.
#
# ### Installation
#
# ```bash
# # Clone the repository
# git clone https://github.com/ml4t/itch-parser.git
# cd itch-parser
#
# # Build release binary
# cargo build --release
# ```
#
# ### Usage
#
# ```bash
# # Parse ITCH file (works with .gz or uncompressed)
# ./target/release/itch_parser <input_file> <output_dir> <MMDDYYYY>
#
# # Example
# ./target/release/itch_parser data/01302020.NASDAQ_ITCH50.gz ./messages 01302020
# ```
#
# Output is identical Parquet files partitioned by message type, compatible with
# the Python code in this notebook and downstream analysis.
#
# ### When to Use Which
#
# | Use Case | Recommendation |
# |----------|---------------|
# | Learning the protocol | Python (this notebook) |
# | Debugging parse issues | Python |
# | Processing a single day | Either |
# | Multi-day backtesting | **Rust** |
# | Production pipeline | **Rust** |
#
# Both parsers write the same Parquet schema, so a day parsed either way feeds every
# notebook that follows without change.
# %% [markdown]
# ## Key Takeaways
#
# 1. **The protocol is message-by-order.** Every message that changes the book names the
# individual order it acts on, stamped to the nanosecond, which is what makes book
# reconstruction possible at all. Session-level messages - `S` for market events, `R`
# for the stock directory, `Q` for the auction crosses - carry no order reference and
# describe the venue rather than a single order.
# 2. **Seven message types carry the book.** `A` and `F` add an order, `E` and `C` execute
# against one, `X` cancels part of one, `D` deletes one, `U` replaces one.
# 3. **Numbers arrive encoded.** Prices are integers with four implied decimal places, and
# timestamps are nanoseconds since midnight, so both need converting before use.
# 4. **Parse in batches, not in one pass.** Buffering by message type and flushing to
# Parquet is what keeps a day that does not fit in memory from having to.
# 5. **Which parser depends on how many days you need.** Python reads clearly and takes
# about twenty minutes per day; Rust emits the same schema in a fraction of that.
#
# ### Known limitations
#
# - The parse covers one venue. NASDAQ ITCH sees NASDAQ-routed activity, not the
# consolidated tape, so counts here are a venue's share of a symbol's trading rather than
# all of it.
# - Message types outside `FMT_DICT` are counted and skipped rather than decoded.
# - The Rust timings quoted above were measured elsewhere, on one machine; this notebook
# times only its own Python parse.
#
# ### Next Steps
#
# - **Order Book Reconstruction**: `02_itch_lob_reconstruction`
# - **Trading Activity Overview**: `05_itch_trading_activity` (includes E/C enrichment)
#
# ---
#
# ## Reference
#
# Bouchaud, J.-P., Bonart, J., Donier, J., & Gould, M. (2018).
# *Trades, Quotes and Prices: Financial Markets Under the Microscope*.
# Cambridge University Press.
# [https://doi.org/10.1017/9781009028943](https://doi.org/10.1017/9781009028943)
```출처의 라이선스에 따라 출처를 표시하고 전문을 공개합니다. 라이선스: MIT
이 요약은 원문을 바탕으로 Stratmill의 리서치 에이전트가 작성했으며, 원문을 복사한 것이 아닙니다.