Skip to content
All library documents

Building a Nautilus Parquet Catalog from Databento Market Data

Code NautilusTrader

Summary

This guide shows how to retrieve historical market data from Databento, save it locally in compressed DBN files, convert it into Nautilus data objects, and store those objects in a Parquet catalog. Its examples cover an E-mini S&P futures order book depth feed and a month of Nasdaq AAPL trades. It also explains that an explicit instrument identifier can skip symbology mapping and speed up loading.

The workflow recommends checking the request cost before retrieving data, avoiding duplicate downloads, and keeping local copies for reuse. It demonstrates reading the downloaded files into a dataframe and querying the completed catalog. The examples illustrate an engineering workflow rather than a trading signal or performance result. The guide notes that catalog creation replaces the existing catalog directory, so running it can delete previously stored catalog data. The examples also depend on Databento access, installed client libraries, and suitable instrument metadata; the order book example is a limited sample rather than a comprehensive dataset.

Key ideas

  • Request a cost estimate and check for an existing local file before downloading historical data.
  • Save retrieved market data in compressed DBN files for reuse.
  • Convert DBN records into Nautilus objects and write them to a Parquet catalog.
  • Providing an instrument identifier can bypass symbology mapping and speed up loading.
  • The example clears the existing catalog directory before creating a new one.

Tags

Full text
# data_catalog_databento.py


```py
# %% [markdown]
# # Data Catalog with Databento
#
# Set up a Nautilus Parquet data catalog with market data from Databento. The
# catalog provides efficient storage and querying for backtests and research.
#
# [View source on GitHub](https://github.com/nautechsystems/nautilus_trader/blob/develop/docs/how_to/data_catalog_databento.py).

# %% [markdown]
# ## Prerequisites
#
# - Python 3.13+
# - [NautilusTrader](https://pypi.org/project/nautilus_trader/) 2.x installed
#   (`pip install -U --pre nautilus_trader`)
# - [databento](https://pypi.org/project/databento/) Python client library (`pip install databento`)
# - [Databento](https://databento.com) account with API key set as `DATABENTO_API_KEY`

# %% [markdown]
# ## Request data
#
# Initialize a Databento historical client. The client reads your API key from
# the `DATABENTO_API_KEY` environment variable by default.

# %%
import databento as db


client = db.Historical()  # Uses the DATABENTO_API_KEY environment variable

# %% [markdown]
# **Every historical streaming request from `timeseries.get_range` incurs a cost (even for the same data), so**:
# - Check the cost before making a request
# - Avoid requesting the same data twice
# - Write responses to disk as zstd compressed DBN files

# %% [markdown]
# Use the metadata [get_cost endpoint](https://databento.com/docs/api-reference-historical/metadata/metadata-get-cost?historical=python&live=python) to quote the cost before each request. Only request data that does not already exist on disk.
#
# The response is in USD, displayed as fractional cents.

# %% [markdown]
# The following request is for a small amount of data (as used in the Databento blog post [Building high-frequency trading signals in Python with Databento and sklearn](https://databento.com/blog/hft-sklearn-python)) to demonstrate the workflow.

# %%
from pathlib import Path

from databento import DBNStore


# %% [markdown]
# We'll prepare a directory for the raw Databento DBN format data, which we'll use for the rest of the tutorial.

# %%
DATABENTO_DATA_DIR = Path("databento")
DATABENTO_DATA_DIR.mkdir(exist_ok=True)

# %%
# Request cost quote (USD) - this endpoint is 'free'
client.metadata.get_cost(
    dataset="GLBX.MDP3",
    symbols=["ES.n.0"],
    stype_in="continuous",
    schema="mbp-10",
    start="2023-12-06T14:30:00",
    end="2023-12-06T20:30:00",
)

# %% [markdown]
# Use the historical API to request the data used in the blog post.

# %%
path = DATABENTO_DATA_DIR / "es-front-glbx-mbp10.dbn.zst"

if not path.exists():
    # Request data
    client.timeseries.get_range(
        dataset="GLBX.MDP3",
        symbols=["ES.n.0"],
        stype_in="continuous",
        schema="mbp-10",
        start="2023-12-06T14:30:00",
        end="2023-12-06T20:30:00",
        path=path,  # <-- Passing a `path` writes the data to disk
    )

# %% [markdown]
# Read the data from disk and convert to a pandas.DataFrame

# %%
data = DBNStore.from_file(path)

df = data.to_df()
df

# %% [markdown]
# ## Write to data catalog
#
# The guide writes the catalog to `catalog/` under the working directory and replaces that directory on each run.

# %%
import shutil
from pathlib import Path

from nautilus_trader.adapters.databento import DatabentoDataLoader
from nautilus_trader.model import InstrumentId
from nautilus_trader.persistence import ParquetDataCatalog


# %%
CATALOG_PATH = Path.cwd() / "catalog"

# Clear if it already exists
if CATALOG_PATH.exists():
    shutil.rmtree(CATALOG_PATH)
CATALOG_PATH.mkdir()

# Create a catalog instance
catalog = ParquetDataCatalog(str(CATALOG_PATH))

# %% [markdown]
# Use a `DatabentoDataLoader` to decode and load the data into Nautilus objects.

# %%
loader = DatabentoDataLoader()

# %% [markdown]
# Passing an `instrument_id` is optional but speeds up loading by skipping symbology mapping. If provided, use the Nautilus `symbol.venue` format (e.g., "ESZ3.GLBX").

# %%
path = DATABENTO_DATA_DIR / "es-front-glbx-mbp10.dbn.zst"

# Option 1 (recommended): Let the loader infer the instrument ID from DBN metadata
depth = loader.load_order_book_depth(filepath=path)

# Option 2: Explicitly specify a valid Nautilus instrument ID (symbol.venue format)
# instrument_id = InstrumentId.from_str("ESZ3.GLBX")  # E-mini S&P December 2023 futures on Globex
# depth = loader.load_order_book_depth(
#     filepath=path,
#     instrument_id=instrument_id,
# )

# %%
# Write data to catalog (this takes ~20 seconds or ~250,000/second for writing MBP-10 at the moment)
catalog.write_order_book_depths(depth)

# %%
# Test reading from catalog
depths = catalog.query_order_book_depths()
len(depths)

# %% [markdown]
# ## Preparing a month of AAPL trades

# %% [markdown]
# Now we'll expand on this workflow by preparing a month of AAPL trades on the Nasdaq exchange using the Databento `trades` schema, which will translate to Nautilus `TradeTick` objects.

# %%
# Request cost quote (USD) - this endpoint is 'free'
client.metadata.get_cost(
    dataset="XNAS.ITCH",
    symbols=["AAPL"],
    schema="trades",
    start="2024-01",
)

# %% [markdown]
# Pass a `path` parameter when requesting historical data to write it to disk.

# %%
path = DATABENTO_DATA_DIR / "aapl-xnas-202401.trades.dbn.zst"

if not path.exists():
    # Request data
    client.timeseries.get_range(
        dataset="XNAS.ITCH",
        symbols=["AAPL"],
        schema="trades",
        start="2024-01",
        path=path,  # <-- Passing a `path` parameter
    )

# %% [markdown]
# Read the data from disk and convert to a pandas.DataFrame

# %%
data = DBNStore.from_file(path)

df = data.to_df()
df

# %% [markdown]
# We'll use an `InstrumentId` of `"AAPL.XNAS"`, where XNAS is the ISO 10383 MIC (Market Identifier Code) for the Nasdaq venue.
#
# Passing an `instrument_id` speeds up loading by skipping symbology mapping.

# %%
instrument_id = InstrumentId.from_str("AAPL.XNAS")

trades = loader.load_trades(
    filepath=path,
    instrument_id=instrument_id,
)

# %% [markdown]
# Here we organize data as one file per month. A file per day works equally well.

# %%
# Write data to catalog
catalog.write_trade_ticks(trades)

# %%
trades = catalog.query_trade_ticks([str(instrument_id)])

# %%
len(trades)

```

Shown in full with attribution under the source's licence. Licence: LGPL-3.0

This summary was written by Stratmill's research agent from the original; it is not a copy of the source.