Skip to content
All library documents

Managing Parquet Market Data for Research and Backtesting

Article NautilusTrader

Summary

This guide explains how NautilusTrader stores and accesses market data through a Parquet catalog backed by a Rust storage layer. It covers local and cloud storage, timestamp precision, compression choices, file organization, typed data queries, and configuration for loading data into backtests. It also describes writing market data, consolidating files by time period, deleting ranges, and promoting records staged in Feather files.

The material is operational rather than a trading strategy or performance study. It gives configuration examples and documents constraints, including required ordering and identity for writes, canonical directory layouts, and the possibility that some SQL readers truncate sub-microsecond timestamps. Overlapping writes are rejected by default, while deletion is permanent. Backtest results also depend on reading the expected catalog layout and selecting the intended instruments and time bounds. The guide does not compare data sources or assess whether a strategy is profitable; its value is in helping researchers manage data consistently for analysis and simulation.

Key ideas

  • The catalog stores typed market data in Parquet files and supports local paths and several remote storage backends.
  • Timestamp columns retain nanosecond precision and UTC metadata in Arrow-aware readers, though some SQL readers can lose precision.
  • Writers expect each call to contain one data identity with records ordered by initialization time.
  • Backtest data configurations select data types, instruments or bars, time bounds, filters, and storage settings.
  • Overlapping writes are rejected by default, and deleting catalog data cannot be undone.

Tags

Full text
# Data catalog


# Data catalog

The data catalog stores NautilusTrader data in [Parquet](https://parquet.apache.org) files for
backtesting, live trading, and research.

## Overview and architecture

`ParquetDataCatalog` is the Python interface to the Rust catalog and DataFusion query engine.
The Rust model and persistence crates define the Arrow schemas for built-in data. Registered custom
data supplies its schema and encode/decode handlers at runtime.

Instant timestamps use `Timestamp(Nanosecond, Some("UTC"))`; durations remain integers. Arrow readers
and Nautilus queries preserve the nanoseconds and UTC annotation. SQL readers that map these columns
to microsecond-precision `TIMESTAMPTZ` can truncate sub-microsecond values.

Parquet provides compressed columnar storage and cross-language access. The catalog stores these
files under one root without requiring a separate database service. A local path or object-store
URI selects the storage backend.

## Initializing

Pass a local path or URI as the first constructor argument:

```python
from pathlib import Path

from nautilus_trader.persistence import ParquetDataCatalog


CATALOG_PATH = Path.cwd() / "catalog"
catalog = ParquetDataCatalog(str(CATALOG_PATH))
```

## Filesystem protocols and storage options

The catalog accepts the storage protocols supported by its Rust object-store backend.

### Supported filesystem protocols

| Storage              | URI schemes        | Common option keys                                                        |
| -------------------- | ------------------ | ------------------------------------------------------------------------- |
| Local filesystem     | Plain path, `file` | None.                                                                     |
| Amazon S3            | `s3`               | `region`, `access_key_id`, `secret_access_key`, `endpoint_url`.           |
| Google Cloud Storage | `gs`, `gcs`        | `service_account_path`, `service_account_key`, `application_credentials`. |
| Azure Blob Storage   | `az`, `abfs`       | `account_name`, `account_key`, `sas_token`.                               |
| HTTP or WebDAV       | `http`, `https`    | `timeout`, `connect_timeout`, `proxy_url`.                                |

Option keys are `object_store` configuration keys for the URI scheme:

- S3, GCS, and Azure keys also take a prefixed form, such as `aws_region`,
  `google_application_credentials`, or `azure_storage_sas_token`.
- Client options such as `timeout` apply to every remote scheme. Set `allow_http` to `true` to
  connect over plain `http://`.
- S3 also accepts the legacy `key` and `secret` names.
- An unknown key fails with an error that names it.

Pass credentials and other backend settings through `storage_options`:

```python
catalog = ParquetDataCatalog(
    "s3://my-bucket/nautilus-data/",
    storage_options={
        "access_key_id": "your-key",
        "secret_access_key": "your-secret",
        "region": "us-east-1",
    },
)

azure_catalog = ParquetDataCatalog(
    "abfs://container@account.dfs.core.windows.net/nautilus-data/",
    storage_options={"account_key": "your-account-key"},
)
```

The base path of a remote URI, after the bucket, container, or host, must not contain spaces,
non-ASCII characters, or characters that object-store paths encode, such as `~`, `%`, `#`, `?`,
`^`, or `|`. The catalog rejects such a URI when it opens, with an error that names it.

## Compression and row groups

`DataCatalogConfig` sets how a configured catalog reads and writes data files:

| Field                | Controls                                    | Default          |
| -------------------- | ------------------------------------------- | ---------------- |
| `batch_size`         | Rows per batch the catalog reads and writes | 10,000           |
| `compression`        | Codec name for written data files           | `zstd` (level 1) |
| `max_row_group_size` | Maximum rows per written row group          | 131,072          |

```python
from nautilus_trader.config import DataCatalogConfig


catalog = DataCatalogConfig(path="./catalog", compression="snappy", max_row_group_size=65_536)
```

`compression` accepts `uncompressed`, `snappy`, `gzip`, `brotli`, `lz4`, `lz4_raw`, or `zstd`,
ignoring case. Construction fails for any other name and for a zero `batch_size` or
`max_row_group_size`. `params` carries only options for an external catalog backend, and the
Parquet catalog rejects any `params` key.

`ParquetDataCatalog` takes the same settings as `batch_size`, `max_row_group_size`, and a numeric
`compression` code: `0` (uncompressed), `1` (Snappy), `2` (gzip), `4` (Brotli), `5` (LZ4), or `6`
(zstd).

:::info
`lz4` and `lz4_raw` name the same codec: the catalog writes Parquet `LZ4_RAW`. The Parquet format
deprecates the older Hadoop-framed `LZ4` codec, so the catalog never writes it, and compression code
`5` writes `LZ4_RAW` even though Parquet numbers `LZ4_RAW` as `7`. Files written earlier with the
Hadoop-framed codec remain readable. LZO is not supported because the Parquet writer cannot produce
it, so the name `lzo` and code `3` fail at construction.
:::

## Writing data

Use the writer for the concrete data type. Instrument definitions and custom data have separate
writers.

```python
catalog.write_instruments([instrument])
catalog.write_quote_ticks(quote_ticks)

catalog.write_trade_ticks(
    trade_ticks,
    start=1704067200000000000,
    end=1704153600000000000,
)

catalog.write_bars(bars, skip_disjoint_check=True)
```

The built-in market-data writers are:

- `write_quote_ticks`
- `write_trade_ticks`
- `write_order_book_deltas`
- `write_order_book_depths`
- `write_bars`
- `write_mark_price_updates`
- `write_index_price_updates`
- `write_option_greeks`

Each writer accepts optional `start` and `end` overrides as UNIX nanoseconds. The data in one call
must have one identity, such as one instrument ID or bar type, and must be ordered by `ts_init`.

## File naming and data organization

The catalog names files from their timestamp range with the pattern
`{start_timestamp}_{end_timestamp}.parquet`. It converts each ISO 8601 timestamp to a filename-safe
form by replacing `:` and `.` with `-`.

Built-in data is organized in directories by data type and identifier. For instrument IDs and bar
types, the catalog removes `/` and replaces `^` with `_` when creating the URI-safe directory name:

```text
catalog/
├── data/
│   ├── quotes/
│   │   └── EURUSD.SIM/
│   │       └── 2024-01-01T00-00-00-000000000Z_2024-01-01T23-59-59-999999999Z.parquet
│   └── trades/
│       └── BTCUSD.BINANCE/
│           └── 2024-01-01T00-00-00-000000000Z_2024-01-01T23-59-59-999999999Z.parquet
```

Custom data uses `data/custom/<type_name>/` with optional identifier path segments.

Instruments use one directory per concrete instrument class, such as `data/currency_pair/` or
`data/equity/`, rather than a shared `instruments` directory.

Catalog queries, and therefore backtests, read only these canonical directories. Legacy layouts
such as `data/trade_tick/`, `data/quote_tick/`, and the Python-written
`data/custom_<snake_case>/` are not read at query time. To upgrade a catalog that still uses a
legacy layout, convert it with
[nautilus catalog migrate-parquet](../../how_to/migrate_parquet_catalog.md).

:::warning[Overlapping writes]
By default, overlapping writes raise an `OSError` to maintain data integrity.
Set `skip_disjoint_check=True` only when the overlap is intentional.
:::

## Reading data

Use a typed query when the expected return type is known. `start` and `end` are UNIX nanoseconds:

```python
quotes = catalog.query_quote_ticks(
    identifiers=["EUR/USD.SIM"],
    start=1704067200000000000,
    end=1704153600000000000,
)

trades = catalog.query_trade_ticks(
    identifiers=["BTC/USD.BINANCE"],
    start=1704067200000000000,
    end=1704153600000000000,
)
```

## `BacktestDataConfig`: backtest data

`BacktestDataConfig` defines the catalog data that a `BacktestNode` loads for one run.

### Core parameters

- `data_type` is a `NautilusDataType` value: `QuoteTick`, `TradeTick`, `Bar`, `OrderBookDelta`,
  `OrderBookDepth`, `MarkPriceUpdate`, `IndexPriceUpdate`, `FundingRateUpdate`, `InstrumentStatus`,
  `OptionGreeks`, `InstrumentClose`, or `Instrument`. `Instrument` loads every instrument class the
  catalog holds for the selected identifiers.
- `catalog_path` identifies the catalog root.
- One of `instrument_id`, `instrument_ids`, or `bar_types` is required.
- `start_time` and `end_time` are optional UNIX nanosecond bounds.
- `filter_expr` is an optional DataFusion SQL predicate.
- `catalog_fs_protocol` prefixes `catalog_path` for remote storage.
- `catalog_fs_rust_storage_options` supplies the Rust backend options. If it is unset,
  `BacktestNode` falls back to `catalog_fs_storage_options`.
- For bars, `bar_spec` combines with the instrument ID to select an `EXTERNAL` bar type. Explicit
  `bar_types` can select internal, external, or composite bars.
- `optimize_file_loading` registers whole directories when possible.
- `batch_deltas` (default `True`) replays `OrderBookDelta` data as `OrderBookDeltas` batches closed
  by `F_LAST`. See [order book delta replay](../backtesting/apis-and-runs.md#order-book-delta-replay).

### Basic usage examples

```python
from nautilus_trader.config import BacktestDataConfig
from nautilus_trader.model import BarAggregation
from nautilus_trader.model import BarSpecification
from nautilus_trader.model import InstrumentId
from nautilus_trader.model import NautilusDataType
from nautilus_trader.model import PriceType

quote_data = BacktestDataConfig(
    data_type=NautilusDataType.QuoteTick,
    catalog_path="/path/to/catalog",
    instrument_id=InstrumentId.from_str("EUR/USD.SIM"),
    start_time=1704067200000000000,
    end_time=1704153600000000000,
)

trade_data = BacktestDataConfig(
    data_type=NautilusDataType.TradeTick,
    catalog_path="/path/to/catalog",
    instrument_ids=[
        InstrumentId.from_str("BTC/USD.BINANCE"),
        InstrumentId.from_str("ETH/USD.BINANCE"),
    ],
)

bar_data = BacktestDataConfig(
    data_type=NautilusDataType.Bar,
    catalog_path="/path/to/catalog",
    instrument_id=InstrumentId.from_str("AAPL.NASDAQ"),
    bar_spec=BarSpecification(5, BarAggregation.MINUTE, PriceType.LAST),
)
```

This bar config selects `AAPL.NASDAQ-5-MINUTE-LAST-EXTERNAL`.

### Cloud storage and filtering

```python
book_data = BacktestDataConfig(
    data_type=NautilusDataType.OrderBookDelta,
    catalog_path="my-bucket/nautilus-data",
    catalog_fs_protocol="s3",
    catalog_fs_rust_storage_options={
        "access_key_id": "your-access-key",
        "secret_access_key": "your-secret-key",
        "region": "us-east-1",
    },
    instrument_id=InstrumentId.from_str("BTC/USD.COINBASE"),
    filter_expr="ts_init >= 1704067200000000000",
)
```

### Integration with BacktestRunConfig

Pass the data configurations to `BacktestRunConfig`:

```python
from decimal import Decimal

from nautilus_trader.config import BacktestDataConfig
from nautilus_trader.config import BacktestRunConfig
from nautilus_trader.config import BacktestVenueConfig
from nautilus_trader.execution import MakerTakerFeeModel
from nautilus_trader.model import AccountType
from nautilus_trader.model import BookType
from nautilus_trader.model import InstrumentId
from nautilus_trader.model import NautilusDataType
from nautilus_trader.model import OmsType

data_configs = [
    BacktestDataConfig(
        data_type=NautilusDataType.QuoteTick,
        catalog_path="/path/to/catalog",
        instrument_id=InstrumentId.from_str("EUR/USD.SIM"),
    ),
]

run_config = BacktestRunConfig(
    venues=[
        BacktestVenueConfig(
            name="SIM",
            oms_type=OmsType.HEDGING,
            account_type=AccountType.MARGIN,
            book_type=BookType.L1_MBP,
            starting_balances=["1_000_000 USD"],
            fee_model=MakerTakerFeeModel(
                maker_rate=Decimal("0"),
                taker_rate=Decimal("0"),
            ),
        ),
    ],
    data=data_configs,
    start=1704067200000000000,
    end=1704153600000000000,
)
```

### Data loading process

When a backtest runs, the `BacktestNode` processes each `BacktestDataConfig`:

1. Create a `ParquetDataCatalog` from the configuration.
1. Load the required instrument definitions while building the engine.
1. Build and run a DataFusion query from the configuration fields.
1. Sort merged data by `ts_init` and add it to the backtest engine.

## Direct catalog access

Use `ParquetDataCatalog` to query or write a catalog directly. Use `BacktestDataConfig` when a
`BacktestNode` should load catalog data for a run. `LiveNodeConfig` has no counterpart for loading
catalog data; request historical data through a configured data client or query the catalog
directly. Its `streaming` field configures feather writing only.

## Querying and filtering

The generic query takes a catalog directory name such as `quotes`, `trades`, or `bars`. Use it when
you need the `files` or `optimize_file_loading` controls:

```python
catalog.query(
    data_type=NautilusDataType.QuoteTick,
    identifiers=["EUR/USD.SIM"],
    start=1704067200000000000,
    end=1704153600000000000,
    where_clause="ts_event <= ts_init",
    files=None,
)
```

Typed methods such as `query_quote_ticks`, `query_trade_ticks`, and `query_bars` return the concrete
model type. `query_custom_data` resolves custom decoders through the runtime registry. `query`, the
typed market-data query methods, and `query_custom_data` use UNIX nanosecond time bounds and accept
a DataFusion SQL `where_clause`.

:::warning[Time-zone database mismatch]
With the current `Cargo.lock`, DataFusion SQL temporal functions resolve named time zones with the
transitive `chrono-tz` 0.10.4 database (IANA 2025b). Rust core time-zone operations use Jiff 0.2.35
with its bundled IANA 2026c database. Zone results can differ when zone rules change or historical
data is corrected after 2025b until DataFusion migrates.

If RustSec files unmaintained advisories for `chrono` or `chrono-tz`, maintain matching documented
ignores in `.cargo/audit.toml` and `deny.toml` until DataFusion migrates.
:::

## Catalog operations

Catalog operations rename, consolidate, or delete data files. Each operation takes a type selector:
a `NautilusDataType`, a `NautilusRecordType`, or a `NautilusInstrumentType`.
`NautilusDataType.Instrument` covers every instrument class; a `NautilusInstrumentType` targets
one. `delete_data_range(...)` takes a `NautilusDataType` alone, because only data families support
ranged deletes.

### Reset file names

Reset Parquet file names to match their content timestamps so filename-based filtering remains
accurate. `reset_all_file_names()` processes the entire catalog; `reset_data_file_names(...)`
targets a data path. Supply an instrument ID for data types partitioned by instrument. Without one,
the operation recursively reads the type directory and renames each file within its own instrument
directory, checking every directory's intervals before renaming any file.

```python
catalog.reset_all_file_names()
catalog.reset_data_file_names(NautilusDataType.QuoteTick, "EUR/USD.SIM")
catalog.reset_data_file_names(NautilusDataType.TradeTick, "BTC/USD.BINANCE")
```

### Recover from overlapping file names

Overlapping file names block writes to the affected directory. A rejected coverage extension never
renames files, so overlap points to an earlier rename, a manual move, or a concurrent writer. When
the file contents are still disjoint, recover the directory with a filename reset:

1. Stop all writers to the catalog.
1. Back up the affected directory. The reset renames files one at a time, so an I/O failure
   partway through can leave some files renamed. On object stores that rename by copying then
   deleting, a failure can also leave a file under both names.
1. Run `reset_data_file_names(...)` for the affected path. For custom data types, use
   `reset_all_file_names()` instead, which covers every leaf directory including custom layouts.
1. If the reset reports that a new name is held by another file, rename that file to an unused
   interval name outside the data range, then run the reset again. Repeat until it succeeds.
1. Confirm each renamed file matches its content range, then write a later disjoint interval to
   confirm the directory accepts writes again.

The reset processes one directory at a time. Before renaming any file in a directory, it reads
each file's `ts_init` range. It fails when the content ranges overlap or when a file's new name is
the current name of another file, and it then leaves every file name in that directory unchanged.
Directories reset before the failing one keep their new names.

Use filename reset only for filename-only damage. When file contents overlap, resetting names
cannot reconcile the data, so the reset fails; rebuild the affected range from source through a
separately validated process instead.

### Consolidate catalog

Combine small Parquet files to reduce file count and query overhead.
With no bounds, `consolidate_catalog()` processes each leaf data directory in the catalog.
`consolidate_data(...)` operates on one directory; supply an instrument ID for data types
partitioned by instrument.

```python
catalog.consolidate_catalog()

catalog.consolidate_catalog(
    start=1704067200000000000,
    end=1704153600000000000,
    ensure_contiguous_files=True,
)

catalog.consolidate_data(
    NautilusDataType.QuoteTick,
    identifier="EUR/USD.SIM",
    start=1704067200000000000,
    end=1706745600000000000,
)
```

### Consolidate catalog by period

Split data files into fixed periods. Durations and time bounds use nanoseconds. Both methods accept
optional bounds. Supply an identifier to the data-type method for data partitioned by instrument.

The catalog-wide method processes quotes, trades, order book deltas, order book depths, bars, index
prices, mark prices, instrument closes, and registered custom types. It logs a warning and skips
other types. The data-type method rejects instrument and record selectors, which have no
period-typed rewrite; use `consolidate_data(...)` for those.

```python
DAY_NS = 86_400_000_000_000
HOUR_NS = 3_600_000_000_000

catalog.consolidate_catalog_by_period(period_nanos=DAY_NS)

catalog.consolidate_catalog_by_period(
    period_nanos=HOUR_NS,
    start=1704067200000000000,
    end=1704153600000000000,
)

catalog.consolidate_data_by_period(
    data_type=NautilusDataType.QuoteTick,
    identifier="EUR/USD.SIM",
    period_nanos=HOUR_NS,
)

catalog.consolidate_data_by_period(
    data_type=NautilusDataType.TradeTick,
    identifier="EUR/USD.SIM",
    period_nanos=HOUR_NS,
    start=1704067200000000000,
    end=1706745600000000000,
)
```

### Delete data range

Delete data within a time range, optionally limited to one data type and instrument. Omitting
`start` extends the range to the beginning; omitting `end` extends it to the end. For
`delete_data_range(...)`, omitting both bounds removes all matching data. Supply an instrument ID
for data partitioned by instrument.

`delete_data_range(...)` supports quotes, trades, bars, order book deltas, order book depth, and
registered custom types. Pass `NautilusDataType.OrderBookDepth` for order book depth and
`NautilusDataType.Custom("MarketTickPython")` for a custom data type.

`delete_catalog_range(...)` continues after unsupported directories, logs a warning, and leaves
their data unchanged. Use `delete_data_range(...)` when you need to confirm that the requested
type is supported. `NautilusDataType.Instrument` is rejected because instrument definitions do not
support ranged deletion.

```python
catalog.delete_catalog_range(
    start=1704067200000000000,
    end=1704153600000000000,
)

catalog.delete_catalog_range(end=1704067200000000000)

catalog.delete_data_range(
    data_type=NautilusDataType.QuoteTick,
    identifier="BTC/USD.BINANCE",
)

catalog.delete_data_range(
    data_type=NautilusDataType.TradeTick,
    identifier="EUR/USD.SIM",
    start=1704067200000000000,
    end=1706745600000000000,
)
```

:::danger[Permanent data removal]
Delete operations cannot be undone. The catalog splits partially overlapping files to preserve data
outside the range.
:::

## Feather streaming and conversion

The runtime can stage records in local Feather files and promote them into a catalog with
`StreamingConfig(writer_path=..., catalog=DataCatalogConfig(...))`. Staged records become available
to catalog queries after promotion succeeds. A staging flush and a catalog commit are separate steps.

Streaming defaults to promotion on close, no interval-based promotion, and retention of committed
Feather sources. A positive `promotion_interval_ms` uses live wall-clock scheduling or checks
against the supplied backtest clock during writes and flushes. See
[stream data into a Parquet catalog](../../how_to/stream_parquet_catalog.md) for defaults, configuration,
query visibility, and recovery.

`StreamingFeatherWriter` remains available for direct staging. Its completed sessions can be converted
manually with `ParquetDataCatalog.convert_stream_to_data()`.

Shown in full with attribution under the source's licence. Licence: LGPL-3.0

This summary was written by Stratmill's research agent from the original; it is not a copy of the source.