Skip to content
All library documents

Building Comparable Futures Options Data for Backtests

Article Quant Q&A · Author: Drew

Summary

The document recommends storing options observations in a time-indexed, long-format table, with one row per contract at each snapshot. Useful fields include timestamp, underlying or forward price, strike, expiry, option type, bid and ask, implied volatility, and relevant Greeks or inputs. This raw record preserves the listed chain as observed and supports filtering by maturity, moneyness, date, and liquidity. Point-in-time underlying prices and rates help avoid look-ahead bias when calculating implied volatility or Greeks; stale quotes make bid/ask, volume, and open interest useful quality checks.

For comparisons across time, the response suggests building a derived surface on fixed tenor and relative-moneyness or delta coordinates. This avoids treating changing absolute strikes as equivalent contracts when the underlying moves or listings change. A labeled multidimensional array or a panel table can hold the resulting surface. The discussion is methodological and provides no empirical backtest results. Interpolation and filtering choices affect the surface, while synchronization, carry assumptions, and contract settlement conventions can affect moneyness and valuation.

Key ideas

  • Store each option contract snapshot as a separate row with timestamp and market inputs.
  • Keep point-in-time underlying prices and rates aligned with option quotes to reduce look-ahead bias.
  • Use relative moneyness or delta and fixed tenors to construct comparable surfaces over time.
  • Filter stale or illiquid quotes using market quality fields such as spreads, volume, and open interest.
  • Account for forward construction and settlement conventions when calculating moneyness and implied volatility.

Tags

Full text
# What is the most convenient data structure for backtesting a model of futures options prices?


# What is the most convenient data structure for backtesting a model of futures options prices?












I have an empirical model for the dynamics of futures prices in a particular market that I have implemented using a long series of the front five contracts. (I account for the roll in my model.) I have never tried to forecast prices of options on these futures, but I have some ideas and I would like to experiment. Is there an analogous data construct used by practitioners to backtest options pricing models? I am thinking of something like a series of a vector containing the prices of the ATM put and two strikes above and below or something. Also, I am thinking of a format that might be constructed conveniently by a canned Bloomberg API, or similar thing.

## Answer by James Cartwright (score 0)

https://quant.stackexchange.com/a/85483

For backtesting option pricing or trading models, the most convenient data structure is typically a time-indexed panel (or “long format”) table where each row corresponds to a specific contract at a specific timestamp.

At minimum, the structure should include: timestamp, underlying price, strike, maturity, option type (call/put), bid/ask or mid price, implied volatility (if available), and any relevant Greeks or model inputs. Using a multi-index (timestamp, contract identifier) in environments like pandas, or a relational table with composite keys in SQL, makes it easier to filter by maturity buckets, moneyness, or trading date.

A long format is generally preferable to a wide format because maturities and strikes vary over time, and wide matrices quickly become sparse and difficult to maintain.

For performance-sensitive backtests, storing the data in columnar formats (e.g., Parquet) and separating static contract metadata from time-series price data can also improve efficiency.

The key consideration is not the specific container, but ensuring that the structure supports consistent alignment between underlying data, option quotes, and model state across time.

## Answer by Ankur Parikh (score 0)

https://quant.stackexchange.com/a/85802

The trap in "ATM put and two strikes above/below" is that it's anchored to absolute strikes, and the strike grid drifts relative to the underlying as the future moves and as new strikes get listed. Your "ATM" contract changes identity over time, and the ±2 strikes aren't comparable from one day to the next. Just as you made your futures series comparable by rolling to a constant-maturity front-five construct, the options analog is to store the chain in relative coordinates and build a constant-maturity, constant-moneyness surface on top of it.

I'd keep two layers.

- Raw layer — point-in-time, long/tidy, one row per contract-snapshot. Store the chain as observed, without reshaping:

date | expiry | strike | cp | bid | ask | mid | iv | delta | volume | oi | underlying | rate

Long, not wide — a wide "strike ladder" breaks the moment the grid changes. Point-in-time: record the underlying/forward and rate as they were at that timestamp, and compute IV/greeks from those, never restated later — otherwise you bake look-ahead into the surface. Keep bid/ask, volume and OI so you can later filter to genuinely tradeable strikes; a large fraction of listed strikes are stale quotes. Columnar storage partitioned by date (parquet) is convenient and is exactly what a Bloomberg/vendor pull drops into. 2. Derived layer — the analysis-ready surface. From the raw table, interpolate each day onto a fixed grid of (tenor, moneyness-or-delta):

date × tenor × moneyness → iv (or price)

Tenors like 30/60/90 calendar days; moneyness as K/F or log-moneyness — or better, by delta (25Δ put / ATM / 25Δ call), which is how practitioners quote surfaces and which stays stable as spot moves. This is the direct analog of your roll-adjusted front-five futures: a constant-maturity, constant-moneyness vol-surface time series. Every cell is now comparable across time, and backtesting a pricing model becomes "compare my predicted surface cell to the realized one." In Python the natural container is an xarray Dataset with dims (date, tenor, moneyness) — a labeled cube that interpolates cleanly; a pandas MultiIndex panel also works. For a first experiment you don't need the full delta-interpolated surface — a small fixed moneyness grid (ATM and ±5%/±10% of the forward, across a couple of tenors) captures most of it and already sidesteps the fixed-strike drift. Just parameterize by moneyness, not absolute strike.

Options-specific pitfalls the futures case doesn't have:

The strike grid expands/contracts as spot moves — never assume a fixed ladder. Synchronise each option snapshot with the contemporaneous future price (same timestamp); a mismatch there silently corrupts every IV. Liquidity: filter by OI/volume and spread before trusting a "price." Carry/dividends and American-vs-European settlement feed the forward you compute moneyness and IV against. Bloomberg (or any vendor) will hand you the raw long table happily; the constant-maturity / constant-moneyness construction is the part you build — and it's the part that makes the backtest honest.

Shown in full with attribution under the source's licence. Licence: CC BY-SA 4.0 (Stack Exchange)

This summary was written by Stratmill's research agent from the original; it is not a copy of the source.