Skip to content
All library documents

Choosing Intraday Sampling Frequency for Market Microstructure Research

Article Quant Q&A · Author: Ana

Summary

The document offers a framework for choosing the sampling interval for intraday market research instead of assuming that five-minute bars are always appropriate. The researcher should match the data frequency to the trading horizon they can act on and to the market effect or question being studied. The frequency itself may also be a research choice when the goal is to compare how findings change with aggregation.

Examples cited span hourly aggregates, half-hour intervals, five-minute intervals, and tick-by-tick observations, including full order-book data. This range shows that no single interval is standard across all research questions. Finer observations can require substantially more data and computing resources, so practical constraints matter alongside analytical goals. The guidance is qualitative: it does not prescribe a universally optimal interval or provide a controlled comparison of results at different frequencies.

Key ideas

  • Choose an interval that fits both the question being studied and the frequency at which a practitioner could trade.
  • The sampling frequency can itself be the subject of analysis when researchers compare aggregation choices.
  • Published examples use intervals ranging from hourly data to tick-by-tick observations.
  • Fine-grained data can sharply increase storage and computing demands.
  • The document gives decision criteria rather than a universal best frequency.

Tags

Full text
# Intraday data frequency


# Intraday data frequency












How should I determine what frequency should I use for doing microstructure research using intraday data? For some reason, there seems to be general consensus of using 5 minute interval, but is there any advantage to this over higher frequencies if e.g. 1 minute data is available?

## Answer by NegativeJo (score 4)

https://quant.stackexchange.com/a/24588

I think the answer is driven by asking yourself a few questions :

- If you are a practitioner at what frequency are you able to trade and want to trade ? (you are limited by this so no need to go to higher frequencies than that in any case)

- What effect(s) do you want to study ?

- What is a common frequency used by practitioners or academics for the question you have. Maybe which frequency is best becomes the question you want to study. You can form an opinion/intuition ex-ante as well.

- Potentially computing time and resources (tick by tick can grow to a lot of data very quickly)

I have seen people use from a 1 hour aggregation to the tick by tick full book level.

A few examples of academic paper with different aggregations : https://www.scheller.gatech.edu/directory/faculty/lee_s/pubs/Jump2-1-12.pdf (analyst estimates impact on price, done at 30 min intervals)

https://www.researchgate.net/profile/Paresh_Narayan/publication/281200352_Intraday_volatility_interaction_between_the_crude_oil_and_equity_markets/links/55dabe1f08aed6a199aaf80c.pdf (oil price impact done at 5 min intervals)

https://www.researchgate.net/profile/Shaojun_Zhang3/publication/228302398_An_Improved_Estimation_Method_and_Empirical_Properties_of_the_Probability_of_Informed_Trading/links/00b495368bafd7961a000000.pdf (probability of informed trading on a tick by tick level)

If you specify which type of questions you want to answer, others might have more specific suggestions on which interval to use.

## Answer by Alifeleti Etuate (score -1)

https://quant.stackexchange.com/a/47015

I think it all about your choice of data you want to research on, this is because you can always switch to finer intervals if 5 min is not working for you. So high frequency will range between tick, 1 sec, 1 min and 5 min bars. As a researcher you have the autonomy to choose the data set you need given the research objectives that you have.

Shown in full with attribution under the source's licence. Licence: CC BY-SA 4.0 (Stack Exchange)

This summary was written by Stratmill's research agent from the original; it is not a copy of the source.