Methods for Analyzing Irregularly Spaced Tick Data
Summary
The discussion outlines ways to study high-frequency quotes and trades when observations arrive at uneven intervals. It recommends first defining the target, such as volatility or correlation, and accounting for how measurement intervals affect the statistics. It points to realized volatility research and methods for unevenly spaced series, while noting that market microstructure noise, including bid-ask bounce, can make estimates from every tick less reliable than estimates from aggregated intervals.
The answers also distinguish clock time from trade time. Trade-time sampling can help compare activity across assets with different trading frequencies, but clock-time patterns such as opening and closing effects may still matter. For order-book data, the analysis may need to include depth and changes across price levels, not just top-of-book prices. The suggestions are introductory rather than a single prescribed method: the right approach depends on the question and whether the data contains top-of-book or full-book information.
Key ideas
- Define the statistic or price behavior of interest before choosing a sampling method.
- Irregular observation times require care because variation depends on the elapsed interval.
- Microstructure noise can make tick-by-tick volatility estimates less useful than estimates from aggregated data.
- Trade-time sampling helps account for different activity rates, while clock-time effects may remain relevant.
- Full order-book data can support analysis across price levels as well as through time.
Tags
Full text
# Analyzing tick data # Analyzing tick data What are some of the commonly used techniques to analyze tick data? I am looking at tick data to see how the quotes/ mid-price evolves due to certain events in the market. Since tick data is asynchronous one can't really apply traditional time series models to explain these price movements. Some people have proposed that I create price bars based on either clock-time or trade-time but I think that tends to miss out on information happening in between the bars. Any suggestions on how I can approach this ? ## Answer by Shane (score 18, accepted) https://quant.stackexchange.com/a/4330 Your question is very vague (e.g. what are you trying to measure, and what "tick data" do you have), but I'll give you some pointers: - In general, when people consider how prices evolve, they will tend to think about things like volatility and correlation dynamics. So I would start by defining exactly what you want to measure. The irregularity of time series data is not a problem in itself, except in so far as you are making assumptions in your calculations about things like dispersion in time. The amount of variation over 1 millisecond will generally be different than over 1 second (and will also vary by asset), so you need to arrange your statistics to account for this. 1.1. There is a vast literature on measuring volatility using high-frequency tick data. Search for papers on realized variance, volatility, and correlation from people like Neil Shepard (see his institute) or Tim Bollerslev. One feature of this literature is that it is actually optimal to not use tick-by-tick data because of what is known as microstructure noise (e.g. bid-ask bounce), and you're generally better making estimates off something like 5-minute data. 1.2 There is also a literature on dealing with unevenly spaced data (see, for instance, papers by Muller and Zumbach). A recent paper on the subject is "Algorithms for Unevenly-Spaced Time Series: Moving Averages and Other Rolling Operators". There is a nice section in Eric Zivot's book on time series analysis that covers this (look for irregularly spaced high frequency data or inhomogeneous operators). - Looking at statistics in clock time or trade time is an important distinction. For instance, the number of quotes or trades can vary dramatically across assets, with illiquid assets only trading a few times a day vs. liquid assets which trade many times each second. Using trade time to measure things like volatility can partly address this problem (as well as things like the significance of your estimate), although you will need to consider whether there are other clock time effects (such as open or close time seasonalities) even when you work in trade time. - For tick data, are you working with level 1 (top of the book quotes and trades) or level 2 (full order book) data? If it's level 2, then you may not only want to consider changes through time, but also across the book. ## Answer by Konsta (score 3) https://quant.stackexchange.com/a/4298 In order to use methods for equidistant time series: - simply disregard timestamps - separate trade and clock time (like 1: clock time increments as time series) - create [sparse] equidistant time series with tiny time increment ([implicitly] repeating prices when necessary) - aggregate equidistant bars Although some above are blatant, they would get you going. Besides that, I have had Engle, Russell, 2004, "Analysis of High Frequency Financial Data" waiting for me to read it for some time now. An Introduction to High-frequency Finance might be relevant, too. ## Answer by shoonya (score 1) https://quant.stackexchange.com/a/4261 In case of Tick Data, you can use the RTAQ package in R. The standard techniques for analyzing tick data can be seen in Haustch or Frederi G. Viens ## Answer by jeff m (score 0) https://quant.stackexchange.com/a/4290 I suggest checking out some of the research from Nanex. You should be able to pick up some methods just by going through some of their event analysis.
Shown in full with attribution under the source's licence. Licence: CC BY-SA 4.0 (Stack Exchange)
This summary was written by Stratmill's research agent from the original; it is not a copy of the source.