Skip to content
All library documents

Preparing Irregular Tick Data for Time-Based Analysis

Article Quant Q&A · Author: user1025852

Summary

The document discusses how to prepare unevenly spaced quote data for technical analysis and other calculations that assume regular time intervals. Several quotes may arrive within the same second, while other seconds have no observations. One suggested treatment is to retain the last quote at each unique timestamp, removing duplicate times. If a complete fixed-frequency series is required, the prior quote can be carried forward across missing intervals.

The responses caution that indexing ticks as if each were an equal time step distorts the market’s pace. Busy periods become stretched across many indexed observations, while quiet periods are compressed, making indicator signals difficult to compare and potentially impossible to execute at the implied speed. The appropriate method depends on the task: time-based analysis may use aggregated fixed-interval OHLC bars, while order book research may require reconstructing the book and sampling snapshots. These are practical suggestions rather than a universal recipe; bar frequency and sampling choices should fit the intended analysis.

Key ideas

  • Duplicate quotes at a timestamp can be reduced to the last quote when that suits the analysis.
  • A prior quote can fill missing intervals when a regularly spaced series is required.
  • Treating each tick as an equal time step hides the difference between busy and quiet periods.
  • Time-window indicators are more interpretable when observations preserve their actual time scale.
  • OHLC aggregation and order book snapshots serve different analytical needs.

Tags

Full text
# Analyze raw tick data


# Analyze raw tick data












I'd like to work with raw tick data and naturally this data is unevenly spaced (for example, a couple of quotes are at the same second etc.)

For example

```
10:12:35 - 14.44
10:12:35 - 14.45
10:12:35 - 14.47
10:12:36 - 14.46
10:12:36 - 14.49
10:12:37 - 14.50
```

My question is regarding how to set this data to "fixed" intervals for valid math calculation. Though I read here couple of suggestions on what should be done, I'm not sure its clear to me:

- Do I have to manipulate the raw data to have "clean" timestamps as the x-axis to work with common technical indicators?

- Or can I refer to the index as the x-axis (and just ignore the timestamp)? Can I look at it as a "stream" of ticks to anaylze?

## Answer by chrisaycock (score 3)

https://quant.stackexchange.com/a/7906

If you just want to run some simplistic technical analysis on quotes, then select the last quote for each unique timestamp. That will ensure that you don't have duplicate timestamps. If you must have it evenly spaced (i.e. no gaps from one second to another), then you can reuse the previous quote to fill-in the missing value.

## Answer by hroptatyr (score 2)

https://quant.stackexchange.com/a/7915

To help you understand why you need to follow recipes (like chrisaycock's) just have a look at your tick data. You will find ticks clustered at some points in time while they seem scarce at others.

If you proceed with your recipe 2, you will lose those clusters of activity and stretch them out. In periods of low activity you will condense the market.

Most indicators you mentioned will expect to work over a window of time, simply because the results are so much more meaningful. In theory, you could apply the same algorithms to stretched or condensed indexed ticks, but the results will also be valid for those periods.

That means for a cluster of activity that makes the price jump from `X` to `X+10` within 1000 ticks but only one second of time your indicator might tell you to sell on the 400th tick and buy on the 800th. But to implement this you would have to execute a round-trip within less than a second.

Now also this indicator's results are less comparable. Because in a period of low market activity the indicator might give you the same result (sell on 400, buy on 800), but now it's stretched over, say, hours, and it's much more likely that you can realise this trade.

## Answer by ast4 (score 0)

https://quant.stackexchange.com/a/7925

It really depends on what you're trying to do, a solution as simple as agging the data to x-sec OHLC bars may suffice and from the sounds of things that's what you need. Now if you need to work with order book dynamics then tick data is fairly crucial, then what I'd do for analysis is reconstruct the orderbook from the ticks then just take snapshots of the book (again depends on your requirements).

Shown in full with attribution under the source's licence. Licence: CC BY-SA 4.0 (Stack Exchange)

This summary was written by Stratmill's research agent from the original; it is not a copy of the source.