Limits of Reconstructing Missing Trade Ticks
Summary
The document describes a lower-resolution trade feed that preserves selected timestamps and prices while aggregating the volumes of trades between observations. Because the omitted trades are unknown, the question asks whether a single incomplete feed can support recovery of the missing events, perhaps through Bayesian modeling or image-processing ideas, and whether two independently incomplete feeds can be combined.
The responses emphasize that missing ticks cannot be recovered with certainty from one feed because the information has been lost. With multiple sources, one can use a primary feed and fill gaps from a secondary feed, while checking timestamps. Interpolation may be possible, but it introduces bias, and the appropriate method depends on the later use of the data. The document mentions linear interpolation as an example, not as a validated universal solution. Reconstructed data should therefore be treated as an estimate whose distortions may affect downstream analysis.
Key ideas
- Aggregated observations do not preserve enough information to recover omitted trades exactly.
- A second feed can be used to fill gaps, with timestamps checked for consistency.
- Interpolation can create estimates but introduces bias into the resulting data.
- The choice of reconstruction method depends on how the altered data will be used.
Tags
Full text
# Recover full tick data from missing tick data
# Recover full tick data from missing tick data
Due to some economics/regime problem, I can only have access to non full-tick data from an exchange.
To make the problem precise, a full tick data $X$ is a series of $(t_i,p_i,v_i)$ for $0 \leq i \leq N$ where $t_i$ is the timestamp, $p_i$ is the price, $v_i$ is the deal volume.
The data that I could only see is a lower resolution $\hat{X}$ of $X$, in the sense that, I can only observe the market in a sequence $j_1 < j_2 < \ldots < j_m$ and get the data like: (the sequence is not necessary deterministic or in fixed interval)
$(\hat{t_{j_k}},\hat{p_{j_k}},\hat{v_{j_k}})$ where $\hat{t_{j_k}} = t_{j_k}, \hat{p_{j_k}} = p_{j_k}$, but $$\hat{v_{j_k}} = \sum_{i=j_{k-1}+1}^{j_k} v_i$$
For instance, if the true data $X$ is:
$(0,100,1) \\ (1,102,2) \\ (2,101,1)$
I may only see the lower resolution one $\hat{X}$ as
$(0,100,1) \\ (2,101,3)$
or
$(1,102,3) \\ (2,101,1)$
The question is..
- Suppose I only have one source of $\hat{X}$, what is the best way to recover most missing tick? I know this may be a bad question, as information has already been lost. I think I need to add some model assumption for this problem from Bayesian point of view, any reference for this?
- Suppose I have two different source of $\hat{X}$, and because of random nature of the missing ticks, two source would be different. Any method to recover it?
P.S. I think I can think the tick data as a one-dimensional image, and lower resolution data is a pixelized version of real image data, and apply some image processing technique on it, any idea?
## Answer by chrisaycock (score 5, accepted)
https://quant.stackexchange.com/a/7316
If you're missing ticks, then no technique will get those ticks back.
If you have two sources, then designate one source as the primary feed and then fill-in gaps from the secondary feed. Of course, you'll have to mind the timestamps when determining whether the secondary feed can be used properly.
## Answer by Alexey Kalmykov (score 5)
https://quant.stackexchange.com/a/7320
Obviously merging two streams is harmless and it should be done. But it's hard to advise you regarding the "interpolation" methods you can use to generate the ticks without knowing why you need this. The reason is that any method will introduce a certain bias to the data. Therefore, it very much depends on what are you going to do with your altered data on the next step.
Some links regarding the interpolation methods that you can find useful:
- Take a look at the book An Introduction to High-Frequency Finance (preview in Google Books available)
- Olsen Data is using linear interpolationShown in full with attribution under the source's licence. Licence: CC BY-SA 4.0 (Stack Exchange)
This summary was written by Stratmill's research agent from the original; it is not a copy of the source.