Why Millisecond Trade Data Can Show Tied Timestamps
Summary
The document discusses repeated trade records that share a displayed timestamp and price but report different quantities. Its main explanation is that the source may round or aggregate event times to millisecond precision, making distinct executions appear simultaneous. Exchange feeds can retain finer timestamp resolution, so the apparent tie may reflect the historical vendor’s data format rather than the actual event sequence. A separate response suggests trades could also differ by venue or trade code.
The discussion distinguishes individual trade messages from broader market-data feeds, which can include order-book changes and administrative updates as well as executions. It argues that vendor datasets may be aggregated or transformed, and describes direct exchange feeds as a route to more detailed messages, subject to licensing and technical requirements. These are general cautions, not a diagnosis of the specific dataset: its provider and specifications are unknown. Researchers should check timestamp precision, aggregation rules, venue coverage, and trade-condition fields before interpreting same-time observations or drawing conclusions about execution order.
Key ideas
- Trades with matching millisecond timestamps may have occurred at distinct times that the data source rounded to the same value.
- Timestamp precision and aggregation practices vary across market-data sources.
- Trade messages can include execution details, while feeds also carry order-book and market-status updates.
- Venue or trade-condition differences may also explain records that appear identical at first glance.
- Researchers should consult the source specifications before inferring execution sequence from timestamps.
Tags
Full text
# Interpreting tick by tick stock data # Interpreting tick by tick stock data I downloaded data regarding the YUM stock traded on the New York Stock Exchange in the year 2014. Basically in this dataset, we have 4 columns: the first column is the day the second column is the time in milliseconds expressed in epoch time, the third column is the transaction price and the fourth column is the volume. At a particular moment, there are instances where transactions, reported in milliseconds, occur at the same price and instant but involve different volumes for example ``` 28,1393603190476,56940,450 28,1393603190476,56940,188 ``` These 2 rows indicate that on the 28th at the instant 1393603190476(milliseconds elapsed from 1 January 1970) there are 2 transactions at the same price with different volume Why are the transactions reported separately if they occur at the same instant? ## Answer by danospanos (score 1, accepted) https://quant.stackexchange.com/a/78602 > Why are the transactions reported separately if they occur at the same instant? In the absence of any other information about your data source, it is most likely that the trades do not occur at the same instant. That is, your data source has timestamped the trades to the nearest milisecond. But trades on NYSE Arca are not executed in milisecond time buckets. There are trades within those miliseconds, and that is why it appears to you that they were executed at the same time. You can check the real-time datafeed specification, to see that the trade message timestamps go down to nanoseconds. Regarding the importance of time resolution, I recall a paper that compares the differences between measurements derived from second vs. milisecond time-stamped TAQ data, definitely a good read: Holden, Craig W., and Stacey Jacobsen. "Liquidity measurement problems in fast, competitive markets: Expensive and cheap solutions." The Journal of Finance 69, no. 4 (2014): 1747-1785. ## Answer by Jim Broiles (score 1) https://quant.stackexchange.com/a/79306 Free data equals bad data. Cheap data equals bad data. You are not dealing with real data here. This data has been aggregated to the 1ms timeframe at best. I have seen other such data that is aggregated to 4ms timeframe. The providers call this real-time data but be careful. In data provider parlance, real-time can be anything that is provided less than 10 minutes from the time the event occurred. Real trade message data from the exchange is timestamped with 1ns resolution. Trade messages typically contain the information for a single transaction. This information includes Exchange Time (when the transaction occurred), Sending Time (when the transaction message was sent to the wire), Price, Quantity Filled, Order Count (Aggressing order filled by N sitting orders), Aggressor Side (Sell the Bid, Buy the Ask). The highest volume of messages in a session comes from order book updates. These consist of new limit orders, sitting order modification and/or cancellation. The volume of these messages can exceed trade messages by 5 to 1. There are also various administrative messages sent to indicate exchange status, market status, instrument status, statistics, etc. Historical data from the exchange is stored as a PCAP (packet capture) file which contains the binary encoded messages for the entire session. The exchanges publish very detailed specifications on the message protocols (usually based on SBE or FIX). Do a bit of google search to find the specifications for the exchange's message protocol to find out exactly what the messages contain and timing of the messages. Search "NYSE FIX Protocol" or "NASDAQ Market Data Message Spec" and go from there. The only place to get this data is from the exchange. Of course, there are other providers of market data, but they offer aggregated data and are often restricted by the terms of their license agreement (in exchange for lower cost) with the exchange from providing the raw data. If you want the real data, the only way is to be directly licensed by the exchange to have access to this data. If you can get the license, which retail traders and researchers usually cannot, you have been granted permission to connect to the market data feed and receive the serialized, encoded messages. Great! Now all you have to do is develop an application that can connect to the exchange session (must pass a rigorous certification process), receive the message packets (over UDP typically), decode them and present them for analysis by a trading strategy. If you want to display the data on a chart on a screen you have to pay more for the licensing fee. There are vendors who offer the "FIX Decoders", which is a great way to accelerate the development of your system to read market data. Of course, you'll want to place orders as well. The order feed is an entirely separate feed and protocol, but the process of getting licensed, certified and able to place orders is similar except for the fact that now other people become concerned with your risk profile and such. But I digress. You also need to be collocated as well. The connection to the exchange servers does not work for servers that are not collocated. In summary, any data you get that is not from the exchange has been aggregated or manipulated in some way by the data provider to fit within their business model, storage capacity, infrastructure, etc. It is not real data. I'm afraid far too many people simply are not aware of this. ## Answer by cdatwork (score 0) https://quant.stackexchange.com/a/78878 Maybe the trades occurred on separate markets/exhanges (e.g., NYSE vs. CBOE) or maybe different trading codes (e.g., "bunched sold trade" vs. "intermarket sweep").
Shown in full with attribution under the source's licence. Licence: CC BY-SA 4.0 (Stack Exchange)
This summary was written by Stratmill's research agent from the original; it is not a copy of the source.