Understanding Trade Ticks, Quote Data, and Aggregated Stock Bars
Summary
The document compares historical AAPL data from trade records, quote snapshots, and minute or multi-minute price bars. Its central question is why a displayed bid and ask can appear inconsistent with the prices in nearby trades or aggregated bars. The included replies explain that trade data may represent activity from only one venue or a particular feed, while widely distributed stock data may combine venues into a national best bid and offer. They also clarify that bar data summarizes observations over an interval rather than estimating each individual transaction.
The discussion emphasizes checking how each vendor defines and aggregates its feed before treating records as complete market history. It cautions that a trade dump may not include every market transaction, and that a venue-specific quote is not directly comparable to a consolidated quote. One reply characterizes the quote data as insight into market participants, but the more concrete explanation identifies it as best bid and offer data; the source-specific details still require verification. The exchange illustrates why feed coverage and aggregation matter in market data analysis.
Key ideas
- Trade records may cover only a particular venue or feed rather than every market transaction.
- Quote snapshots and trade prints represent different kinds of market observations.
- Price bars aggregate activity over intervals instead of estimating individual trades.
- Consolidated national quotes can differ from quotes reported by a single venue.
- Researchers should verify each data provider’s coverage and field definitions before analysis.
Tags
Full text
# Bid/Ask vs Low/High # Bid/Ask vs Low/High I am trying to gather historical data for experimental reasons (intellectual curiosity) and am having trouble understanding how that data is calculated. First some data gathering on AAPL from Feb. 10th, 2015 at opening. dataA = http://hopey.netfonds.no/tradedump.php?date=20150210&paper=AAPL.O&csv_format=txt dataB = http://hopey.netfonds.no/posdump.php?date=20150210&paper=AAPL.O&csv_format=txt dataC = http://www.google.com/finance/getprices?i=60&p=4d&f=d,o,l,h,c,v&df=cpct&q=AAPL dataD = http://chartapi.finance.yahoo.com/instrument/1.0/AAPL/chartdata;type=quote;range=4d/csv DataA seems to provide every transaction that took place during the prescribed day; Is that correct or am I reading the data wrong? If I take the first line of dataC (close,high,low,open,volume)=(120.3,120.31,120.16,120.17,646886), then it corresponds to the first few introductory transactions in dataA. Likewise, dataD also corresponds to the transactions of dataA, but over several minutes. In other words, dataC and dataD seem like estimations (using close,high,low,open,volume) of dataA. Is this correct? If this is true, then dataA is "raw data" and awesome for analytical reasons. However, I am confused by dataB. I suppose dataB is the bid/ask spread, but if I go to the following line: 20150210T150001 120.54 300 300 120.55 4600 4600 then the bid/ask seems to be 120.54/120.55 which seems entirely inaccurate compared to dataA (the raw data of actual transactions)? Even google indicates that the (c,h,l,o,v) is (120.39,120.58,120.25,120.3,576584) during the first minute of opening, which doesn't seem close to the 120.54/120.55 spread? What am I misunderstanding/misreading? ## Answer by Joshua Ulrich (score 1) https://quant.stackexchange.com/a/16592 Data set A does look like transactions, but I would hesitate to say that it is every transaction. You would need to investigate the data source and how transactions are defined. Data set B looks like BBO (best bid and offer). Data sets C and D are not estimations; they're aggregations to a higher periodicity. You need to investigate the data sources for data sets A and B. The US stock market is a distributed system. There are many trading venues. A and B could be from a specific venue, or a specific aggregation of venues, while the data on Google and Yahoo is likely from the NBBO (national BBO). In short, stock market data is complex. ## Answer by Rime (score 1) https://quant.stackexchange.com/a/16599 The data from hopey.netfonds is only data from the exchange. In this case all transactions you see there are NASDAQ quotations , hence the "O" after AAPL. It fails to provide transactions from other venues such as BATS etc. which is what free data providers usually use as Google finance ## Answer by Eric (score 0) https://quant.stackexchange.com/a/25501 I know this is an old post, but I just came across it while researching for the same data. As such, I though I might provide another explanation to the above question. The netfonds website provides both a 'tradedump' and a 'posdump'. The trade dump is simply tick level data that shows the movement of the asset from trade to trade. The posdump, 'dataB' in your case, provides insights into the market participants. This is important because you can identify irregularities in bid and ask prices and therefore capitalize on the difference in supply and demand. For further reading, investopedia provides and excellent explanation. ``` http://www.investopedia.com/terms/o/order-book.asp ```
Shown in full with attribution under the source's licence. Licence: CC BY-SA 4.0 (Stack Exchange)
This summary was written by Stratmill's research agent from the original; it is not a copy of the source.