How Tick, Volume, and Dollar Bars Differ from Resampling Methods
Summary
The document distinguishes financial bars from statistical resampling methods. Time bars divide observations by clock intervals, while tick, volume, and dollar bars aggregate irregular transaction data until a chosen trade count, traded volume, or traded value is reached. In this framing, bars convert aperiodic tick data into a periodic series that is easier to use with some analyses; they are an aggregation scheme, rather than random selection from the data.
Cross-validation partitions data to assess a predictive model’s sensitivity to overfitting, while bootstrap methods draw observations with replacement to assess a statistic. The response says it sees no special obstacle to using different bar definitions as inputs to those procedures. That is a brief conceptual answer, not a detailed validation protocol. It does not discuss time dependence, leakage, or how to preserve the chronological structure of market observations when forming folds or bootstrap samples, so those practical questions remain outside the exchange’s treatment.
Key ideas
- Tick, volume, and dollar bars aggregate irregular trades using different activity thresholds.
- Bars regularize tick data for analyses that work more easily with periodic observations.
- Cross-validation partitions data to assess predictive model sensitivity to overfitting.
- Bootstrap methods use sampling with replacement to assess a statistic.
- The response does not provide detailed guidance on preserving time structure during validation.
Tags
Full text
# Sampling and cross-validating with tick, volume and dollar bars # Sampling and cross-validating with tick, volume and dollar bars Financial data is usually structured with time bars. Other sampling techniques include: - tick bars - volume bars - dollar bars. These are so-called sampling techniques to better identify signals and trends in financial data. But if they are sampling techniques, then how do they compare to sampling techniques used to test the robustness of a model, such as cross-validation and bootstrapping/bagging (bootstrap aggregation)? Or are they incomparable and completely different animals? I'm thinking that bars are simply aggregation techniques, and are not higher-level sampling methods like CV and BS. Am I right to differentiate between sampling and aggregation and that they are different things? With this categorization aside, are there any issues one must consider when cross-validating or bagging with tick, volume, or dollar bars vs. time bars? ## Answer by chrisaycock (score 1, accepted) https://quant.stackexchange.com/a/54652 Your thinking is correct: bars are simply for aggregation. Within the realm of aperiodic data (tick data), bars make the data periodic, which in turn makes certain types of analysis easier to perform. I wouldn't call bars a "sampling" technique since that implies random selection. Cross-validation checks a predictive model's sensitivity to overfitting by using random subsets or partitions of the data. Bootstrapping assesses the accuracy of a statistic by using random drawings (with replacement) of the data. I can't imagine there would be any gotchas for using different definitions of bars as inputs for cross-validation or bootstrap.
Shown in full with attribution under the source's licence. Licence: CC BY-SA 4.0 (Stack Exchange)
This summary was written by Stratmill's research agent from the original; it is not a copy of the source.