Skip to content
All library documents

Bulk Volume Classification Accuracy Across Bar Sizes

Article Quant Q&A · Author: babelproofreader

Summary

The document considers whether Bulk Volume Classification (BVC), which estimates the buy-trade share within a bar, should be applied to short or long bars. One response argues that larger bars provide more trades and therefore tend to yield more accurate estimates. It cites a comparison across time bars and volume bars in three futures markets: reported classification accuracy rises from roughly 70% for one-minute bars to roughly 90% for one-hour bars, with accuracy also increasing across the tested volume-bar sizes.

The practical choice of bar size still depends on the trading strategy’s time horizon and required signal frequency. Short bars may be useful when decisions must be made quickly, but the cited results suggest accepting lower classification accuracy. Another response recommends prototyping and choosing time frames according to the classifier’s objective, while noting the possibility of using features from multiple horizons. The document does not provide the study’s full methodology, so its figures should be treated as reported evidence rather than a universal accuracy guarantee.

Key ideas

  • BVC estimates the proportion of buy trades within an aggregated bar.
  • The cited comparison reports better classification accuracy for larger time bars and volume bars.
  • Bar size should reflect the strategy’s decision horizon as well as the desired classification accuracy.
  • Shorter bars may support faster signals while producing noisier trade-flow estimates.
  • The reported results cover three futures markets and do not establish performance for all instruments or datasets.

Tags

Full text
# Bulk Volume Classification Algorithm


# Bulk Volume Classification Algorithm












I'm thinking of implementing the Bulk Volume Classification algorithm on my data of hourly OHLC bars and associated volume, but my sense of it is that hourly granularity is insufficient and that really this bulk volume classification should be applied to bars of a lower time frame, e.g. 1 or 5 minute bars. Is my intuition correct?

## Answer by MikeRand (score 4)

https://quant.stackexchange.com/a/44249

I'd argue your intuition is backwards.

Think about what Bulk Volume Classification is doing: it's estimating a proportion of Buy trades for a given bar. All else being equal, for bigger bars (i.e. larger sample sizes), an estimate of the proportion will be more accurate.

In the follow-up paper where they assess the accuracy of Bulk Volume Classification vs. aggregate tick rules here, the authors assess it across time bars (1, 2, 3, 5, 10, 20, 30, 40, 50 and 60 minutes) and volume bars (1000-25000 contracts). Indeed, their trade classification accuracy for 1-minute bars is ~70% but ~90% for 1-hour bars (83% for 1000 contract bars and 94% for 25000 contract bars). This trend is seen across all 3 futures tested.

From a pure trade classification perspective, BVC gets better (not worse) the larger each bar is. The decision to use lower time frame bars should be made for other reasons (frequency of trading strategy, etc.) and take into account that the BVC will get worse.

You probably want to consider whether to use BVC at all. I will note that in his recent book, one of the authors makes little-to-no use of BVC that I could determine, but uses tick rules quite frequently when discussing microstructural features.

## Answer by Emma Marcier (score 4)

https://quant.stackexchange.com/a/44267

It highly depends on the goal of your classifier, or your Bulk Volume Classification Algorithm in this case.

MikeRand might be right, if you would be in any search of one hour or few hours candles information from time-frame viewpoint.

But, if your time-frame might be lower (e.g., `1 min`, `5 min`), which may appear so, then your intuition is right, yet you will face accuracy issues, as MikeRand very well points out. However, if, at this time, accuracy may not be an issue for you, and you are interested in prototyping a BVCA, that is fine then.

Trading algorithms are very pertinent to speed processing systems. If I were you, I would first prototype one BVCA (e.g., any time-frame), then redesign and re-architect for expanding into a `asyn` system of algorithms in tangent based on only BVCA algorithm or other practical algorithms such that I could implement it to lower time-frames (`1 second`, `5 seconds`, ... , `1 min`, `5 mins`, ...) as well as higher time-frames (e.g., `1 hour`, `3 hours`, ..., `1 day`). Because in algorithmic trading systems, you can extract various features from each time-frame, which are important and connected to information from other time-frames, and that is what practical traders do by sitting and watching the real-time charts.

I would then optimize for precision based on processed data from each time-frame, according to my objective function, however that would be. Because, computational costs may not be a factor or issue in algorithmic trading systems, not to mention you may be able to rent super-computing CPU services, once you are done with prototyping and scale-up, in production level, which will help you to perform your entire computations in `0.5<t<5 seconds`.

Great question!

Image Courtesy: Illinois State Government

Comparing Trade Flow Classification Algorithms in the Electronic Era: The Good, the Bad, and the Uninformative

Parallel Algorithms

Shown in full with attribution under the source's licence. Licence: CC BY-SA 4.0 (Stack Exchange)

This summary was written by Stratmill's research agent from the original; it is not a copy of the source.