Approximating Dollar Bars from Minute OHLCV Data
Summary
The document considers how to construct dollar bars when only minute OHLCV data is available. One proposal estimates each minute’s traded value by multiplying volume by an average price, then combines minutes until their estimated dollar value reaches a chosen threshold. This creates bars that may overshoot or undershoot the target because the minute data does not reveal the order or value of individual trades.
A response suggests estimating the minute’s representative price from its high and low, since averaging all four OHLC values imposes extra assumptions about price movement within the minute. It also notes that a package provides standard and information driven bar methods that can operate on minute timestamps, and proposes empirical investigation of whether dollar bars improve return distribution properties. A second answer sketches a more detailed approximation by assigning an assumed sequence and timing to OHLC prices and splitting volume among them. These are approximations: minute data cannot recover true trade level dollar bars, and the claimed statistical benefits require testing on the data and use case at hand.
Key ideas
- Dollar bars group trading activity by a target dollar value rather than by elapsed time.
- Minute OHLCV data can approximate traded value using a representative price multiplied by volume.
- Using the high and low midpoint avoids some assumptions required by averaging all four OHLC prices.
- Minute aggregation can overshoot or undershoot the intended dollar threshold.
- Any expected improvements in stationarity or return normality should be evaluated empirically.
Tags
Full text
# How can I approximate Dollar Bars from Minute Data instead of Tick Data? # How can I approximate Dollar Bars from Minute Data instead of Tick Data? Having been influenced by de Prado's Advances in Machine learning book, I've set out to build the dollar bars (in which each bar represents a set dollar amount of transactions in the security) that he endorses as a superior data structure to conventional time-based bars, mostly for its more stationary, iid, and statistically useful properties. Unfortunately, I just don't have the tick data necessary to really put the idea to use. I do, however, have an abundance of 1-minute data, which has me wondering the most faithful method I might use to approximate true dollar bars. My plan is to: - take the average of the OHLC of each minute bar, - multiply that by the volume of that bar, - assign that dollar value to the bar, - and then begin aggregating the bars to the desired dollar amount from the start of the original time series to its end. I realize, though, that this might introduce slightly over/undershooting the target dollar amount for each bar, depending on that target dollar amount per bar. Is such an approach problematic or otherwise unworthy, given de Prado's intentions for the dollar bar? Is there a better way to go about it? ## Answer by Jacques Joubert (score 1, accepted) https://quant.stackexchange.com/a/50946 The following python package, mlfinlab, provides an implementation for both standard and information-driven bars. The good news is that you won't have to implement the techniques from scratch and they will also work on minute time stamps. Regarding how to approximate the VWAP of a minute bar: - Perhaps it's better to take the average (midpoint) of only the low and high. If you take the average of OHLC then you add additional assumptions about price evolution. Applying dollar bars to minute data may make your data less heteroscedastic and you would probably see a return to normality in the returns. An empirical study would prove useful. ## Answer by 13ue (score 0) https://quant.stackexchange.com/a/78343 I am no pro on the topic but i developed the following algorithm to solve exactly your problem. It might be a bit naive, but i'd still love to learn the ins and outs. I am basically making an educated guess of the position in time of the open, high, low and close value of the candle: If `open < close` the order is `[open, low, high, close]` else `[open, high, low, close]`. I then estimate the timestamp by adding a quarter timeframe meaning 15 seconds for 1min bars to each "tick" after the open. To calculate dollar bars i split the candles volume by 4. Do you think this makes sense? For implementing: `pd.melt` does the job :)
Shown in full with attribution under the source's licence. Licence: CC BY-SA 4.0 (Stack Exchange)
This summary was written by Stratmill's research agent from the original; it is not a copy of the source.