Skip to content
All library documents

Why Tick Imbalance Bar Thresholds Depend on Initialization and Updates

Article Quant Q&A · Author: the_dude

Summary

The document examines a practical difficulty in implementing tick imbalance bars from a financial machine-learning text. The proposed threshold depends on the expected number of ticks per bar and the expected imbalance in tick signs. The questioner reports that an implementation using exponentially weighted averages drives the expected imbalance toward zero, causing the threshold to shrink, and asks whether this behavior is expected.

The response highlights sensitivity to initialization: the expected ticks per bar, the unconditional expected tick sign, and the smoothing rates for both estimates all affect the bars produced. It cautions that following the suggested updating procedure closely may leave the amount of information per bar and its evolution largely outside the implementer’s control. Other implementations described in the response cap the expected ticks per bar or use a fixed imbalance threshold. The material offers an implementation caveat rather than a definitive formula or empirical comparison; parameter choices and any modifications should be assessed in the intended data setting.

Key ideas

  • Tick imbalance bars use estimates of expected bar length and expected tick-sign imbalance to set a threshold.
  • Initialization choices and exponential smoothing rates can materially affect the resulting bars.
  • The expected amount of information per bar may be difficult to control under the suggested updating procedure.
  • Alternative implementations can cap expected bar length or choose a fixed imbalance threshold.
  • The discussion does not establish which parameterization or modification performs best empirically.

Tags

Full text
# Information driven bars - exploding threshold level


# Information driven bars - exploding threshold level












It's tough trying to figure this out myself and I after a few hours I thought I'd ask for help:

I'm trying to do tick imbalance bars from 'Advances in Financial Machine Learning', using equations:

- $b_t = \begin{cases} b_{t-1}, & \text{if } \Delta p_t = 0 \\ \\ \dfrac{|\Delta p_t |}{\Delta p_t }, & \text{if } \Delta p_t \neq 0 \end{cases}$

- $\theta_t = \sum_{t=1}^{T} b_t$

- $T^{*} = \arg \displaystyle\min_{T} \{ |\theta_t| \geq E_0[T] \vert 2P[b_t = 1] - 1 \vert \} $

and either get $E_0[T]$ converging to $1$ or the other term for calculating the threshold converging to $0$ or exploding.

Probably calculating them wrong. I've been staring at these pages for a few hours now and can't figure it out, this is what I'm doing:

For $E_0[T]$ each time the threshold is reached, I calculate EWMA of all the $T^{*}$ (first one being arbitrarily set) and use that for the 'next' threshold.

For the other component, each time the threshold is reached, I calculate the absolute value of the EWMA of all the $b_t$ up till now and use that for the 'next' threshold.

For EWMA, I use $a\cdot x(t)+(1-a)\cdot x(t-1)$.

Basically, mostly what happens is that the second term converges towards $0$ as time goes by, forcing the whole threshold value down.

I honestly think, that if we're trying to get the same amount of information in each 'bucket', maybe the threshold value should be stable over time, not dependent of observations? It's inevitable that unless there's a strong trend, the up/down probability is going to converge to zero. But the book can't be wrong, so it's got to work somehow.

Can anyone help?

## Answer by Charlie (score 2)

https://quant.stackexchange.com/a/53205

I explored this topic some time ago and I wrote an informal review on Medium. You can find my article here.

In a few points:

- Tick Imbalance bars are just the simplest application of information driven bars and we should build on that (rather than take its details too seriously) to uncover interesting insights. I believe Marcos Lopez De Prado has just showed us the way.

- The proposed bar generating mechanism is heavily affected by how you initialize its parameters. In fact, to produce imbalance bars you must initialize the expected number of ticks per bar, the unconditional expectation of the tick sign and the alphas that define the two exponential averages used to update our expectations.

- If you closely follow the suggested implementation, the amount of information we enclose in each bar and its dynamics are almost out of our control. And this is bad.

- Other researchers (see this post) modified the suggested implementation either putting a cap to the expected number of ticks per bar or directly choosing a fixed imbalance threshold.

Hope this helps.

Shown in full with attribution under the source's licence. Licence: CC BY-SA 4.0 (Stack Exchange)

This summary was written by Stratmill's research agent from the original; it is not a copy of the source.