Skip to content
All library documents

Kalman Filtering OHLC Data by Separating Midpoint and Range

Article Quant Q&A · Author: ABK

Summary

The document discusses a data-quality issue that can arise when a Kalman filter is applied independently to the high and low prices in minute-bar OHLC data. Because the filtered estimates are produced separately, a bar’s estimated high can fall below its estimated low, violating the usual ordering of those fields. Simply swapping the values may restore the ordering, but it does not preserve the original information represented by the two series.

The proposed alternative is to transform high and low into a midpoint and a range, then filter those quantities instead. This represents the bar’s central level separately from its spread and avoids the specific inversion problem described. The response offers a preprocessing suggestion rather than a validated filtering model; it does not discuss filter parameter choices, whether the range needs constraints, or how this approach affects downstream features.

Key ideas

  • Filtering high and low prices independently can produce estimates where the high is below the low.
  • Represent high and low using a midpoint and a range before filtering.
  • Filtering midpoint and range can avoid the described ordering problem.
  • The suggestion does not specify filter settings or validate effects on downstream analysis.

Tags

Full text
# OHLC prices after filtering


# OHLC prices after filtering












Assume we have minute-bars of OHLC stock prices. Then, applying Kalman filter to those prices separately, we can remove a measurement noise and obtain the estimates of the states of the price processes.

The observations is: after Kalman filtering, some high prices become smaller than low prices. Can it cause some problems?

In my opinion, it is not a problem. In a feature creation after filtering, it it happens that $High < Low$ then I would just flip them.

## Answer by Inthematrix (score 2, accepted)

https://quant.stackexchange.com/a/50390

how about preprocessing the data and create mid_price = (H+L)/2 and range = H-L. And your Kalman filter apply to mid_price and range. Then you will not have the problem as you described

Shown in full with attribution under the source's licence. Licence: CC BY-SA 4.0 (Stack Exchange)

This summary was written by Stratmill's research agent from the original; it is not a copy of the source.