Skip to content
All library documents

Handling Tied Tick Timestamps in Hawkes Process Calibration

Article Quant Q&A · Author: Timcy Arora

Summary

The document examines why maximum-likelihood estimates for an exponential Hawkes process can drive the decay rate and excitation parameter toward very large values when tick records share millisecond timestamps. In the stated likelihood recursion, zero inter-event gaps create repeated excitation without elapsed time, allowing the likelihood to keep increasing while the branching ratio remains bounded. A rate constraint or random timestamp jitter can therefore determine the apparent estimate instead of identifying a decay rate from the data.

The answer proposes treating each timestamp as one marked event, with a mark such as fill count or volume, or modeling arrivals at their observed time resolution with a binned count process. It reports a comparison of raw, jittered, aggregated, and tie-collapsed data, including a case where the marked fit has an interior maximum. These examples support the diagnosis but depend on the chosen data and model. The response also cautions that a single exponential kernel may poorly represent market activity across multiple time scales; its branching-ratio estimates are sensitive to preprocessing and kernel choice.

Key ideas

  • Tied event times can make the exponential Hawkes likelihood increase without bound as decay and excitation parameters grow together.
  • A bound or jitter width can set the estimated decay rate rather than reveal it from the observations.
  • Aggregating simultaneous records into marked events preserves their shared timestamp while representing event multiplicity.
  • A binned count-process model is another way to account for finite timestamp resolution.
  • Kernel choice affects estimated branching ratios, especially when market activity spans several time scales.

Tags

Full text
# Answer by dikovaxi (score 0)


# Hawkes process MLE calibration diverges (β → bound) on tick data with millisecond-tied timestamps. Is timestamp jittering the standard fix?












I'm fitting a univariate Hawkes process with exponential kernel to real tick-level trade data (Binance BTCUSDT, ~1.25M trades/day, millisecond timestamps). The intensity is:

$$\lambda(t) = \mu + \sum_{t_i < t} \alpha e^{-\beta(t - t_i)}$$

I'm maximizing the exact log-likelihood via the standard recursive form (Ozaki 1979):

$$R_i = \sum_{j<i} e^{-\beta(t_i - t_j)}, \qquad R_i = e^{-\beta(t_i - t_{i-1})}\left(1 + R_{i-1}\right)$$

$$\log L = \sum_i \log\left(\mu + \alpha R_i\right) - \mu T - \frac{\alpha}{\beta}\sum_i\left(1 - e^{-\beta(T - t_i)}\right)$$

optimized with `scipy.optimize.minimize` (L-BFGS-B), enforcing $\mu, \alpha, \beta > 0$ and $\alpha < \beta$ (stationarity) via bounds/penalty.

Problem: without a bound on $\beta$, the optimizer diverges $\alpha, \beta$ both run to $10^6$+ while $\alpha/\beta$ (branching ratio) stays sensible (~0.6). With a bound on $\beta$ (e.g. $\beta \le 100$), the fit simply pins $\beta$ at exactly the bound rather than converging to an interior value indicating the optimizer still wants to go higher, not that 100 is the right answer.

I believe the cause is tied event times: many trades share the same millisecond timestamp (one market order filling against several resting limit orders logged separately), which violates the continuous-time assumption underlying the likelihood, effectively rewarding the optimizer for treating simultaneous events as an infinitely fast excitation burst.

My questions:

- Is tie-breaking (e.g. jittering same-timestamp events by a small uniform offset within the timestamp resolution) the standard/correct fix here, or is there a more principled likelihood correction for tied arrivals in Hawkes estimation?

- Is there a standard reference for handling this in the market-microstructure Hawkes literature (tick data specifically), as opposed to the general point-process literature?

- Is a hard upper bound on $\beta$ ever appropriate as anything beyond a stopgap, or does its necessity always indicate a data problem (like ties) rather than a legitimate model constraint?

## Answer by dikovaxi (score 0)

https://quant.stackexchange.com/a/85819

Short version. No — jittering is not the fix, and the bound is not a stopgap either: with tied timestamps the exponential-kernel MLE does not exist, so any device that makes the optimiser stop (a bound, a jitter width) simply becomes the estimate. The principled fix is to make the timestamp define the event: collapse records that share a timestamp into one event with a mark (fill count or volume), which is also what physically happened — one taker order matched against several resting orders. After that the likelihood has an interior maximum. I reproduced your situation on the same market and public data (Binance USD-M BTCUSDT `trades` and `aggTrades`, 2024-03-26, data.binance.vision) so the numbers below can be re-run.

Why the likelihood diverges. In the Ozaki recursion a tie ($t_i=t_{i-1}$) gives $R_i = 1+R_{i-1}$ exactly, so for a tied event $\log(\mu+\alpha R_i)\ge\log\alpha$, while the compensator $\frac{\alpha}{\beta}\sum_i\big(1-e^{-\beta(T-t_i)}\big)\approx \frac{\alpha}{\beta}\,N$ stays bounded when $\alpha,\beta\to\infty$ at fixed $\alpha/\beta$. The log-likelihood therefore increases without limit along that ray — which is precisely the "$\alpha,\beta\to10^6$ with a sensible branching ratio" you observe. It is not the optimiser; there is no maximum to find.

How tied the data really is (whole day, 4.47 M `trades`, 1.80 M `aggTrades`): 84% of trade records share their millisecond with another record, 74% of inter-event gaps are exactly zero (groups average 3.8 fills, p99 = 38, max 371). `aggTrades` does not solve it — 66% still tied, 56% zero gaps — because a market order that walks several price levels produces one aggregated record per level, all with the same timestamp.

Fits on a 20-minute window (08:00–08:20 UTC; exact likelihood, L-BFGS-B on log-parameters, $\beta$ free up to $10^7$):

| data | events | $\hat\beta$ (1/s) | $\hat\alpha/\hat\beta$ |
| raw trades | 115,771 | runs to the bound | 0.73 |
| raw trades, $\beta\le100$ imposed | 115,771 | 100 (pinned) | 0.87 |
| raw trades + jitter U(0, 1 ms) | 115,771 | 4,830 | 0.76 |
| raw trades + jitter U(0, 10 ms) | 115,771 | 623 | 0.79 |
| aggTrades | 44,146 | runs to the bound | 0.57 |
| one event per ms, mark = fill count | 19,105 | 52.5 (interior) | 0.21 |
| same, 3-exponential kernel | 19,105 | 177 / 6.6 / 0.11 | 0.69 |

The profile log-likelihood over $\beta$ (with $\mu,\alpha$ re-optimised at each value) tells the same story: for raw trades and for `aggTrades` it rises monotonically from $\beta=1$ to $10^6$ (by 8.5×10⁵ and 2.6×10⁵ log-likelihood units respectively), whereas for the tie-collapsed marked data it peaks at $\beta\approx10^2$ and falls by 1,200 by $\beta=10^3$. Two things to read off the table:

- Jitter sets the answer. $\hat\beta$ is about 5 / (jitter width): 1 ms gives ~5,000, 10 ms gives ~600. You are not estimating a decay rate, you are estimating your own uniform distribution. The same is true of the bound: at $\beta\le100$ you get "0.87" for the branching ratio, a number with no content.

- Collapsing ties gives a well-posed problem — but note what happens to the branching ratio: 0.21 with one exponential, 0.69 with three (time scales ≈ 6 ms, 150 ms, 9 s). A single exponential fitted to tick data is badly misspecified — real kernels are slowly decaying, roughly power-law over many decades (Bacry, Dayri & Muzy 2012; Hardiman, Bercot & Bouchaud 2013) — and a single scale absorbs whichever part of the kernel the data forces on it. With ties present, that is the zero-lag spike; without them, the sub-second cluster, and the slow mass is missed. If the branching ratio is the quantity of interest, use a sum of exponentials (three or four scales from ms to minutes) or a non-parametric kernel; the exponential value is not comparable across papers or preprocessing choices.

Your three questions.

- Jitter or something more principled? Jitter is used (and is what several toolkits do quietly), but it is a prior on $\beta$ dressed as preprocessing. The principled options are (a) aggregate same-timestamp records into one marked event — the intensity becomes $\lambda(t)=\mu+\sum_{t_j<t}\alpha\,m_j e^{-\beta(t-t_j)}$ and the recursion is $R_i=e^{-\beta(t_i-t_{i-1})}(m_{i-1}+R_{i-1})$; or (b) accept that the data are discrete and fit at the resolution you have — bin the timeline and estimate the kernel as an INAR(p) (Kirchner 2017), which handles simultaneity by construction. Sub-millisecond excitation is not identifiable from millisecond stamps under any method.

- References for tick data specifically: Bowsher (2007, J. Econometrics) on trades/quotes as multivariate Hawkes with same-timestamp handling; Filimonov & Sornette (2015, Quantitative Finance) "Apparent criticality and calibration issues in the Hawkes self-excited point process model", which treats timestamp discretisation and the resulting bias in $\hat\beta$ and the branching ratio; Lallouache & Challet (2016, Quantitative Finance) "The limits of statistical significance of Hawkes processes fitted to financial data", on how the timestamp resolution bounds the shortest identifiable kernel scale; Bacry, Mastromatteo & Muzy (2015) for the review; Bacry, Dayri & Muzy (2012) and Hardiman, Bercot & Bouchaud (2013) for the kernel shape.

- Is a hard bound ever legitimate? Only as a diagnostic: if $\hat\beta$ sits on any bound you choose, the MLE does not exist for that data/model pair and the estimate is the bound. It always means ties (or a zero-lag spike the exponential cannot represent), never a model constraint.

```
# marked, tie-free recursion (t in seconds, m = fills per timestamp)
t, m = np.unique(t_ms, return_counts=True); t = t / 1e3
R = 0.0; ll = 0.0
for i in range(len(t)):
    if i: R = np.exp(-beta * (t[i] - t[i-1])) * (m[i-1] + R)
    ll += np.log(mu + alpha * R)
ll -= mu * T + (alpha / beta) * np.sum(m * (1 - np.exp(-beta * (T - t))))
```

Shown in full with attribution under the source's licence. Licence: CC BY-SA 4.0 (Stack Exchange)

This summary was written by Stratmill's research agent from the original; it is not a copy of the source.