Estimating Avellaneda–Stoikov Fill-Rate Parameters from Order-Book Events
Summary
The document discusses empirical estimation of the Avellaneda–Stoikov order-arrival curve, commonly written as a fill intensity that declines exponentially with quote distance. It relates the decay parameter k to the market-order size distribution and the relationship between order size and price impact. For estimation, it outlines recording event times and quote distances, grouping observations by distance, and fitting exponential interarrival times to estimate arrival rates. Those rates can then be used to fit the curve’s scale A and slope k, either with a regression or with two selected observations.
The answer highlights practical pitfalls: recording the time a later price level is reached measures jump rates rather than necessarily measuring fills; rates should decline with quote distance; and sign conventions in log-linear fitting can produce a negative slope unless handled correctly. The resulting estimates describe historical behavior and may not predict future fills or market moves. The document does not establish that one fitting approach is universally superior; the two-point method is sensitive to the chosen observations.
Key ideas
- Estimate arrival intensity by grouping observed events according to quote distance and fitting interarrival times.
- Fit the exponential intensity curve to obtain its scale A and distance-decay parameter k.
- Events defined by a subsequent price-level move can measure jump rates instead of fill rates.
- Arrival rates should decline as quote distance increases for the exponential model to be coherent.
- Historical estimates describe past activity and may not forecast future market behavior.
Tags
Full text
# Avellaneda-Stoikov empirical estimation verification
# Avellaneda-Stoikov empirical estimation verification
The solution of the model contains constant: $k = \alpha K$, it relates to: (i) probability of getting a fill ($\alpha$) and (ii) market impact ($K$).
Estimating (i). The author proposes that the market order sizes follow a power law distribution:
$$f(x)^Q \propto x^{-1-\alpha}$$
So no problem estimating this.
Estimating (ii). There are two equations of interest here:
$$\Delta p \propto \ln Q \tag{11}$$
$$\begin{align} \lambda (\delta) &= \Lambda \mathbb{P} [\Delta p > \delta] \\ &= \Lambda \mathbb{P}[\ln Q > K \delta] \\ &= \Lambda \mathbb{P} [Q > \exp \left( K \delta \right)] \\ &= \Lambda \int_{\exp (K \delta)}^{\infty} x^{-1-\alpha} dx \\ &= A \exp \left( -k \delta \right) \end{align} \tag{12}$$
According to the above two:
Line 1 in (12) says $\Delta p > \delta$, but we know that $\Delta p = c \ln Q$, therefore we can re-write:
$$\begin{align} \mathbb{P} [\Delta p > \delta] &= \mathbb{P} [c \ln Q > \delta] \\ &= \mathbb{P} [\ln Q > \frac{1}{c} \delta] \end{align}$$
Therefore, $K = \frac{1}{c}$, the inverse of the proportionality constant that we find in relation $\Delta p \propto \ln Q$.
Is this correct? (The estimation of $K$).
Note: I am aware of Sophie Laruelle's paper. But, unfortunately, I do not know French.
Super Note: My understanding of this answer relating to Sophie's paper:
Definitions:
- $t_0$ - start time (this is the start time point in the interarrival time slots in a Poisson Process). Simply put, if you get hit, that is a new time: $t_1$, which becomes your new start time.
- $\delta P$ - distance of your order from the mid price. If the mid price is \$100, and your bid is at \$85, then this quantity is \$15. Same for asks, obviously.
- $P^m (t) $ - is your mid price at time $t$.
Procedure:
- Record the time $t_1$ when the bid/ask is hit and the $\delta P$s. The reason I have added plural "s" ending in there is because you might get hit at a number of levels. An example is asked for here. This is your order book (OB):
```
|--- qty ---| --- size --- |
asks
$105 10
$104 5
bids
$100 6
$99 5
```
Imagine a market order comes in that eats 8 units on the bid size (so it was a market sell). You record the change in price, $\delta P = \$102 - \$100 = \$2$ at first level and the corresponding time of this trade, $t_1$. You also record $\delta P = \$102 - \$99 = \$3$ at the same time, $t_1$ (in this case). If there was a market order that only took portion of the best ask/bid, or full best ask/bid and did not go deeper, we would only collect a single $\delta P$ with its corresponding $t$. Note that you can define the point at which the order is hit differently, this is up to you. - You now have time lengths and corresponding sizes. Interarrival times in a Poisson Process are exponential:
$$\mathbb{P}[X_1 > t] = \mathbb{P}\left[ \texttt{no arrival in time (0, t]} \right]= e^{-\lambda t} $$
- So now for each bucket of price changes like: $[\$1, \$1.5]$ i.e. all of the $\delta P$ that are between 1 and 1.5 bucks, you have a list of interarrival times: $[0.3, 0.2, 0.5, 0.7, 1.1]$ in seconds. Fit the above exponential distribution to each data bucket to obtain some empirical estimate of $\lambda$.
- You now have figure 1 from Sophie's paper. Well done.
Problems ahead:
- What now? So the question is, if you fit the regression to this, what is your $k$, what is your $A$?
- Not so important. What is this equation referring to, what is $P_1$, what is $P_2$, what is $P$:
$$k = \mathbb{E}_{P1,P2}\left(\frac{\log\lambda(\Delta P1) - \log(\lambda(\Delta P2)}{\Delta P1 - \Delta P2}\right),\quad A=\mathbb{E}_{P}(\lambda(\Delta P) \exp k \Delta P)$$
## Answer by wildbunny (score 3, accepted)
https://quant.stackexchange.com/a/45280
Your post updates correctly address most of the problem.
The key is producing a graph like the one in Sophie's paper, once you have that you can do a regression to find $A$ and $k$ in
$$ \lambda(δ) = Ae^{(-kδ)}\, $$
Intuitively, what you are graphing is a realisation of the above equation, so all you need to do is determine the coefficients.
edit: I started with a log-linear regression and although it seemed ok at first, it was very poor for small tick sizes, it might actually be better to use the other method, which is
$$k = \left(\frac{\log(\lambda(δ2) / \lambda(δ1))}{δ2 - δ1}\right),\quad A=\lambda(δ1) e^{(kδ1)}$$
Where δ1 and λ(δ1) are tick value 1 and rate value 1, so you're essentially just building an exponential curve guaranteed to pass through the values corresponding to the two ticks you arbitrarily pick.
Some things to watch out for:
1) If you record a fill when the next price level gets hit, you're not recording fill rates, but jump rates
2) The fill rates must be decreasing with $δ$ otherwise the regression wont make any sense and you'll get bad values out.
3) $k$ must be positive, so you'll need to negate it if you do a log level regression
4) What you're recording here is essentially historical alpha, which may or may not be useful in any way for predicting what the market is going to do, rather than what it has done.Shown in full with attribution under the source's licence. Licence: CC BY-SA 4.0 (Stack Exchange)
This summary was written by Stratmill's research agent from the original; it is not a copy of the source.