Skip to content
All library documents

Modeling Sparse Stock Trades with Duration and Hawkes Processes

Article Quant Q&A · Author: Qbik

Summary

The document considers how to build a warning system for a stock index and its futures using data from both frequently and infrequently traded constituent stocks. For sparse trade data, it points to models that account for event timing as well as prices: the uncertainty-zones approach associated with Rosenbaum and Robert, and Hawkes processes, which can represent clustering in trade arrivals. It also notes that dependence estimates across asynchronously traded assets may be distorted by the Epps effect, so methods for handling that issue matter when measuring correlations.

The discussion suggests supplementing transaction data with quote information, potentially using the best limit order to form a synthetic price proxy between trades, and considering order-flow analysis. These are proposed directions rather than a tested warning-system design: the document supplies no fitted model, empirical results, or thresholds for identifying unusual activity. The intended application is to investigate whether synchronized changes in trading intensity, volume, volatility, or co-movement precede large index moves, while recognizing that the appropriate signal and its predictive value remain unresolved.

Key ideas

  • Trade durations can carry information when transaction observations are sparse and irregularly spaced.
  • Hawkes processes are suggested for modeling clustered trade arrivals.
  • Asynchronous observations can bias measured cross-asset dependence through the Epps effect.
  • Quote data and order flow may help construct price proxies between trades.
  • The proposed warning signals are exploratory and are not validated in the document.

Tags

Full text
# How to model time series of illiquid stocks - 400 observations (transactions) per 8 hours?


# How to model time series of illiquid stocks - 400 observations (transactions) per 8 hours?












How to model time series which are illiquid - 400 observations (transactions) per 8 hours ? Are there models suitable for this situation which incorporate not only size of the transactions but also their timing ? (or mayby incorporate even volume of transactions)

I'm going to edit this in the next two days and give some details about models for irregulary spaced time-series which I know, but mayby someone give interesting point based only on below remarks.

Ok, I'm going to add more details in next two days (from the statistical and econometric side), but now I have to say a little about the goal of the model. There is a stock index (call it X20) based on behavior of 8 liquid (~2200 transactions per 8 hours) and 12 (~200-400 transactions per 8 hours) illiquid stosks and there is derivative(future - FX20) based on this stock index. I need to built "warning system" for the FX20 and I want to create it using data from that 20 stocks - let leave alone other approaches based only on FX20/X20 time-series there will aslo be used. This 12 illiquid stocks make about 35% of value of index, for the 8 liquid stocks it's quite easy to produce results using common econometrics techniques GARCH/VAR for time-series (with 1-5 minute intervals) after seasonal decomposition (data for liquid stock exhibit strict U-shape daily volatility pattern) or moving averages. The goal is warning system which going to produce signals for trading system, at this point I don't know if signals from this warning system are going to be good indicators of trend reverse or mayby for spotting periods of anomalous markets activity at which trading system don't do well. And I'm expecially interested in spotting periods which precede large swings of index value. All this I want to do using data of 20 stocks. For 12 illiquid stocks synchronized increase in activity - shorter periods between transaction, increasing volume, increasing price volatility, increased correlation of price changes are signs of some sort of movement that is going to happen, and now how to quantife is it anomaly from statistical point of view (in 5 minutes time horizont) ? Loosely speaking, from statistical point of view, it's anomaly if we need extremely unlikely realization of random variables to fit that new data.

## Answer by lehalle (score 6, accepted)

https://quant.stackexchange.com/a/3535

From an academic viewpoint you do not have a lot of choices:

- The Rosenbaum-Robert approach, the price model with uncertainty zones is a model of trades and duration between trades (implicitly). It is worthwhile to try it.

- You can also use an Hawkes process, it will have the nice effect of capturing clustering effects on trades.

- if you want to use correlation / dependency measurements, you will face the Epps effect, so you should read at least one Yoshida paper on the topic.

Now if you want to go further, you should have to look also at the quotes on your illiquid stocks. You can probably find a way to build a synthetic price using the first limit that will allow you to have a proxy of the price at the next trade if there will be a trade soon on the illiquid stocks. On order-flow viewpoint can be useful that for. There is a very good suite of papers by Rama Cont and Adrien de Larrad on this topic.

Shown in full with attribution under the source's licence. Licence: CC BY-SA 4.0 (Stack Exchange)

This summary was written by Stratmill's research agent from the original; it is not a copy of the source.