Skip to content
All library documents

Testing Return Forecasts Against Bid-Ask Bounce

Article Quant Q&A · Author: cryo111

Summary

The document asks how to determine whether a one-minute return forecasting model trained on trade prices has predictive value beyond bid-ask bounce. The proposed benchmark simulates normally distributed mid-price returns using observed volatility, adds an alternating half-spread component based on the observed mean spread, then trains and evaluates an AR(1) model on repeated synthetic samples. The motivation is that the model’s lagged return predictor has high importance and a negative partial dependence pattern, which may reflect transaction-price reversals.

The proposal is framed as a question rather than an answered method: no results or validation are provided. Its benchmark depends on simplifying assumptions, including normally distributed mid returns and an unusually strong deterministic alternating bounce. Those assumptions may not represent actual trade direction, spread variation, serial dependence, or other microstructure effects. The document therefore identifies a useful diagnostic goal but leaves open how to build a representative null model, compare performance statistically, and separate genuine forecasting information from artifacts in trade data.

Key ideas

  • A negative relation between consecutive trade returns may arise from bid-ask bounce rather than underlying price predictability.
  • The proposed diagnostic simulates mid returns and adds an alternating half-spread component to create synthetic trade returns.
  • An AR(1) model on synthetic returns is intended to provide a benchmark for bounce-driven forecast performance.
  • The usefulness of the comparison depends on whether the synthetic return and spread assumptions resemble the observed market.

Tags

Full text
# Assess forecasting performance of model in presence of bid-ask bounce


# Assess forecasting performance of model in presence of bid-ask bounce












I have a forecasting model for 1-minute asset returns $y_t$ derived from trade data. (The assets are not very liquid.) The predictors of the model include the lagged target variable $y_{t-1}$, which also takes one of the top spots in the model's variable importance table. The partial dependence plot of $y_{t-1}$ shows a pronounced negative slope. My understanding is that this effect is most likely due to the bid-ask bounce, but I am not sure because I do not have a lot of experience with market microstructure. Since I am stuck with trade data at the moment, I wanted to assess to what degree the bid-ask bounce is responsible for my model's performance, or -to be more precise- to check whether my model has some predictive power left after accounting for the bid-ask bounce.

My idea is as follows:

- Generate normally-distributed synthetic mid price returns with a volatility derived from real data

- Add a bid-ask bounce from a very "strong" bid-ask bounce process on top of the mid price - I was thinking of a time series alternating between -BASpread/2 and +BASpread/2 (using the mean BASpread derived from real data)

- Train an AR(1) model on the returns of the synthetic data (with BA bounce included).

- Record the AR(1) model's out-of-sample performance on new synthetic data (same parameters as in points 1 & 2). By definition the AR(1) model will do really well in modeling the bid-ask bounce. Therefore, this should give me some sort of "benchmark" that I can use to compare my model against.

Repeat steps 1-4 for, let's say, 100 times and calculate the mean out-of-sample performance of the AR(1) models. Finally, compare the AR(1) performance to my model's performance. The assessment of whether my model has a "significantly" better performance than the AR(1) model is worth another question, I guess. For now, I hope that the outcome of this experiment is clear enough so that I can make a gut-feeling decision.

Questions

- Do you think this is a valid approach or do I miss something? If so, please elaborate.

- Given the limitations of my data, can I do better than the above in order to see whether my model has predictive power after accounting for the bid-ask bounce?

Shown in full with attribution under the source's licence. Licence: CC BY-SA 4.0 (Stack Exchange)

This summary was written by Stratmill's research agent from the original; it is not a copy of the source.