Skip to content
All library documents

Choosing LASSO Regularization Windows for ETF-Hedged Statistical Arbitrage

Article Quant Q&A · Author: Eaglez

Summary

The document describes a replication of an equity statistical-arbitrage strategy in which a stock is hedged with a basket of ETFs. Regression residuals provide the trading signal, and regularized regression estimates the ETF hedge coefficients. The referenced approach re-estimates those coefficients periodically to account for changing relationships. The author chooses LASSO to favor sparse hedges and uses expanding-window time-series cross-validation to tune its regularization strength.

The central question is whether the tuning process can use a longer history than the window used to estimate the trading betas, given the limited sample available for each coefficient fit. The document offers no answer or empirical comparison, so it establishes neither the best tuning frequency nor whether a longer tuning window improves results. Any evaluation would need to preserve time ordering and consider changing relationships between stocks and ETFs, as well as the risk of selecting parameters that do not remain useful out of sample.

Key ideas

  • The strategy uses regression residuals to generate signals for stocks hedged with ETFs.
  • LASSO is selected to encourage sparse ETF hedge baskets.
  • The author tunes the regularization strength with expanding-window time-series cross-validation.
  • The document raises, but does not resolve, whether tuning can use a longer window than beta estimation.
  • Changing stock–ETF relationships may limit how well a tuned parameter carries forward.

Tags

Full text
# How often to tune the regularisation parameter in LASSO?


# How often to tune the regularisation parameter in LASSO?












I'm trying to implement the following paper: Avellaneda & Lee (2010), Statistical Arbitrage in the US equities market.

To build the strategy, the idea is to trade a stock and hedge using a basket of ETFs (the signal is based on the residuals of the stocks vs ETFs regression).

In order to estimate the betas on the ETFs, the authors suggest that a regularised regression framework could be used, re-estimating the parameters every 60 days (because these relationships will change over time).

In my attempt to replicate their results, I have chosen LASSO over Ridge/Elastic Net (to build sparse models and reduce to the minimum the ETFs to trade when hedging).

In order to tune the regularisation parameter (alpha) of the LASSO, I have used time-series cross validation (timeseries split in scikit learn, i.e. expanding window).

The problem is that given the 60 days of daily returns used to estimate the regression coefficients, I'm not sure the amount of data would be sufficient to appropriately tune alpha.

Do you think I could tune the alpha parameter on a longer window than the one used to estimate the regression coefficients?

Would be nice to hear any thoughts on this.

Thank in you in advance for your help.

Shown in full with attribution under the source's licence. Licence: CC BY-SA 4.0 (Stack Exchange)

This summary was written by Stratmill's research agent from the original; it is not a copy of the source.