Skip to content
All library documents

Why High-Frequency Beta Regressions Can Have Low R-Squared

Article Quant Q&A · Author: joesyc

Summary

The document considers whether a low R-squared is surprising when estimating stock beta from intraday returns, using a regression of Google returns on S&P returns as its example. The responses explain that R-squared measures how much of one return series’ variation is accounted for by the linear relationship with the other; a low value therefore indicates a weak fit and limits how well the estimated beta describes those observations.

The discussion points to volatility changes and jumps as features of high-frequency returns that a simple linear CAPM regression may not capture. It also notes that CAPM can fit poorly at lower frequencies, so changing to intraday data does not guarantee a better relationship. Trying other regression specifications may improve fit, but risks overfitting. The exchange offers conceptual guidance rather than a tested alternative model or empirical comparison, and it does not establish that a low R-squared makes beta unusable for every purpose.

Key ideas

  • R-squared describes how much variation in one return series the regression explains.
  • A low R-squared indicates that a simple linear model fits the paired returns poorly.
  • Intraday volatility changes and jumps can weaken the fit of a basic CAPM regression.
  • Low-frequency CAPM fit can also be poor, so lower sampling frequency is not a guaranteed remedy.
  • Alternative regression forms may improve fit but can introduce overfitting.

Tags

Full text
# What do you do with low r-squared when calculating high-frequency beta


# What do you do with low r-squared when calculating high-frequency beta












I am calculating a high-frequency beta. For example I have 90 days of data of the S&P and GOOGLE and I have 10-minute percent returns for each instrument. Each day has 34 10-minute percent returns so my data set is 2 vectors that are both 3060 in length (90 days x 34 10-miunute percent returns each day) = 3060 data points for the S&p and 3060 data points for GOOGLE.

Next in R I run a regression

```
reg= lm(google~snp) # both the google and snp vector have lenght = 3060
summary(reg)
```

My question is that sometimes I get low R-squared for the regression. Is this expected...should/can anything thing be done about it?

I know beta is usually done using daily data NOT 10-minute data but even with daily data sometimes the r-squared is low. What is the significance of beta with a low r-squared?

Thank you.

## Answer by user32416 (score 3)

https://quant.stackexchange.com/a/20837

This simply suggests the linear model is a poor fit in high frequency. But is this that surprising, even before you crunch the numbers? I argue not, for the following reasons:

- Even at low frequencies (i.e. monthly or annually), it is known that the classical CAPM (which is what you're running, albeit at a much higher frequency) does not fit well. It'd be a truly miracle that the CAPM model that performs poorly in low frequencies would even work any better in high frequencies.

- It is also well known that high frequency financial econometrics exhibit a lot of behavior that are not found in their low frequency equivalents. Just a few to think about --- stochastic volatility, jumps, etc., and these properties of a single asset itself, without regard to how it co-moves with another asset (say S&P). All these imply that a simple linear model is expected to fail.

## Answer by amsh (score 2)

https://quant.stackexchange.com/a/20836

A high R-squared (1.0) means that you can explain the movements of one time series using the other. The lower your R-squared is, the worse your explanation is -- that includes the 'quality' of your beta.

You can try to improve your R-squared score using different regression types. Beware of overfitting.

Shown in full with attribution under the source's licence. Licence: CC BY-SA 4.0 (Stack Exchange)

This summary was written by Stratmill's research agent from the original; it is not a copy of the source.