Why High-Frequency Beta Regressions Can Have Low R-Squared
Summary
The document considers whether a low R-squared is surprising when estimating stock beta from intraday returns, using a regression of Google returns on S&P returns as its example. The responses explain that R-squared measures how much of one return series’ variation is accounted for by the linear relationship with the other; a low value therefore indicates a weak fit and limits how well the estimated beta describes those observations.
The discussion points to volatility changes and jumps as features of high-frequency returns that a simple linear CAPM regression may not capture. It also notes that CAPM can fit poorly at lower frequencies, so changing to intraday data does not guarantee a better relationship. Trying other regression specifications may improve fit, but risks overfitting. The exchange offers conceptual guidance rather than a tested alternative model or empirical comparison, and it does not establish that a low R-squared makes beta unusable for every purpose.
Key ideas
- R-squared describes how much variation in one return series the regression explains.
- A low R-squared indicates that a simple linear model fits the paired returns poorly.
- Intraday volatility changes and jumps can weaken the fit of a basic CAPM regression.
- Low-frequency CAPM fit can also be poor, so lower sampling frequency is not a guaranteed remedy.
- Alternative regression forms may improve fit but can introduce overfitting.
Tags
Full text
# What do you do with low r-squared when calculating high-frequency beta # What do you do with low r-squared when calculating high-frequency beta I am calculating a high-frequency beta. For example I have 90 days of data of the S&P and GOOGLE and I have 10-minute percent returns for each instrument. Each day has 34 10-minute percent returns so my data set is 2 vectors that are both 3060 in length (90 days x 34 10-miunute percent returns each day) = 3060 data points for the S&p and 3060 data points for GOOGLE. Next in R I run a regression ``` reg= lm(google~snp) # both the google and snp vector have lenght = 3060 summary(reg) ``` My question is that sometimes I get low R-squared for the regression. Is this expected...should/can anything thing be done about it? I know beta is usually done using daily data NOT 10-minute data but even with daily data sometimes the r-squared is low. What is the significance of beta with a low r-squared? Thank you. ## Answer by user32416 (score 3) https://quant.stackexchange.com/a/20837 This simply suggests the linear model is a poor fit in high frequency. But is this that surprising, even before you crunch the numbers? I argue not, for the following reasons: - Even at low frequencies (i.e. monthly or annually), it is known that the classical CAPM (which is what you're running, albeit at a much higher frequency) does not fit well. It'd be a truly miracle that the CAPM model that performs poorly in low frequencies would even work any better in high frequencies. - It is also well known that high frequency financial econometrics exhibit a lot of behavior that are not found in their low frequency equivalents. Just a few to think about --- stochastic volatility, jumps, etc., and these properties of a single asset itself, without regard to how it co-moves with another asset (say S&P). All these imply that a simple linear model is expected to fail. ## Answer by amsh (score 2) https://quant.stackexchange.com/a/20836 A high R-squared (1.0) means that you can explain the movements of one time series using the other. The lower your R-squared is, the worse your explanation is -- that includes the 'quality' of your beta. You can try to improve your R-squared score using different regression types. Beware of overfitting.
Shown in full with attribution under the source's licence. Licence: CC BY-SA 4.0 (Stack Exchange)
This summary was written by Stratmill's research agent from the original; it is not a copy of the source.