Skip to content
All library documents

Interpreting Statistical Significance in Financial Time Series

Article Quant Q&A · Author: user3138766

Summary

The document asks how to interpret p-values and statistical significance when working with financial time series, where observations are successive measurements of the same variables rather than an obvious random sample from a fixed population. It also describes calculating rolling regression slopes over short windows and testing each one, raising the question of what such tests say about a population relationship.

The response cautions that ordinary least squares results may be spurious unless the series are cointegrated or the classical time-series assumptions, especially absence of serial correlation, are met. It frames observed data as one realized path among many possible paths and notes that the stationarity of financial variables such as interest rates can be disputed. The exchange gives conceptual guidance and references standard econometrics texts, but it does not inspect the author’s data or specify a full testing procedure. Its warning is therefore general rather than a diagnosis of the particular rolling regressions.

Key ideas

  • Financial time-series observations are repeated measurements over time, not a straightforward random sample of distinct individuals.
  • Ordinary least squares inference can be spurious when time-series assumptions are violated.
  • The answer identifies cointegration and suitable classical time-series conditions as contexts for applying OLS.
  • Serial correlation can undermine standard regression significance tests.
  • A p-value from rolling regressions requires careful interpretation in light of the data-generating process.

Tags

Full text
# Statistical significance in the context of financial data?


# Statistical significance in the context of financial data?












I understand statistical significance in the general sense: we take a sample from a population and compute some parameter from the sample to infer what is the propulsion parameter to some degree of confidence, usually 95%. So if I want to find the p-value for a slope between two financial datasets, say interest rates each quarter, then I’d compute the slope between the two sample datasets and find the p-value. Less than 5%, we conclude the slope arrived at is very unlikely to have occurred by random chance. Otherwise, we fail to reject the hypothesis that the population slope is 0.

Here’s my question… if we are dealing with 200 quarters of data, what exactly is our “population?” In the commonly used example of IQs, the population is easily defined as the IQs of ALL the people. With financial data, considering quarterly data, it isn’t clear to me what is the population. Is it all historical data dating back to as long as our interest rates, in our example, existed? So if there were actually 500 quarters where our interest rates existed, that’s the population of data? Is it all of the data measured more granularly? Our interest rate measured at every second, millisecond, etc.?

I’m asking because I’m a bit perplexed how a p-value can be interpreted if we find the p-value for a slope between two datasets of interest rates, both 200 quarters worth of data. The sampling distribution which is assumed to be normal makes sense when thinking of pulling a sample of IQs from the population of all IQs.. how does it work for pulling sample last from the population of our interest rate data?

Additionally and most importantly, the reason I’m interested in this is because I am calculating the rolling slopes 24 quarters at a time, moving by one quarter at a time. If I run a t test on each slope, most are highly insignificant. If I think of these slopes as not trying to ascertain the population slope, but as simply representing the sample slope, then I think the idea of significance goes out the window here. Right? Confusing!

Sorry for the long winded question… I want to be as clear as possible where I’m confused. As always thank you all.

## Answer by AKdemy (score 3)

https://quant.stackexchange.com/a/65767

Welcome to the world of time series data.

Without knowing exactly what you did, your results are almost certainly spurious. There are only two cases when OLS is applicable:

- cointegration

- Classic linear assumptions for time series data: mainly no serial correlation needs to be satisfied

Instead of a random sample from a population, you have (many) observations of the same object over time. The observed data is also interpreted as one of infinitely many possible paths that never materialized. Many variables in finance are modelled as stochastics, and whether interest rates are stationary or not is disputed.

I can recommend a few brilliant books:

- Hamilton

- Hayashi and

- Greene

This answer has a few more details.

Shown in full with attribution under the source's licence. Licence: CC BY-SA 4.0 (Stack Exchange)

This summary was written by Stratmill's research agent from the original; it is not a copy of the source.