Indexed Data in Regression: Scaling, Stationarity, and Omitted Variables
Summary
The document considers whether regression results depend on the base year of an index or on mixing indexed and non-indexed series. It says the choice of index base is generally a rescaling and does not by itself determine whether a regression is appropriate. The discussion shifts to time-series properties and model specification as more important concerns.
A simple regression coefficient can differ from the corresponding multiple-regression coefficient when an omitted variable affects the outcome and is correlated with the included regressor. The answer derives this relationship and identifies two sample conditions under which the estimates coincide: the omitted variable has no estimated effect, or it is uncorrelated with the included variable. It also cautions that time-series regressors and outcomes often fail stationarity or exogeneity assumptions, potentially affecting inference. The response offers a workflow centered on defining the research objective, selecting and checking data, examining breaks and collinearity, and interpreting results against theory. It does not provide a worked empirical example, and its broad statement about base-year irrelevance should be read as a scaling point rather than a remedy for misspecification.
Key ideas
- Changing an index base rescales its values, so coefficient units change even when the underlying series does not.
- Mixing indexed and unindexed variables is possible, but the regression still needs a suitable specification.
- An omitted variable can distort a simple regression when it both affects the outcome and correlates with the included regressor.
- Simple and multiple regression estimates coincide in the stated sample cases when the omitted variable has zero estimated effect or is uncorrelated with the included variable.
- Time-series stationarity, exogeneity, structural breaks, and collinearity require attention before interpreting estimates.
Tags
Full text
# Regressing indexed data with non-indexed data, and varying base years?
# Regressing indexed data with non-indexed data, and varying base years?
Is it acceptable to run a regression with several independent variable datasets whose base years are different? I.e., predicting some variable y using Q4 2007 = 100 vs. Q1 1980 = 100, not in a multiple regression but using each indexed dataset as the independent variable of y one at a time? Or do you want to always have same base year?
Is it acceptable to regress indexed values with non indexed values? Q4 2007 = 100 data regressed with average quarterly data? Any issues?
My thinking is the index base years should be irrelevant. I can interpret a regression coefficient in the same way with dependent variable y and independent variable some index. A one unit increase in the particular index used corresponds to an x unit increase/decrease in y. Correct? In this sense the base years of any indexes used as independent variables should not matter, and it also should not matter if y is an index or not.
## Answer by AKdemy (score 2)
https://quant.stackexchange.com/a/65820
I pointed out some issues with regression in time series data in a previous answer.
While my HIV example used there is about causality, I did not actually think too much about causality with regards to writing the general answer.
Whatever series you use, it must be stationary. There is an interesting towardsdatascience article called How (not) to use Machine Learning for time series forecasting: Avoiding the pitfalls which shows some of the issues.
Setting TS issues aside, comparing simple and multiple regressions estimates is in itself interesting. There are only two special cases where simple regression of $y$ on $x_1$ will produce the same OLS estimate on $x_1$ as regressing $y$ on $x_1$ and $x_2$. Let's see why.
$\tilde{y}=\tilde{\beta_o}+\tilde{\beta_1}x_1$ and the multiple regression analog $\hat{y}=\hat{\beta_o}+\hat{\beta_1}x_1+\hat{\beta_2}x_2$. There is the following relationship between $\tilde {\beta_1}$ and $\hat{\beta_1}$: $$\tilde{\beta_1}=\hat{\beta_1}+\hat{\beta_2}\tilde{\phi_1}$$ where $\tilde{\phi_1}$ is the slope coefficient of the simple regression of $x_{i2}$ on $x_{i1}$, $i=1,...n$. Therefore, $\tilde{\beta_1}$ differs from the partial effect of $x_1$ on $\hat{y}$. The confounding term is the partial effect of $x_2$ on $\hat{y}$ times the slope in the sample regression of $x_2$ on $x_1$.
There are two distinct cases where they are equal:
- the partial effect of $x_2$ on $\hat{y}$ is zero in the sample ($\hat{\beta_2}=0$)
- $x_1$ and $x_2$ are uncorrelated in the sample ($\tilde{\phi_1}=0$)
Showing this omitted variable bias in general requires a bit of matrix algebra and is not important here. All I want to show is that if you assume that both play a role, leaving either out, will lead in biased estimates. That is why defining an appropriate model is actually quite difficult. Not because of causality (alone) but really because correlation alone is not what matters in regression analysis. The simple regression result of $$\frac{sample \ covariance \ of \ x \ and \ y}{sample \ variance \ of \ x}$$ only works if the two conditions above are fulfilled. Otherwise, your estimator is biased.
Base year should not matter much generally speaking as this is mainly a transformation of the data set only (as far as I know).
For interest rate modelling, I think An Investigation into Interest Rate Modelling: PCA and Vasicek is an interesting read. The best explanation of PCA I came across is found here. It shows nicely in a dynamic chart how PCA minimizes the error orthogonal (perpendicular) to the model line. OLS residuals are orthogonal to the regressors, which is an implication of the strict exogeneity assumption $E(\epsilon_i|x_1,....,x_n) =0$ which is not restrictive as long as the regressors include a constant term. It means the cross moment $E(x,y)$ of two random variables x and y is zero (which means x is orthogonal to y and vice versa). In time series (TS), this is rephrased that the regressors are orthogonal to the past, current and future regressors. For the vast majority of time series models, this condition is not satisfied. This is mainly impacting finite sample theory and it can be shown that the estimator still possesses good large-sample properties.
Edit
If your independent variable is Wilshire 5000 as the index, I would say that in itself is a big concern. This is identical to the previous question. Correlation will almost always exist and vary over time. However, this is not regression analysis.
Generally, posing a question and answering it with statistics is a delicate and complex task. I usually follow something along the line of:
- What do I try to achieve? a forecast? explain past movements in variables? value a property or firm?
- What is my hypothesis? What is the (economic) theory behind it?
- Has this been asked before somewhere? If so, what did they use and why?
- What kind of model will I need to use? GLM (OLS), ML, ARIMA,... and what functional form is best suited for this.
- What data will be needed to answer this (and satisfy the assumptions of the model of choice)
- How do I need to clean, transform and check my data before I can use it. Is it stationary? Is it noisy? Any structural breaks (Paul Volker, Great Moderation, Dot.com bubble, Subprime crisis, Covid crises to name a few)? How do I account for these regime switches?
- Is there a different factor influencing this.
- Am I at risk of omitted variable bias? Or multicollinearity?
- What tests are best suited to check for the problems above? How to check for stationarity, collinearity, ...
- Once a model is setup, data is collected and appropriately transformed, you check your results. For example, is my error term uncorrelated with me explanatory variable(s). If not, what did I miss. Redo the above.
- Am I overdoing it now (data mining)?
- How do the results compare to existing findings? How do I interpret them?Shown in full with attribution under the source's licence. Licence: CC BY-SA 4.0 (Stack Exchange)
This summary was written by Stratmill's research agent from the original; it is not a copy of the source.