Transforming Macroeconomic Variables for Return Regressions
Summary
The document considers how to transform quarterly macroeconomic predictors when regressing quarterly index returns on inflation, interest rates, unemployment, and exchange rates. The accepted response recommends considering percentage changes for variables and return transformations for the dependent series, while comparing candidate models using measures such as fit and information criteria. It also suggests that logarithmic returns may better support regression assumptions in some circumstances and notes the relative interpretability of log-log specifications.
A second response emphasizes stationarity, consistent variable definitions, and checks for multicollinearity, serial correlation, and heteroskedasticity. The discussion warns that a quarterly sample with 96 observations is small and may limit reliability, suggesting a longer or higher-frequency dataset where feasible. These are broad recommendations rather than a demonstrated empirical result. Transformation choices should depend on each variable’s meaning and time-series properties; blanket percentage changes may be inappropriate for rates or other series, and normality is an assumption to diagnose rather than guarantee through a return transformation.
Key ideas
- The question concerns selecting transformations for macroeconomic predictors of quarterly index returns.
- The responses suggest testing percentage changes and return-based specifications, then comparing model performance.
- Stationarity and regression diagnostics such as serial correlation and heteroskedasticity should be examined.
- A sample of 96 quarterly observations may provide limited evidence for model selection.
- Transformations should reflect each series’ properties, since percentage changes are not automatically suitable for every variable.
Tags
Full text
# Transforming Variables in Regression # Transforming Variables in Regression I have a very simple problem that hopefully someone could help me with or at least point me in the right direction. I am testing to see which factors affect index returns the most and would like to find the correct way to transform variables used in the multiple regression model. The data sample covers the time period 1990 - 2013 with a quarterly frequency (96 observations). For my dependent variable, I have calculated quarterly index returns. For the independent variables set, I have collected quarterly data for: CPI%, 3 year rate, 10 year rate, unemployment rate, and a couple of exchange rates. Should I calculate quarterly % changes for all of my variables(since I used quarterly returns for my dependent), leave them unchanged(and use index values instead of returns), or change them in a different way? Thank you, Here is a screenshot of my current data set: ## Answer by Quantopik (score 1, accepted) https://quant.stackexchange.com/a/17536 I suggest you to implement all the analysis you cited above and analyze the results choosing the best model on model performance measures, as, for instance, the model $R^2$ value, the AIC (BIC) value, etc.; this should be the proper way to develop a model. As regards your question particularly, the literature about the topic suggest to develop the model on the basis of the change in percentages of all variables you consider developing the factor model. In the matter of the transformation of the dependent variable, you should consider the possibility of transforming that in returns to fulfill the assumptions of the linear regression model; indeed, the logarithmic returns are distributed according to the Normal distribution if they are in the range [-0.04; +0.04] according to the Taylor approximation, and it is more likely to produce normally distributed residuals; the values outside this range should be considered as outliers. The same for independent variable; you should transform those in change in percentage variable, if they're not yet. Lastly, it is easier to analyze log-log regression model results than level-log or log-level model. As last hint, I firstly suggest you to develop a model taking into account the linear regression model assumption and testing for the reliability of those hypothesis after running the model (test for homoschedasticity, autocorrelation, multicollinearity,...). Moreover, try to build a reliable dataset; using quarterly data you can have less observation and get biased results. So, if possible, use higher frequency data than yours, as, for instance, monthly, weekly, daily data. EDIT: since your dataset is pretty small (~96 observations), you should try to get more data to improve your analysis by collecting data previous to the 1990 or later of 2013. ## Answer by Barnaby (score 2) https://quant.stackexchange.com/a/17524 Basic rules for what you want to follow to have a meaningfull result ``` 1 All your variables should be in the same terms 2.All your regression variables should be stationary (weak sense stationary) - check if there are long term dependencies 3.Test for multi-collinearity / Serial correlation / Homocedasticity each variable against the predictor. ```
Shown in full with attribution under the source's licence. Licence: CC BY-SA 4.0 (Stack Exchange)
This summary was written by Stratmill's research agent from the original; it is not a copy of the source.