Skip to content
All library documents

Applying Linear Regression to Stock Returns in Python and R

Article QuantInsti blog

Summary

The article demonstrates simple and multiple linear regression on historical returns for Coca-Cola, PepsiCo, the S&P 500 ETF, and the US Dollar Index. It first uses pairwise correlations, then fits a single-predictor model for Coca-Cola returns using the S&P 500 ETF and a multiple-predictor model adding PepsiCo and the dollar index. Reported regression tables give coefficients, significance statistics, and fit measures for the Python statsmodels implementation. The tutorial also describes equivalent prediction-focused fits with scikit-learn and reproduces the workflow in R.

The comparison explains that statsmodels emphasizes statistical inference and detailed summaries, while scikit-learn is oriented toward fitting and prediction on unseen observations. The numerical outputs are an illustration on a historical sample, not evidence that the relationships will persist or yield a profitable strategy. The author explicitly flags regression assumptions and limitations but defers their discussion; readers should therefore examine those assumptions, data differences, and out-of-sample performance before using the models for forecasting or trading.

Key ideas

  • The tutorial models Coca-Cola returns using market, competitor, and dollar-index returns as explanatory variables.
  • Pairwise correlations provide an initial view of relationships, while regression estimates conditional linear associations.
  • statsmodels supplies inference-oriented summaries, whereas scikit-learn is presented as prediction-focused.
  • Python and R implementations are compared, with small output differences attributed to data differences across interfaces.
  • Historical fit statistics do not establish forecasting power, and regression assumptions require separate scrutiny.

Tags

This summary was written by Stratmill's research agent from the original; it is not a copy of the source.