Skip to content
All library documents

Why Regression Interactions Use Products of Predictors

Article Quant Q&A · Author: adarshad

Summary

The note explains that a product term in a regression represents an interaction: the association between one predictor and the outcome changes with the value of another predictor. For two predictors, including their product allows the fitted slope for one predictor to vary with the other. An additive model alone cannot express that changing slope.

It also shows conceptually that specifying an interaction in formula notation is equivalent to explicitly creating the product variable and including it alongside both main effects. The discussion cautions that main-effect coefficients in an interaction model are conditional on the other predictor being zero or at its reference level. Centering continuous predictors can make those coefficients easier to interpret. The examples are simulated and illustrate model equivalence; they do not establish that an interaction is appropriate for a particular volatility study. A higher R-squared by itself is not evidence that the interaction is substantively justified.

Key ideas

  • A product term models how the association of one predictor depends on another predictor’s value.
  • An additive model cannot represent a slope that changes with another predictor.
  • Formula notation for an interaction is equivalent to explicitly including the product variable.
  • Main-effect coefficients are conditional on the other interacting predictors’ values.
  • Centering continuous predictors can make main effects easier to interpret.

Tags

Full text
# Cross Effect in OLS


# Cross Effect in OLS












I am using cross effect in OLS regression for a time series problem for a multivariate regression. I want to quote reference for use of cross effect. Secondly, I want to explain why better to use multiplication /product in regression rather than addition for calculation of cross effect. My regression equation is:

$Y=intercept+ ax_1 + bx_2 + a*bx_3$+error term

I need to see a reference for interaction based cross effect using multiplication to detail a regression effect in a volatility study. I hope you can understand and reply to my query.

## Answer by Joe King (score 1)

https://quant.stackexchange.com/a/83837

> I want to know why multiplication or interaction is better than just additive regression. Is there any literature review.Like we assume that interaction exists when we have higher Rsquared value for model with multiplication or cross effect than simple additive model? can you explain?

If I am understanding your question, you are asking why interaction terms are multiplicative rather than additive. Interaction terms are multiplicative because they capture how the association of one predictor with the outcome (it might help to consider this association as a slope, which of course it is) depends on the value of another predictor (or another interaction in the case of 3-way and higher order interactions). If we had a binary variable (i.e., a grouping variable) interacting with continuous one, and if we plotted the data for both groups we would see that they have different slopes. This can only happen with a multiplicative interaction term $x_1\times x_2$.

We are not adding effects here; we are modulating the effect of one variable by the value of the other. That's why it’s multiplicative — we literally include the product `$x_1x_2$ in the model.

A simple example might help. The idea is to take a simple (simulated) dataset with an outcome and 2 variables. We first fit a model where the interaction is specified in the formula (e.g., in R, we specify the interaction using `x1:x2`) and then we fit a model excluding the interaction term but including a new variable that is the product of `x1` and `x2`. We will see that the models are identical,

First in R:

```
set.seed(15)
n <- 100
x1 <- rnorm(n)
x2 <- rnorm(n)
y  <- 1 + 2*x1 + 3*x2 + 4*x1*x2 + rnorm(n, sd = 0.5)

mydata <- data.frame(x1, x2, y)

# Fit model with interaction using formula notation
model1 <- lm(y ~ x1 + x2 + x1:x2, data = mydata)

# Create explicit interaction variable
mydata$x3 <- mydata$x1 * mydata$x2

# Fit model using x1, x2, and explicit interaction term
model2 <- lm(y ~ x1 + x2 + x3, data = mydata)

# Compare the two models (if you want)
# summary(model1)
# summary(model2)

# Confirm they are equivalent
all.equal(unname(coef(model1)), unname(coef(model2)))

Which produces

> # Confirm they are equivalent
> all.equal(unname(coef(model1)), unname(coef(model2)))
[1] TRUE
```

And for the Pythonistas:

```
import numpy as np
import pandas as pd
import statsmodels.formula.api as smf

# Simulate data
np.random.seed(123)
n = 100
x1 = np.random.normal(size=n)
x2 = np.random.normal(size=n)
y = 1 + 2*x1 + 3*x2 + 4*x1*x2 + np.random.normal(scale=0.5, size=n)

# Create DataFrame
df = pd.DataFrame({'x1': x1, 'x2': x2, 'y': y})

# Fit model with interaction using formula
model1 = smf.ols('y ~ x1 * x2', data=df).fit()

# Create explicit interaction term
df['x3'] = df['x1'] * df['x2']

# Fit model with explicit interaction variable
model2 = smf.ols('y ~ x1 + x2 + x3', data=df).fit()

# Compare both models if you want
#print(model1.summary())
#print(model2.summary())

# Check if coefficients are equal
print(np.allclose(model1.params, model2.params))
```

which produces `true`

Lastly it is worth noting that, while the same arguments as above apply, things get a little more tricky when we have continuous x continuous interactions, higher order (i.e., 3-way and above) interactions, and it is often advised to centre the continuous variables that are involved in an interaction, because the interpretation of the main effects change when a variable is involved in an interaction - the interpretation of the main effectof one variable is conditional on the other variable it is interacting with is held at zero (or at it's reference level if it is categorical), which would not make sense in many scenarios (like a bond price being zero)

Shown in full with attribution under the source's licence. Licence: CC BY-SA 4.0 (Stack Exchange)

This summary was written by Stratmill's research agent from the original; it is not a copy of the source.