Skip to content
All library documents

Using Nonlinear Factor Models to Estimate Portfolio Risk

Article Quant Q&A · Author: Winger 14

Summary

The document considers how machine learning could extend cross-sectional factor models used to explain security returns and estimate portfolio variance. It suggests modeling factor exposures as functions of observable security characteristics, such as size, valuation, or dividend yield, while keeping returns as a combination of factor exposures, factor realizations, and residual risk. In that setup, portfolio variance retains the familiar factor covariance structure, but the factor covariance is estimated from the learned characteristic functions.

For risk prediction, the learning objective can explicitly include out-of-sample risk performance. The discussion also distinguishes unsupervised nonlinear factor discovery from supervised models that use stock characteristics, and recommends examining residuals after fitting linear factors to identify additional structure. It proposes variance-based objectives and, potentially, objectives that capture skewness or other distributional features. The source offers modeling guidance and points to published asset-pricing research, but provides no empirical comparison, implementation details, or evidence that nonlinear factors improve risk forecasts. It also cautions that nonlinear models can reduce transparency and complicate covariance estimation.

Key ideas

  • Model factor exposures as functions of security characteristics rather than treating all exposures as fixed linear loadings.
  • Portfolio variance can preserve a factor covariance form when learned characteristic functions define the exposures.
  • A training objective can target out-of-sample portfolio risk as well as return explanation.
  • Nonlinear factors can be discovered without labels or supervised using security characteristics as inputs.
  • Fitting nonlinear structure to residuals from linear factors helps test whether it adds explanatory value.
  • Risk objectives may account for distributional properties beyond variance, though the source gives no performance evidence.

Tags

Full text
# ML in Factor Models


# ML in Factor Models












I have recently learned about (implicit) factor models of the form:

$$ R = Xf + \epsilon $$

where $R \in \mathbb{R}^{n}$ are security returns, $X \in \mathbb{R}^{n \times F}$ are factor loadings for each security and each of $F$ factors and we fit a regression to get the estimated $f$.

This is also called cross-sectional regression.

Then, we compute factor covariances $\Omega := Cov(F_i,F_j)_{i,j=1,...,F}$ and can compute the volatility of the return of a portfolio $P$ with weights $w \in \mathbb{R}^{n}$ as:

$$Var(R_P) = w^T X^T \Omega X w + Var(\epsilon)$$

Someone mentioned that we could use ML-methods to fit more sophisticated models $$R = \phi(X,f) + \epsilon$$ I guess this means that $\phi$ might be a neural network or some other model class, mapping data input $x_i \in \mathbb{R}^F$ to estimated returns $R_i$ via to-be-found parameters $f$ (i.e. weights and biases of a neural network).

However, I am wondering how this might help getting a more precise estimate on $Var(R_P)$, after all $Var(\phi(X,f))$ is not easy to compute for complicated (nonlinear) $\phi$ and transparency might be an issue.

## Answer by lehalle (score 4)

https://quant.stackexchange.com/a/68171

Not sure try to fit $\phi(X,f)$ makes sense: how would you define $X$? For a linear model $X$ is naturally defined as the beta of a linear regression.

Probably you would rather need to go "one layer deeper in the definition of factors", i.e. to use characteristics of the stock (like its market cap, P/E, dividend yield, etc) that you will name $C_1,\ldots,C_N$ and now the goal is to find a neural network (or any machine learning algo) such that

$$R = \sum_{1\leq i\leq F} X_i\cdot\psi_i(C_1,\ldots,C_N)+\epsilon.$$

This is in brief what is proposed in "Deep learning in asset pricing" by Chen, Pelger, and Zhu (2020).

That being said, what you have in mind now is specifically to focus on out of sample risk prediction. This can be done by adding a term in your loss function. What you note now is that in this better formulation, the risk of your portfolio continues to have this kind of shape

$$Var(R_P) = w^T X^T \Omega X w + Var(\epsilon)$$

where the expression of $\Omega$ is now different: $$\Omega := Cov(\psi_i(C_1,\ldots,C_N),\psi_j(C_1,\ldots,C_N))_{i,j=1,...,F}.$$

Note that the covariance is a scalar product, and hence if your learning algorithms $(\psi_i)_i$ are kernel based, you could use the "kernel trick" (I am not saying that it would be easy... I never seriously thought about doing it).

[EDIT following the comments to address the question of the best way to formulate Non-linear factors] The linear formulation of factor has the advantage of defining simultaneously the factors and the loadings (here I take a case with more than one factor to express the problem a more generic way, and I explicitly figure the time $t$): $$R_i(t)=\sum_{i,j}X_i f_j(t) +\epsilon_i(t).$$

It is a well-posed problem in the sense that under quite generic assumptions you can

- use a PCA on returns of a lot of stocks $(R_1,\ldots,R_N)$ to find $K$ factors (see for instance Factors That Fit the Time Series and Cross-Section of Stock Returns by Lettau and Pelger)

- use a linear regression in a second stage to find the loadings.

If you want to replace the factors with a non linear counterpart, my advice is

- work on the residuals of the linear factors, ie on $\epsilon$, because at least it will guarantee you that you are doing something beyond the natural linear factors

- either you use any non-linear PCA like Self Organizing Maps, ie you are in a non supervised mode, and in such a case you have no inputs and the data are again $(R_1,\ldots,R_N)$. In such a case your loss function should be something like explained variance of the cross section of returns like for a PCA

- either you use a *supervised algorithm (the most famous being neural nets, like perceptrons) and you need exogenous inputs, like characteristics of the stock. In such a case your loss function will be the variance of the residuals of the explained returns.

It is interesting to notice

- that 1 and 2 respectively correspond to PCA vs characteristics of factors in the linear case (ie there is already 2 approaches that corresponds to 2 different philosophies)

- to really exploit the non linearity of the approach you may decide to go beyond L2 criteria. For instance for case 1 you can try to explain not only the variance of the cross section but its potential skewness or any other non L2 property.

Shown in full with attribution under the source's licence. Licence: CC BY-SA 4.0 (Stack Exchange)

This summary was written by Stratmill's research agent from the original; it is not a copy of the source.