Skip to content
All library documents

Selecting Predictors and Modeling Nonlinear Trading PnL

Article Quant Q&A · Author: silencer

Summary

The discussion frames the task as explaining changes in a strategy’s daily profit and loss using candidate variables that were not used to generate its trading signals. It recommends first considering how the predictors relate to one another: dimensionality reduction can help when they are strongly correlated, while sparse or stepwise selection may be more informative when they are not. These approaches can help narrow the set of variables, but linear selection may miss nonlinear effects.

For nonlinear modeling, the response suggests constraining the range of candidate relationships with prior knowledge, such as assuming effects are additive or smooth. It lists several model families, including neural networks, tree-based methods, support vector regression, and spline methods; some can combine variable selection with nonlinear fitting. The discussion offers methodological guidance rather than an empirical demonstration: it provides no dataset, comparison, or validation results. Any apparent explanation of the equity curve may reflect overfitting or patterns without a meaningful economic basis, so interpretation depends on careful validation and informed restrictions on the model.

Key ideas

  • Model changes in daily profit and loss as a function of candidate explanatory variables.
  • Use dimensionality reduction when predictors are highly correlated, and consider sparse or stepwise selection when they are not.
  • Linear selection methods may overlook predictors whose effects are nonlinear.
  • Restricting candidate relationships using assumptions such as smoothness or additivity can make nonlinear fitting more manageable.
  • Nonlinear model families differ, and some can select variables during fitting.
  • Treat fitted explanations cautiously because flexible models can overfit or find spurious relationships.

Tags

Full text
# How to better understand trading signals?


# How to better understand trading signals?












I am looking to get a better understanding of an output from a trading strategy. Basically I have a daily equity curve lets call it $Y_t$. I have defined a bunch of independent variables $X_{it}$ that I think can explain the movement in the daily PnL. The independent variables are not used in trading signal generation directly.

1) How can I go about deducing which independent variables explain my $Y_t$ , assuming that the relationship could be non-linear? I can start out with PCA but from what I understand it assumes a linear relationship.

2) Using a reduced independent variable set $X_{it}$ from 1) how do I go about defining a non-linear relationship with $Y_t$ . Neural Nets maybe?

I understand that doing 1) and 2) might result in overfitting but I just want to understand the equity curve better.

## Answer by MichaelJ (score 3)

https://quant.stackexchange.com/a/12843

This is an pretty general question

You are essentially asking how to estimate the regression function

$$Y[t]-Y[t-1] = m(X_i [t], ..., X_p[t]) + \epsilon[t]$$

without any additional structure. Here is a basic list of questions to consider. Keep in mind that the more you can guide these procedures with domain knowledge the better your results will likely be.

- A couple of your questions suggest that you suspect many of your $X_i$ are irrelevant to $Y$. You suggest PCA. I think PCA is most useful when the $X_i$ are extremely correlated. If your $X_i$ are not correlated then selection techniques such as LASSO and stepwise regression may provide more insight.

- There are many methods for fitting non-linear forms for $m( )$. Before you launch into these it is useful if you can further restrict the function space. Remember, the set of all non-linear functions in $p$ variables can be vast and many of these functions will fit the data despite being nonsense. Common restrictions include additivity and smoothness.

- Some suggestions for fitting $m( )$ include: neural nets, regression trees, boosting, random forests, support vector regression, and multivariate adaptive regression splines. Some of these methods can automatically select variables as part of the fitting process. This avoids the problem with step (1) you point out, namely that those methods are linear and ignore potentially important non-linear relationships.

Shown in full with attribution under the source's licence. Licence: CC BY-SA 4.0 (Stack Exchange)

This summary was written by Stratmill's research agent from the original; it is not a copy of the source.