Skip to content
All library documents

Methods for Optimizing Predictions for Rank Correlation

Article Quant Q&A · Author: quant122

Summary

The document considers prediction of a continuous variable when correctly ordering observations matters more than matching their values in a way that maximizes Pearson correlation. One simple approach is to transform the target into ranks and fit an ordinary least squares model to those ranks. Another is to optimize a model’s parameters directly against Spearman rank correlation, though that requires choosing a model specification and solving an optimization problem.

The response also suggests comparing regression models with cumulative gains or lift charts, or predicting discrete target quantiles instead of raw values. For example, a classifier can estimate the probability of an observation being in an upward category, and those probabilities can then be used to rank predictions. These are alternative practical approaches rather than a controlled comparison: the document reports no empirical results and does not discuss optimization difficulties, tie handling, validation design, or how to choose among the methods for a particular dataset.

Key ideas

  • Regressing on ranked target values is a simple way to emphasize ordering rather than raw magnitudes.
  • Model parameters can be optimized directly for Spearman correlation, subject to the selected model specification.
  • Cumulative gains and lift charts can help compare candidate regression models by ranking performance.
  • Predicting target categories or quantiles can produce probability scores for ordering observations.
  • The document proposes methods but supplies no empirical comparison or selection rule.

Tags

Full text
# Rank Correlation Based Prediction


# Rank Correlation Based Prediction












Are there any methods of prediction (machine learning, regression, etc.) which are designed to maximize the rank correlation (spearman correlation, kendall's tau, etc.) of your prediction with your independent variable?

I am trying to predict a continuous numerical variable with a given set of inputs, and I am more concerned that the rank-correlation of predictions is high than the pearson correlation.

## Answer by Ram Ahluwalia (score 3, accepted)

https://quant.stackexchange.com/a/4489

The simplest way is to transform your dependent variable from returns into rank-space and then use ordinary least squares regression.

A more complex technique would involve setting up an optimization problem where you maximize the spearman correlation between your vector of predictions and actuals. More explicitly, the objective function in the optimizer is the spearman rank correlation function. The optimizer will search over the set of parameters to some model specification (for example, if you use a linear factor model it will identify a set of betas) that maximize the spearman correlation. Of course, there are many potential model specifications so this approach is a bit open ended.

Another simple approach would be to use a cumulative gains and lift chart to select amongst competing regression models.

Or you can use a discrete dependent variable to attempt to predict which quantiles an instrument is likely to fall within. A simple example would be to develop a logistic regression model where you predict 1 or 0 (let's say 1 means Up, and 0 means down). Then you can use the probability outputs generated when you score data using your logistic model to rank-order your predictions (vs. using regression outputs which may be more sensitive to outliers).

Shown in full with attribution under the source's licence. Licence: CC BY-SA 4.0 (Stack Exchange)

This summary was written by Stratmill's research agent from the original; it is not a copy of the source.