Skip to content
All library documents

Testing Whether an Out-of-Time Strategy Beats Its Benchmark

Article Quant Q&A · Author: BGa

Summary

The document presents a validation question for a machine-learning portfolio strategy evaluated out of time against a strong benchmark. Its custom rolling metric reflects portfolio value and additional state variables such as inventory and time. The observed performance gap appears to widen steadily, but the author wants a statistical test to assess whether the strategy consistently outperforms.

The metric’s distribution is unknown, and increments in the two performance series appear correlated. The document raises a t-test as a possibility but provides no answer, proposed test, assumptions, or empirical results. The paired and time-dependent nature of the increments is therefore an important unresolved issue: a simple test may not be appropriate without accounting for serial dependence and how the metric is constructed. The material identifies the inference problem rather than offering a validated procedure.

Key ideas

  • The evaluation compares an out-of-time machine-learning strategy with a strong benchmark.
  • The custom performance metric includes portfolio value and other variables such as inventory and time.
  • The apparent widening performance gap motivates a formal test of outperformance.
  • The metric’s distribution is unknown, and strategy and benchmark increments may be correlated.
  • The document poses the testing question but does not identify or validate a specific statistical test.

Tags

Full text
# Statistical testing of out-of-time portfolio performance (measured via a custom metric)


# Statistical testing of out-of-time portfolio performance (measured via a custom metric)












I'm testing (out-of-time) my machine learning (ML) based strategy against a strong benchmark. As a performance metric, I'm using a custom rolling metric $M(t)$ which takes into account the portfolio value at time $t$ as well as other variables (inventory, time, etc). Visually, it is clear that the ML strategy is beating the benchmark since the performance difference is increasing in time quite steadily. However, I would like to back my findings up with an appropriate statistical test. I know nothing about the distribution of this metric, but the $M(t)$ increments (i.e. a differenced time series) corresponding to my ML strategy seem to be correlated with the $M(t)$ increments corresponding to the benchmark. In essence, I would like to prove that my strategy is consistently beating the benchmark, so I would like to find an appropriate statistical test. Perhaps something as simple as a t-test would do it?

Shown in full with attribution under the source's licence. Licence: CC BY-SA 4.0 (Stack Exchange)

This summary was written by Stratmill's research agent from the original; it is not a copy of the source.