Skip to content
All library documents

Evaluating Trading Algorithms with Task-Specific Metrics

Article Quant Q&A · Author: RoyalGoose

Summary

The document argues that no single absolute benchmark can establish the quality of every trading algorithm. The appropriate evaluation depends on the algorithm’s role: a signal model, portfolio-sizing process, execution method, or complete automated strategy. Sharpe ratio can help assess the returns of an integrated strategy, but its interpretation depends on the chosen reference return and the risks being measured.

It also suggests examining losing-period duration, the conditions associated with losses, and extreme downside outcomes, including through extreme-value analysis of backtest data. For alpha models, it mentions Giacomini–Rossi fluctuation tests as a way to study forecast performance when the better-performing model may change over time. These are suggestions rather than a worked evaluation or empirical comparison. The post gives no universal reference values and does not specify data, validation procedures, or implementation details; its central lesson is to match measures to the system’s objective and assess risk as well as returns.

Key ideas

  • Algorithm evaluation should reflect whether a component generates signals, sizes positions, executes trades, or runs a complete strategy.
  • Sharpe ratio evaluates returns relative to a selected benchmark and is not an absolute quality score.
  • Loss duration, loss conditions, and extreme downside outcomes can add information about strategy risk.
  • Extreme-value analysis of backtest results is proposed for examining unusually severe losses.
  • Giacomini–Rossi fluctuation tests can assess forecast performance under changing model rankings.

Tags

Full text
# The quality of an trading algorithm


# The quality of an trading algorithm












Good day. I am currently writing a term paper on the creation of trading algorithms in the foreign exchange market (by an algorithm I mean the one that follows the alpha model, for example, signals from some kind of technical analysis indicator). And I wondered, what indicator can be considered reliable when assessing the quality of an algorithm? Indicators such as Sharp ratio or LR correlation are, as I understand it, relative, i.e. the bigger, the better. But how can we give absolute analytics of the efficiency of a single algorithm? This leads to the question, are there any reference values for these indicators for comparison?

## Answer by Stéphane (score 1, accepted)

https://quant.stackexchange.com/a/59098

As regards a commentary made by another user, the Sharp Ratio is not a standalone measure because it requires the selection of a benchmark rate of return. While it is common practice in financial economics to use some proxy for the risk-free rate as a benchmark, it is worth nothing that financial economists are usually focusing on data as it relates to representative agents. Your average trader will probably find other rates than the yield on US Treasury bonds more meaningful.

Now, how do you evaluate a trading algorithm depends on what the algorithm is tasked to achieve. As @noob2 pointed out, algorithms are used to automate tasks and trading involves many tasks. It is not uncommon, from what I understand, for one or many algorithms to focus exclusively on identifying trading signals. Alpha models are usually designed to flag short, long or neutral trading opportunities -- again, that's my understanding. Those signals can then be fed into another algorithm used to determine a new target portfolio which poses the problem of sizing your target positions and designing an execution strategy which presumably would split trades into blocks to avoid moving prices in too much of an adverse direction. Each step in there can be automated with more or less sophisticated methods and every solution can be more or less integrated. Your entire strategy consist in identifying desirable changes to your portfolio and chasing after this moving target, hopefully turning a profit large enough to justify operational costs and to compensate the risks involved.

Now, how do you evaluate it? Well, in the end, if you automated the whole thing, you could use a Sharp ratio to evaluate the WHOLE strategy. You'd also potentially look at the length of time you can spend loosing money, under which condition you seem to be loosing money, and how much money you can expect to loose. For the later, I would personally use my back testing data as input for an EVT kind of analysis to see what happens with the extreme lows. Basically, you need an idea of whether the approach works well enough to be profitable, what kind of bank roll you need to ensure any Law of Large Numbers effects kick in, and I'd also look into the Giacomini-Rossi fluctuation tests for my alpha models to make sure things aren't bouncing around too much -- they check for forecasting performance under model instability (i.e., when the best model changes over time).

I'm sure you can easily find meaningful metrics and methods to do all of the above very easily.

Shown in full with attribution under the source's licence. Licence: CC BY-SA 4.0 (Stack Exchange)

This summary was written by Stratmill's research agent from the original; it is not a copy of the source.