Skip to content
All library documents

Cross-Validating SVM Stock-Movement Classifiers

Article Quant Q&A · Author: Lee Schmidt

Summary

The document asks how to evaluate a binary support vector machine that predicts NASDAQ stock-price direction using V-fold cross-validation. The described procedure trains on all but one subset, predicts the held-out subset, and repeats until every subset has served as the validation set. It recommends averaging the prediction-error estimates across folds rather than selecting the minimum, with roughly equal-sized, randomly sampled, non-overlapping subsets.

The response says the cross-validation procedure is not specific to SVMs and that each fold need not contain a minimum proportion of both outcome classes if sampling is random. It offers no empirical comparison, dataset details, or discussion of time dependence. In financial time series, random folds can mix observations across time and may yield optimistic estimates when nearby observations are dependent; the guidance should therefore be applied with care when forecasting future market data.

Key ideas

  • Estimate out-of-sample prediction error on each held-out fold and average the fold estimates.
  • Use roughly equal-sized, randomly sampled folds that do not overlap.
  • The answer does not require a minimum class proportion in every fold under random sampling.
  • Random-fold validation may be unsuitable for time-dependent financial data if it leaks temporal information.

Tags

Full text
# How to properly cross-validate when optimizing SVM classification?


# How to properly cross-validate when optimizing SVM classification?












I'm using SVM binary classification to predict movement of NASDAQ stock prices. My question is regarding cross-validation. I will divide the training data into V subsets. Training will be performed on (V-1) subsets and then prediction on the V-th subset is recorded. This will be done V times.

(1) Is the best measure of accuracy the average of all V outcomes? Or perhaps the minimum?

(2) Should subsets be equal length? Random?

(2a) Can subsets overlap?

(3) Do I need to ensure that each subset has at least some minimum percentage of each class outcome (1 or 0) for the results to be valid?

## Answer by Ram Ahluwalia (score 1, accepted)

https://quant.stackexchange.com/a/4281

The cross-validation procedure does not turn on the choice of algorithm.

- Yes - calculate the prediction error of the fitted models when predicting the V'th part of the data. Combine the V estimates of prediction average using a simple average.

- Subsets should be randomly sampled (roughly equally sized). 2a. Subsets should not overlap.

- No. As long as the sampling is random you are OK.

Shown in full with attribution under the source's licence. Licence: CC BY-SA 4.0 (Stack Exchange)

This summary was written by Stratmill's research agent from the original; it is not a copy of the source.