SVM Price Direction Models and the Risks of In-Sample Accuracy
Summary
This article outlines a three-class SVM model that labels future price movement as up, down, or sideways. It describes candidate inputs from recent price changes, volatility, RSI, time of day, trade flow, and order-book depth, and discusses feature relevance, redundancy, and stability. It also proposes examining feature importance, mutual information, and a confusion matrix to understand model behavior.
The reported results illustrate why accuracy can mislead: sideways cases dominate one example, while a later 100% result comes from a small, highly related sample that was also used for evaluation. The article recognizes this as overfitting and cautions against treating either result as a live expectation. It presents order-flow and book-imbalance features as a planned expansion, but the code and examples do not establish out-of-sample profitability. Proper time-separated evaluation and realistic trading costs would be needed to assess whether the approach generalizes.
Key ideas
- The model classifies future movement into upward, downward, and sideways outcomes using engineered market features.
- Candidate features include price and volatility measures, trade imbalance, order-book pressure, and time variables.
- Overall accuracy can be inflated when sideways labels dominate the sample.
- Evaluating a model on its training data can produce perfect-looking results through overfitting.
- Feature importance and confusion matrices help diagnose model behavior but do not prove live profitability.
Tags
This summary was written by Stratmill's research agent from the original; it is not a copy of the source.