Skip to content
All library documents

Selecting Inputs, States, and Validation for HMM Index Timing

Article SuperMind

Summary

The article discusses design choices for a hidden Markov model used to time exposure to a Chinese equity index. It proposes screening a broad set of candidate features, including technical indicators, with feature-selection or extraction methods such as PCA, ICA, or random forests. It also recommends testing alternative numbers of hidden states instead of assuming that states correspond directly to rising, falling, or sideways markets. The example applies the approach to the CSI 300 and uses an ETF position when the predicted next-day state is classified as favorable, otherwise holding cash.

A central validation point is to predict recursively one trading day at a time, avoiding the future-data leakage that can arise from passing future observations into a single prediction. Favorable states are identified from their training-period performance using a chosen criterion, and the author proposes comparing feature and state-count choices on training and validation data. These selection procedures can overfit if repeatedly optimized against performance. The article offers a research outline rather than detailed results, and notes that suitability for short-term, high-frequency, or individual-stock timing is uncertain.

Key ideas

  • Candidate HMM inputs can be screened for useful information and redundancy rather than chosen without justification.
  • The number of hidden states should be tested because states may capture patterns other than simple market direction.
  • Recursive next-day predictions help prevent future information from leaking into a backtest.
  • State labels for exposure decisions can be based on training-period performance, but that choice risks overfitting.

Tags

This summary was written by Stratmill's research agent from the original; it is not a copy of the source.