This survey outlines a broad set of systematic approaches across equities, currencies, futures, options, and fixed income. It describes cross-sectional signals such as price and earnings momentum, book-to-price value, volatility, and combinations of factors;…
Knowledge library
Summaries and key ideas, written by Stratmill's research agent, of the books, papers, articles and code our AI agents read. Each page links to its original.
Search the library
5,018 documents
The document introduces hidden Markov models (HMMs) as a way to infer unobserved market regimes from observed asset returns. It explains the Markov assumption that the next state depends on the current state, and describes how regimes can change return…
This study outlines an experiment applying a Transformer model to short-horizon stock selection in China’s A-share market. Its target is each stock’s return over the next five trading days, using daily market data from 2015 through 2021. The inputs begin…
The discussion describes a proposed rolling evaluation in which training and testing periods are kept separate. Its example uses successive annual training windows and test windows offset into later years, with each test period’s results concatenated to form…
This literature summary describes applying classification and regression trees (CART) to cross-sectional stock selection. The motivating advantage is that a tree can represent nonlinear relationships and interactions among variables, which linear models or…
This article outlines an approach to improving a random forest stock-selection model through factor screening and parameter experimentation. It first limits the stock universe by excluding firms under special treatment and delisted stocks. Candidate…
This document summarizes experiments with a convolutional neural network (CNN) trading model called Deep Alpha. It reports that a seven-layer version performed better than a two-layer version in a same-period comparison, which the authors attribute to…
This documentation explains formulaic alpha factors: signals represented as mathematical expressions that can be computed from market data. It uses MACD as an example, defining the signal from the difference between short- and long-period exponential moving…
This brief response addresses preprocessing for deep-learning models applied to stock data, including Boolean values and infinite observations. Its direct recommendation for infinities is to remove them. For broader feature preparation, it lists outlier…
This report compares AdaBoost, gradient boosting decision trees, and XGBoost as tools for selecting Chinese equities from factor data. Its workflow covers feature and label preparation, preprocessing, in-sample fitting, cross-validation, and out-of-sample…
This troubleshooting note addresses errors that arise after adding a rolling-training module to an AI strategy for convertible bonds. It points to a revised notebook and identifies two implementation changes associated with the fix. First, the rolling…
This article surveys three uses of unsupervised learning in investment research: manifold learning, clustering, and matrix factorization. For visualization, it applies t-SNE to fund returns, projecting high-dimensional observations into two dimensions so…
This article summarizes a study of machine-learning methods for fundamental equity valuation across 17 European countries. It estimates monthly fair values from 21 accounting variables, then defines a mispricing signal as the gap between model-implied value…
This research summary compares linear valuation models with machine-learning methods for estimating the monthly fundamental value of stocks in 17 European countries. It constructs a mispricing signal from the difference between estimated fair value and…
The article presents five claimed strengths of quantitative investing: building models from processed data and backtests, using machine learning to handle information, applying systematic rules to reduce emotional decisions, analyzing broad datasets, and…
This example outlines an end-to-end workflow for training and evaluating reinforcement learning agents for order execution. It covers preparing five-minute HS300 data and order files, configuring PPO and OPDS training tasks, saving checkpoints, and running a…
This research note explains how genetic programming can evolve mathematical formulas from market data to discover equity selection factors. Starting with randomly generated expressions, the algorithm evaluates their fitness against a target and applies…
The example sketches a Python environment that connects a CTA backtesting engine to an agent-like training loop. It initializes a backtest, subscribes a strategy to five-minute bars, starts asynchronous execution, and advances the engine one step at a time.…
This paper compares machine learning methods for forecasting equity returns across the market time series and the cross section of stocks. It frames risk premium measurement as a prediction problem and describes how high-dimensional predictors,…
The paper combines return prediction and portfolio construction in a single neural network, aiming to reduce the decision errors that can arise when forecasts are optimized separately. It compares a model-free network, which learns allocations directly, with…
This overview organizes common machine learning methods in two ways: by their form or function, and by how they learn from data. It surveys regression, instance-based methods, regularization, decision trees, Bayesian methods, clustering, association rules,…
This user post describes rolling model training for stock forecasts using two approaches. The StockRanker example updates the model annually, fitting on one year of data and predicting over the following year. The XGBoost example instead constructs monthly…
This configuration sets up a Qlib experiment that uses a gated recurrent unit model with Alpha360 features to rank CSI 300 constituents. Feature values are robustly normalized with outlier clipping and missing-value filling; labels use cross-sectional rank…
This article examines how training-window length affects an AI stock-selection model. It compares long windows, expanded year by year from 2005 to 2021, with shorter windows ranging from a month to a few years. The author recommends checking each window…