Quant Strategy Q&A on Factor Analysis and ETF Machine Learning
Summary
This Chinese-language post collects questions about factor research and strategy implementation, but does not provide answers. Topics include adding data factors, scheduling entries and exits, evaluating and combining factors, handling collinearity, distinguishing temporary factor drawdowns from lasting failure, and adding profit-taking or stop-loss rules. The questions identify practical areas a quant researcher would need to address, but they do not teach procedures for doing so.
The post also describes a proposed ETF selection pipeline: use constituent stocks and their factors, train an LSTM to produce features, feed those features into XGBoost, then aggregate stock scores by ETF holdings and rebalance among the top-ranked funds. The author reports that early tests tracked the index without clear excess returns, while a score threshold appeared to improve short tests but left the portfolio uninvested much of the time. These are preliminary observations, not robust evidence. The post itself raises unresolved concerns about training history, data volume, model behavior, and evaluation, with no final analysis or validation.
Key ideas
- The post raises questions about factor evaluation, combination, collinearity, and interpreting factor drawdowns.
- It outlines an LSTM-to-XGBoost pipeline that scores ETF constituents and aggregates scores to rank ETFs.
- The author reports preliminary tests that mostly tracked the index and lacked demonstrated excess returns.
- The apparent benefit of a score threshold is based on short tests and came with frequent underinvestment.
- The post leaves its methodological and data sufficiency questions unanswered.
Tags
This summary was written by Stratmill's research agent from the original; it is not a copy of the source.