Structural Limits of LLMs in Trading Research
Summary
The article argues that language models are unreliable for discovering trading edges because of three problems: conventional trading advice dominates their training data, models struggle to retrieve the latest value after repeated updates, and their open-ended answers converge on familiar patterns. It describes research reporting proactive interference in sequential-value tasks across tested models and high similarity among responses from the same and different models. The author connects these findings to dynamic market conditions and says retrieval-augmented generation did not prevent generic advice from appearing in his own system.
The proposed workflow is to have people form and assess an economic explanation for an edge, while using AI for implementation tasks such as coding and data handling. The article cautions that the cited studies concern general model behavior, not direct tests of trading performance, and acknowledges that future architectural changes could alter the picture. Its central practical claim is that attractive backtests and fluent model output do not establish a persistent edge.
Key ideas
- Language models may reproduce common trading advice because it is prevalent in their training data.
- Research described in the article finds that models can confuse recent values with earlier updates as context accumulates.
- The cited study on open-ended responses reports substantial convergence within and across models.
- The author argues that these tendencies may suppress unusual, mechanism-based hypotheses about trading edges.
- The suggested division of labor is for humans to develop and evaluate edge theories while AI helps implement them.
Tags
This summary was written by Stratmill's research agent from the original; it is not a copy of the source.