Risks of LLM-Generated Code in Quantitative Research
Summary
The document discusses using language models to assist with data cleaning and backtesting, focusing on the difficulty of verifying research logic. Its example is a rolling signal intended to use observations through the prior period; generated code accidentally included the current period, creating look-ahead bias despite running successfully. The author argues that plausible output and a successful backtest do not establish that code is aligned correctly or free of leakage.
The accepted answer takes a strongly skeptical position. It cites concerns including hallucinations, inefficient suggestions, privacy, licensing, and regulatory compliance, and argues that data preparation and research quality require substantial human expertise. These points are presented as opinion and supported mainly by external examples and commentary rather than controlled evidence. The document offers no concrete workflow for validating generated research code, so its clearest practical lesson is to independently inspect data timing and logic before trusting results.
Key ideas
- Code that runs successfully can still introduce look-ahead bias through incorrect time alignment.
- Backtest performance alone cannot reveal whether the research process contains leakage or other logic errors.
- The answer raises privacy, licensing, accuracy, and compliance concerns about public language models.
- The document argues for expert oversight but does not lay out a detailed validation workflow.
Tags
Full text
# Best practices for using LLM coding assistants safely in quant research # Best practices for using LLM coding assistants safely in quant research I am relatively new to quantitative trading and I have been using AI coding assistants (Codex) to help me write Python/Pandas code for data cleaning and backtesting. While the AI is great for syntax, I am finding it dangerous for logic. It can introduces subtle stochastic mistakes even if I explicitly instructed otherwise. Just to name one, in a recent backtesting I explicitly asked it to generate signals depending on a rolling window from $t-N$ to $t-1$. The generated code looked perfect and ran without errors, but I later discovered it had subtly included index $t$ (a look-ahead bias). In 2026, my experience is that AI-assisted coding is fairly reliable in tasks where the output is directly verifiable, like building a scraper. Quant research is different: you often cannot tell from a backtest result alone whether the strategy is good or whether there is hidden leakage or misalignment. That is why AI-assisted coding in quant industry is very different from that in other areas. I would still like to use LLMs for repetitive programming work so I can spend more time on alpha research, but I want a workflow that helps ensure the AI-generated parts are trustworthy. ## Answer by AKdemy (score 5, accepted) https://quant.stackexchange.com/a/85459 I'd argue you can never safely use LLMs. See What are some factually incorrect quantitative finance answers generated by AI? for a community wiki on this topic. More generally, Nick Patterson gives a good overview about what they do at Rentec (the whole podcast starts at 16:40, Rentec starts at 29:55 - a sentence before that is helpful). He states that you need the smartest people to do the simple things right, that's why they employ several PHDs to just clean data. The use of public LLMs is outright banned at many companies (see https://www.techzine.eu/news/applications/103629/several-companies-forbid-employees-to-use-chatgpt/), for various reasons including - data security / privacy issues - (new) employees using poor quality responses - hallucinations - inefficient code suggestions - copyright and licensing issues - lack of regulatory standards - potential non compliance with data laws like GDPR ... It's a great tool for simple school stuff, but it's mostly inefficient when it comes to actual work. That's why all use of generative AI (e.g., ChatGPT and other LLMs) is banned on Stack Overflow, see https://meta.stackoverflow.com/q/421831 which states: > Overall, because the average rate of getting correct answers from ChatGPT and other generative AI technologies is too low, the posting of content created by ChatGPT and other generative AI technologies is substantially harmful to the site and to users who are asking questions and looking for correct answers. Below is what ChatGPT "thinks" of itself here. A few lines: - I can't experience things like being "wrong" or "right." - I don't truly understand the context or meaning of the information I provide. My responses are based on patterns in the data, which may lead to incorrect or nonsensical answers if the context is ambiguous or complex. - Although I can generate text, my responses are limited to patterns and data seen during training. I cannot provide genuinely creative or novel insights. - Remember that I'm a tool designed to assist and provide information to the best of my abilities based on the data I was trained on. For critical decisions or sensitive topics, it's always best to consult with qualified human experts. The only large company I know of who was initially very keen on using these models is Citadel, but they also largely changed their mind by now, see https://fortune.com/2024/07/02/ken-griffin-citadel-generative-ai-hype-openai-mira-murati-nvidia-jobs/. https://www.bloomberg.com/news/articles/2025-10-15/ken-griffin-says-genai-fails-to-help-hedge-funds-produce-alpha Same for coding. Initially, Devin AI was hyped a lot, but it's essentially a failure, see https://futurism.com/first-ai-software-engineer-devin-bungling-tasks It's bad at reusing and modifying existing code, https://stackoverflow.blog/2024/03/22/is-ai-making-your-code-worse/ Causing downtime and security issues, https://www.techrepublic.com/article/ai-generated-code-outages/, or https://arxiv.org/abs/2211.03622 Even if the answers were always correct and public LLMs worked flawlessly, you could take the argument further and say that you cannot rely on tools and data everyone else has access to if you’re trying to do research or find something that generates profit. That's what Graham Giller refers to in https://www.youtube.com/watch?v=qUmRQCC61Vw&t=623s. More fundamentally, though, computers cannot even drive cars properly yet. That’s something most grown-ups can do. Yet the number of people working successfully as quants, traders, and developers is significantly lower.
Shown in full with attribution under the source's licence. Licence: CC BY-SA 4.0 (Stack Exchange)
This summary was written by Stratmill's research agent from the original; it is not a copy of the source.