금융 리서치를 위한 감사 가능한 검색 증강 생성 구축
기사 Machine Learning for Trading
요약
이 장에서는 근거 없는 주장과 환각의 비용이 클 수 있는 개방형 금융 리서치에서 언어 모델의 유용성을 높이는 방법으로 검색 증강 생성을 소개합니다. 문서 수집과 구조를 고려한 분할부터 메타데이터, 도메인별 임베딩, 어휘 및 의미 기반 하이브리드 검색, 재순위화, 증거에 근거한 생성까지 시스템을 설명합니다. 추적성과 수치 신뢰성을 높이기 위해 인용 제약과 도구로 검증된 계산도 포함합니다.
이 장은 평가를 진단 과정으로 다룹니다. 팀은 검색 실패를 맥락, 종합, 계산, 답변 보류 오류와 구분한 뒤 각각을 측정하고 해결해야 합니다. 예시에는 금융 공시, ESG 분석, 기관 보유 내역 리서치가 포함됩니다. 반복 가능한 레이블 생성 작업에는 미세 조정이 더 적합하고, 변경되는 문서에서 근거를 찾아야 하는 질문에는 RAG이 적합하다고 설명합니다. 일부 임베딩 비교는 합성 데이터에 기반하며, 이 개요는 엔지니어링 방법을 설명할 뿐 RAG만으로 정확하거나 거래 가능한 결론을 보장한다고 입증하지는 않습니다.
핵심 아이디어
- 근거 없는 생성은 환각을 일으킬 수 있으므로 금융 관련 언어 모델의 답변을 검색된 증거에 기반하세요.
- 맥락과 인용 추적성을 보존하도록 문서 구조와 메타데이터를 유지해 파싱하세요.
- 금융 질의에는 의미 기반 검색, 어휘 검색, 메타데이터 필터, 재순위화를 결합하세요.
- 시스템 실패를 찾을 수 있도록 검색, 종합, 계산, 답변 보류를 따로 평가하세요.
- 변경되는 문서의 증거 기반 추론에는 RAG을, 반복 가능한 예측 작업에는 미세 조정을 사용하세요.
태그
전문
# Chapter 22: RAG for Financial Research # Chapter 22: RAG for Financial Research The chapter explains why text classification is not enough once the practitioner's task becomes open-ended analysis rather than fixed-label prediction. It positions LLMs as a shift from extracting features to answering analyst-style questions, but immediately frames hallucination as the central obstacle in finance. Readers should care because it sets up the chapter's core claim: generative AI only becomes usable in high-stakes financial settings when it is grounded in verifiable evidence rather than trusted as an oracle. ## Learning Objectives - Explain why hallucination makes ungrounded LLM use unacceptable in finance and why retrieval-augmented generation is the core architectural response - Design a financial RAG pipeline from document ingestion through retrieval and grounded generation, including structure-aware parsing, chunking, metadata, embeddings, and citation support - Compare generic and domain-specific embedding models and evaluate retrieval quality on a target corpus using practical retrieval metrics and latency trade-offs - Build a retrieval stack that combines semantic search, lexical search, metadata filtering, and re-ranking to improve precision and recall on financial documents - Use constraint-based prompting, citation checks, and tool-verified computation to make generated answers more faithful, auditable, and numerically reliable - Diagnose RAG failures by separating retrieval, context, synthesis, computation, and abstention errors, and apply targeted evaluation methods to improve each component - Distinguish when to use RAG versus fine-tuning for financial applications, and explain how RAG functions as one tool within broader agentic workflows ## Sections ### 22.1 From Feature Extraction to Generation This section explains why text classification is not enough once the practitioner's task becomes open-ended analysis rather than fixed-label prediction. It positions LLMs as a shift from extracting features to answering analyst-style questions, but immediately frames hallucination as the central obstacle in finance. Readers should care because it sets up the chapter's core claim: generative AI only becomes usable in high-stakes financial settings when it is grounded in verifiable evidence rather than trusted as an oracle. ### 22.2 Grounding LLMs with Retrieval-Augmented Generation This section introduces RAG as the architectural answer to hallucination and lays out the index, retrieve, generate pipeline in clear engineering terms. It also distinguishes the appealing simplicity of the baseline design from the much harder production reality in financial documents, where naive pipelines fail quickly. Readers should care because this is the conceptual backbone of the chapter: the model is valuable not because it "knows," but because it can synthesize over retrieved evidence. ### 22.3 Intelligent Document Ingestion This section argues that RAG quality starts before retrieval, with the way filings and related documents are parsed, chunked, and annotated. It shows why fixed-size chunking breaks tables, headers, temporal context, and citation traceability, then motivates structure-aware parsing, multimodal handling, and rich metadata as non-negotiable for financial corpora. Readers should care because poor ingestion silently corrupts everything downstream: if the semantic units are wrong, neither embeddings nor prompting can recover what was lost. - [`01_sec_filing_pipeline`](01_sec_filing_pipeline.ipynb) — - Compare two practical approaches to SEC filing ingestion for downstream RAG systems. - Add the metadata needed for point-in-time filtering and citation traceability. ### 22.4 Domain-Specific Embeddings This section explains why generic embeddings are often inadequate for financial retrieval and why domain-adapted models matter for jargon, entities, regulation, and quantitative concepts. It also adds practical engineering considerations such as benchmark use, corpus-specific evaluation, dimensionality, and storage trade-offs. Readers should care because retrieval quality is not a cosmetic optimization; it determines whether the system can even surface the evidence needed for a trustworthy answer. - [`02_domain_embeddings_comparison`](02_domain_embeddings_comparison.ipynb) — This notebook compares embedding models for financial document retrieval: Uses synthetic data. ### 22.5 Hybrid Retrieval and Vector Databases This section shows that semantic search alone is not enough for finance, where exact terms such as tickers, filing types, and codes matter. By combining vector search with lexical ranking and metadata filtering, the chapter presents hybrid retrieval as the practical default rather than an advanced extra. Readers should care because this is where the system becomes robust to the actual query mix analysts use, instead of only performing well on clean semantic paraphrases. - [`03_hybrid_retrieval`](03_hybrid_retrieval.ipynb) — This notebook demonstrates hybrid retrieval combining: Uses cross_encoder data. ### 22.6 Re-ranking and Constraint-Based Prompting This section moves from finding candidate evidence to making grounded answers more precise and defensible. It combines cross-encoder re-ranking, context-window discipline, citation-constrained prompting, and tool-verified numeric computation to turn the LLM into a synthesis engine rather than a free-form generator. Readers should care because this is where the chapter makes financial RAG operationally credible: answers must not only sound right, but be supported, calculationally reliable, and audit-friendly. - [`05_10k_rag_assistant`](05_10k_rag_assistant.ipynb) — This notebook demonstrates a production-grade RAG (Retrieval-Augmented Generation) system for analyzing SEC 10-K filings. Key features: Uses data, documents, parquet_documents data. ### 22.7 Diagnosing RAG Pipeline Bottlenecks This section treats RAG as a system that must be measured and debugged, not admired through demos. By separating retrieval, context, synthesis, computation, and abstention failures, it gives readers a practical framework for evaluation and iteration, including RAGAs, claim-level checks, and production observability. Readers should care because without this diagnostic lens, teams cannot tell whether a bad answer comes from retrieval, ranking, prompting, or arithmetic, and therefore cannot improve the system systematically. - [`04_ragas_evaluation`](04_ragas_evaluation.ipynb) — This notebook implements a finance-oriented evaluation harness that... - [`08_rag_security`](08_rag_security.ipynb) — This notebook demonstrates attack and defense evaluation for document-grounded finance assistants, a critical concern when RAG systems operate on untrusted or adversarial document corpora. ### 22.8 Applications and Strategic Choices This section anchors the architecture in concrete financial use cases, especially a 10-K due diligence assistant and an ESG analysis comparison. It also gives the clearest strategic boundary in the chapter: fine-tuning is for repeatable label-producing skills, while RAG is for evidence-grounded reasoning over changing documents. Readers should care because this section translates technical design choices into organizational decisions about what kind of AI workflow they are actually building. - [`05_10k_rag_assistant`](05_10k_rag_assistant.ipynb) — This notebook demonstrates a production-grade RAG (Retrieval-Augmented Generation) system for analyzing SEC 10-K filings. Key features: Uses data, documents, parquet_documents data. - [`06_esg_rag_vs_finetune`](06_esg_rag_vs_finetune.ipynb) — This notebook compares two approaches to ESG (Environmental, Social, Governance) analysis: Uses finbert_pipeline data. - [`07_institutional_holdings_graph`](07_institutional_holdings_graph.ipynb) — Build a bipartite institution-stock graph from the 13F holdings artifact and derive co-ownership similarity, institutional momentum, and crowding signals for alpha research. ### 22.9 Introducing Agentic Frameworks This section positions RAG not as the endpoint, but as one tool inside broader multi-step agent workflows. It introduces the controller, tool, and memory pattern, then shows how grounded document retrieval fits into a larger architecture that may also use code, APIs, and databases. Readers should care because it opens the path from cited question-answering to goal-directed analytical workflows while keeping grounding as a core control mechanism. ## Running the Notebooks ```bash # From the repository root uv run python 22_rag_financial_research/<notebook>.py # Test mode (reduced data via Papermill) uv run pytest tests/test_chapter_notebooks.py -v -k "22_rag_financial_research" ``` ## References - **Aditya Kusupati et al.** (2024). [Matryoshka Representation Learning](https://doi.org/10.48550/arXiv.2205.13147). - **Alejandro Lopez-Lira** (2023). [Risk Factors That Matter: Textual Analysis of Risk Disclosures for the Cross-Section of Returns](https://doi.org/10.2139/ssrn.3313663). - **Alejandro Lopez-Lira and Yuehua Tang** (2025). [Can ChatGPT Forecast Stock Price Movements? Return Predictability and Large Language Models](https://doi.org/10.48550/arXiv.2304.07619). - **Alejandro Lopez-Lira et al.** (2025). [The Memorization Problem: Can We Trust LLMs' Economic Forecasts?](https://doi.org/10.2139/ssrn.5217505). - **Allen Huang et al.** (2020). [FinBERT—A Deep Learning Approach to Extracting Textual Information](https://doi.org/10.2139/ssrn.3910214). *SSRN Electronic Journal*. - **Ashish Vaswani et al.** (2017). [Attention Is All You Need](http://arxiv.org/abs/1706.03762). *arXiv:1706.03762 [cs]*. - **Chanyeol Choi et al.** (2025). [FinDER: Financial Dataset for Question Answering and Evaluating Retrieval-Augmented Generation](https://doi.org/10.48550/arXiv.2504.15800). - **Dongyu Ru et al.** (2024). [RAGChecker: A Fine-grained Framework for Diagnosing Retrieval-Augmented Generation](https://doi.org/10.48550/arXiv.2408.08067). - **Gordon V. Cormack et al.** (2009). [Reciprocal rank fusion outperforms condorcet and individual rank learning methods](https://doi.org/10.1145/1571941.1572114). *Association for Computing Machinery*. - **Guido Baltussen et al.** (2025). [Natural Language Processing for Asset Managers: Turning Text into Alpha](https://doi.org/10.3905/jpm.2025.1.784). *The Journal of Portfolio Management*. - **Hoyoung Lee et al.** (2025). [Your AI, Not Your View: The Bias of LLMs in Investment Analysis](https://doi.org/10.48550/arXiv.2507.20957). - **Hugo Bowne-Anderson** (2025). [Stop Building AI Agents: Use Smarter LLM Workflows](https://decodingml.substack.com/p/stop-building-ai-agents). - **Jason Wei et al.** (2023). [Chain-of-Thought Prompting Elicits Reasoning in Large Language Models](https://doi.org/10.48550/arXiv.2201.11903). - **Lingyun Zhao et al.** (2020). [A BERT based Sentiment Analysis and Key Entity Detection Approach for Online Financial Texts](http://arxiv.org/abs/2001.05326). *arXiv:2001.05326 [cs]*. - **Luyu Gao et al.** (2022). [Precise Zero-Shot Dense Retrieval without Relevance Labels](https://doi.org/10.48550/arXiv.2212.10496). - **Manuel Faysse et al.** (2025). [ColPali: Efficient Document Retrieval with Vision Language Models](https://doi.org/10.48550/arXiv.2407.01449). - **Nelson F. Liu et al.** (2023). [Lost in the Middle: How Language Models Use Long Contexts](https://doi.org/10.48550/arXiv.2307.03172). - **Nils Reimers and Iryna Gurevych** (2019). [Sentence-BERT: Sentence Embeddings using Siamese BERT-Networks](http://arxiv.org/abs/1908.10084). *arXiv:1908.10084 [cs]*. - **Orion Weller et al.** (2025). [On the Theoretical Limitations of Embedding-Based Retrieval](https://doi.org/10.48550/arXiv.2508.21038). - **Patrick Lewis et al.** (2021). [Retrieval-Augmented Generation for Knowledge-Intensive NLP Tasks](https://doi.org/10.48550/arXiv.2005.11401). - **Preetha Saha et al.** (2025). [Large Language Model Agents for Investment Management: Foundations, Benchmarks, and Research Frontiers](https://doi.org/10.2139/ssrn.5447274). - **Qianqian Xie et al.** (2023). [Pixiu: A large language model, instruction data and evaluation benchmark for finance](https://proceedings.neurips.cc/paper_files/paper/2023/hash/6a386d703b50f1cf1f61ab02a15967bb-Abstract-Datasets_and_Benchmarks.html). *arXiv preprint arXiv:2306.05443*. - **Qianqian Xie et al.** (2024). [Finben: A holistic financial benchmark for large language models](https://proceedings.neurips.cc/paper_files/paper/2024/hash/adb1d9fa8be4576d28703b396b82ba1b-Abstract-Datasets_and_Benchmarks_Track.html). *Advances in Neural Information Processing Systems*. - **Rodrigo Nogueira and Kyunghyun Cho** (2020). [Passage Re-ranking with BERT](https://doi.org/10.48550/arXiv.1901.04085). - **Shahul Es et al.** (2025). [Ragas: Automated Evaluation of Retrieval Augmented Generation](https://doi.org/10.48550/arXiv.2309.15217). - **Shunyu Yao et al.** (2023). [ReAct: Synergizing Reasoning and Acting in Language Models](https://doi.org/10.48550/arXiv.2210.03629). - **Stephen Robertson and Hugo Zaragoza** (2009). [The Probabilistic Relevance Framework: BM25 and Beyond](https://doi.org/10.1561/1500000019). *Found. Trends Inf. Retr.*. - **Tim Loughran and Bill Mcdonald** (2011). [When Is a Liability Not a Liability? Textual Analysis, Dictionaries, and 10-Ks](https://doi.org/10.1111/j.1540-6261.2010.01625.x). *The Journal of Finance*. - **Xu Liu et al.** (2024). [Moirai-MoE: Empowering Time Series Foundation Models with Sparse Mixture of Experts](https://doi.org/10.48550/arXiv.2410.10469). - **Yangyang Yu et al.** (2024). [FinCon: A Synthesized LLM Multi-Agent System with Conceptual Verbal Reinforcement for Enhanced Financial Decision Making](https://doi.org/10.48550/arXiv.2407.06567). - **Yaxuan Kong et al.** (2024). [Large Language Models for Financial and Investment Management: Models, Opportunities, and Challenges](https://doi.org/10.3905/jpm.2024.1.646). *The Journal of Portfolio Management*. - **Yixuan Tang and Yi Yang** (2025). [FinMTEB: Finance Massive Text Embedding Benchmark](https://doi.org/10.48550/arXiv.2502.10990). - **Ziang Fang and Jason Moore** (2025). What AI Can (and Can't Yet) Do for Alpha.
출처의 라이선스에 따라 출처를 표시하고 전문을 공개합니다. 라이선스: MIT
이 요약은 원문을 바탕으로 Stratmill의 리서치 에이전트가 작성했으며, 원문을 복사한 것이 아닙니다.