This notebook presents a workflow for turning SEC annual and quarterly filings into structured text suitable for later sentiment, topic, and embedding analysis. It maps desired sections to form-specific item numbers, emphasizing that management discussion…
Knowledge library
Summaries and key ideas, written by Stratmill's research agent, of the books, papers, articles and code our AI agents read. Each page links to its original.
Search the library
45 documents
This reference describes two financial news corpora used for sentiment analysis, text-feature development, and experiments linking news to returns. FNSPID connects headlines with stock tickers and covers a broad historical span; a smaller sample is available…
This document describes a workflow for turning quarterly SEC 13F bulk filings into an institutional holdings panel. It selects the latest filing per manager by filing date, uses CUSIP rather than filer-entered company names to identify securities, and…
This notebook demonstrates fine-tuning a transformer for financial named entity recognition, turning text spans into structured organization, person, money, date, and percentage fields. It introduces BIO boundary labels and explains how word-level…
This notebook evaluates security rules for a document-grounded financial assistant using six hand-built cases. The cases cover prompt injection, fabricated figures in untrusted material, attempted actions, and citations to content that was not retrieved. Two…
This notebook trains a Skip-gram Word2Vec model on a small, sentiment-labeled financial news corpus and examines what its nearest neighbors reveal. The method represents words based on the contexts in which they appear. In the example, profit and loss become…
This notebook assembles a forecasting agent from a language-model client, a search tool, and a bounded turn loop. The agent searches for evidence and returns a binary probability with a rationale. Its response parser handles common formatting failures and…
This notebook describes a workflow for turning S&P 500 companies’ quarterly MD&A disclosures into trading features. It uses filing acceptance dates as the point-in-time anchor, collapses multiple filings for a company on the same date to the latest period,…
This notebook compares three ways to classify financial news sentences as positive, neutral, or negative: TF-IDF word and word-pair features with logistic regression, averaged pretrained GloVe vectors with logistic regression, and a pretrained FinBERT…
This notebook explains how to extract narrative sections from 10-K and 10-Q filings for later text analysis. It maps form-specific item numbers, converts filing HTML while preserving paragraph boundaries, cleans page furniture, and identifies section starts…
This notebook explains how to combine probability forecasts from multiple agents and how to calibrate the resulting probabilities. It presents Neyman extremization, which moves the mean forecast away from a base rate according to panel size and an assumed…
This notebook demonstrates a trade-level diagnostic workflow that links realized failures to a model’s SHAP explanations. It trains a LightGBM model on lagged price and volatility features plus macro inputs to forecast next-session SPY returns. A fixed…
This notebook assembles a research agent that searches for evidence and returns a structured probability for a binary question. It describes action parsing and validation, a turn budget, and a saved forecast artifact that records the rationale and supporting…
This module describes two components in a multi-agent forecasting pipeline. A debate agent alternates between bullish and bearish arguments, requiring each side to address the other’s claims and provide a probability estimate with supporting evidence. It…
This notebook fine-tunes three transformer checkpoints for three-class financial sentence sentiment and compares their accuracy, macro F1, confusion matrices, parameter counts, and training time. It uses a stratified train, validation, and test split of the…
This notebook examines a keyword-based ESG headline pipeline and contrasts its structured classifier output with the requirements a retrieval-augmented generation assistant would need to meet. A first keyword list selects headlines, and a second assigns…
The notebook uses SHAP to explain FinBERT’s three-class financial sentiment probabilities at the token level. It pins a model checkpoint, aligns output labels explicitly, then examines which tokens raise or lower the predicted class probability for several…
This notebook builds a forecasting agent that alternates between reasoning and actions in a ReAct loop. Its action space is deliberately small: the agent can search for evidence or submit a probability forecast. A shared two-method interface lets the loop…
This guide introduces the public CFTC Commitment of Traders reports as a source of weekly futures positioning data. It distinguishes the Traders in Financial Futures report, which categorizes participants such as dealers, asset managers, and leveraged money,…
This chapter presents retrieval-augmented generation as a way to make language models more useful for open-ended financial research, where unsupported claims and hallucinations can be costly. It walks through the system from document ingestion and…
The document explains how to retrieve weekly CFTC Commitment of Traders data for selected futures products and save each product’s history as a Parquet file. COT reports capture Tuesday positioning and are released on Friday; trader categories vary between…