Building a Reproducible SEC Filings Pipeline from Public APIs
Summary
The document outlines a small public-data pipeline that maps a stock ticker to a Central Index Key, retrieves the company's recent SEC submissions, and filters for 10-K, 10-Q, and 8-K filings. It explains that the submissions endpoint stores recent filing details in parallel column arrays, so each selected field must be read using the same index. The example returns filing dates, accession identifiers, and primary document names, with a sample of the first ten matching records.
Three quality controls address ticker-to-identifier ambiguity, consistent indexing across arrays, and responsible API access through identification, throttling, caching, retries, and failure records. It also explains limits of company-level XBRL facts: custom tags may be absent, and units, periods, dimensions, restatements, and company definitions can prevent direct comparison. The document is a data retrieval and validation guide, not an investment strategy; it provides code but no trading signals, backtest, or evidence that filings data alone predicts returns. Its stated access guidance and endpoint behavior may change and should be checked against current SEC documentation.
Key ideas
- Map tickers to CIKs carefully because the association can be incomplete or ambiguous.
- Filter filing types while preserving matching indices across the submissions column arrays.
- Use a declared client identity and conservative access controls, caching, and failure handling.
- Company Facts may omit custom tags and does not guarantee mechanical comparability across observations.
- The pipeline retrieves public filing metadata and does not generate trading signals or evaluate a strategy.
Tags
This summary was written by Stratmill's research agent from the original; it is not a copy of the source.