Finding SEC XBRL Filings with Quarterly Full Index Files
Summary
The document describes how to locate company XBRL filings in the SEC’s archive without searching company directories by trying to infer accession-number patterns. It points to quarterly full-index files, which list filing paths and can be downloaded in compressed form. The example traces a company’s annual filing in the first quarter of a year from its index entry to the corresponding archive text file.
The answers also mention alternative ways to obtain and parse filing data, including SEC XBRL endpoints and datasets, code libraries, and a real-time filing feed. These options differ in scope: a real-time feed may not provide historical records, while crawler packages may be outdated or require separate evaluation. The document offers discovery routes rather than a complete data pipeline, and users still need to select the filing type and parse the relevant XBRL facts for their research.
Key ideas
- Quarterly SEC full-index files can identify the archive path for XBRL filings.
- Compressed index files provide a practical way to retrieve filing listings in bulk.
- Accession numbers identify individual submissions, but searching index entries avoids guessing archive naming patterns.
- SEC endpoints, datasets, and third-party tools offer alternative retrieval and parsing routes.
- Real-time filing feeds may not cover historical filings.
Tags
Full text
# Automated 10-K XBRL data grab using the SEC file structure
# Automated 10-K XBRL data grab using the SEC file structure
I would like to write a program that takes as input a list of CIK/year/quarter entries. The program should iterate through the list and, for each entry, grab XBRL financial data from the SEC website for the given CIK/year/quarter combination.
I can decipher some parts of the SEC file structure, but not all. For example, post fixing `Archives/edgar/data/1288776/11/` to the SEC base address gives a directory listing of all filings for the year 2011, for the company with CIK 1288776. Unfortunately I cannot make sense of the naming convention within this directory.
One way around this is to simply use the SEC's search tool. However, this requires that I use a web crawler and I would prefer to use ftp directly.
Can anyone clarify how accession numbers are assigned? How do others go about pulling financials from the SEC website?
## Answer by chrisaycock (score 7)
https://quant.stackexchange.com/a/3314
Look in
```
edgar/full-index/{YYYY}/QTR{N}/xbrl.idx
```
You can grab the compressed version too:
```
xbrl.{Z,sit,gz,zip}
```
This will state what file you want.
For example, I want AOL's 10-K that was filed in the first quarter of 2012. So I download
```
edgar/full-index/2012/QTR1/xbrl.gz
```
After decompressing, I see that AOL's 10-K is available at
```
edgar/data/1468516/0001193125-12-076633.txt
```
## Answer by Dominik Fischer (score 1)
https://quant.stackexchange.com/a/33883
Check out this XBRL-Crawler: https://github.com/eliangcs/pystock-crawler It runs on Python but might be outdated.
You also can download the Text-Files using this Crawler: https://pypi.python.org/pypi/SECEdgar
I will test the first one, but the second one works fine.
## Answer by John F (score 1)
https://quant.stackexchange.com/a/80881
Accession number is the cik of the filer (not necessarily the company), the year filed, and the count of submitted filings for that year from that cik.
For more details about accession number look here.
XBRL is best accessed through the XBRL endpoint, example here.
If you want to download every XBRL and parse them into tables you can use the datamule python package, and then subset by year/quarter. It should take 10-20 minutes to download and parse every XBRL. Disclaimer: I'm the developer.
If you're looking for a more targeted approach try this page which details the different XBRL endpoints.
## Answer by Dimitri Vulis (score 1)
https://quant.stackexchange.com/a/85603
SEC's Division of Economic and Risk Analysis (DERA) has some helpful-looking examples on Github for accessing and analyzing SEC's XBRL Data Sets in Python / pandas / Jupyter:
https://github.com/sec-gov/python-for-dera-financial-datasets
## Answer by Jay (score 0)
https://quant.stackexchange.com/a/43219
sec-api (https://www.npmjs.com/package/sec-api) provides a websocket-based real-time API using Socket.IO - works with server-side (eg Node.js) and client-side (eg React, React Native, Angular, Vue) JavaScript.
The API returns new filings in JSON format, eg:
```
{
companyName: 'WALT DISNEY CO/ (0001001039) (Issuer)',
type: '4',
description: 'FORM 4',
linkToFilingDetails: 'https://www.sec.gov/...',
linkToHtmlAnnouncement: 'https://www.sec.gov/...',
announcedAt: '2018-12-21T20:02:07-05:00'
}
```
You can hook-up to the websocket channel, and ignore all filings where type doesn't equal '10-K'.
However, the API doesn't return historical filings.
The Node.js implementation seems to be very simple:
```
const api = require('sec-api')();
api.on('filing', filing => console.log(filing));
```Shown in full with attribution under the source's licence. Licence: CC BY-SA 4.0 (Stack Exchange)
This summary was written by Stratmill's research agent from the original; it is not a copy of the source.