Sources and Methods for Obtaining S&P 500 10-K Data
Summary
The document collects suggestions for obtaining annual filing data for an analysis of S&P 500 companies. Responses point to financial data packages and providers, including tidyquant, Quandl, and Compustat, as well as tools for downloading filings from the SEC’s EDGAR database. One response distinguishes accounting data already organized in a database from the underlying filing documents, which may need substantial processing before they are usable in a regression.
The answers describe several possible routes rather than comparing them systematically. Compustat is presented as a canonical source for academic accounting research, while direct filing extraction offers access to source documents but can require extensive programming and testing. Other responses mention APIs, open-source download tools, and extraction services for selected filing sections. The document gives no assessment of data coverage, licensing, historical consistency, or accuracy, and its package and service suggestions may change over time. Researchers should check whether a source supplies the specific fields and filing periods their analysis requires.
Key ideas
- Financial data packages and commercial databases can provide accounting information derived from company filings.
- Compustat is identified as a widely used source for academic accounting data.
- Downloading and extracting information directly from SEC filings can require substantial engineering and validation.
- Some tools focus on retrieving filings, while others extract selected sections or provide structured financial data.
- The answers do not compare coverage, cost, or suitability for a particular regression.
Tags
Full text
# How to download all 10-K reports for all companies listed on S&P 500?
# How to download all 10-K reports for all companies listed on S&P 500?
I am doing a regression analysis of all companies listed on s&p 500. It requires their 10-k reports. Where can I download all of them once?
## Answer by phlsmk (score 3, accepted)
https://quant.stackexchange.com/a/35248
Tidyquant also has a nice function tq_get() for getting all sorts of equity data from freely available sources including financial statements.
http://www.business-science.io/code-tools/2017/01/01/tidyquant-introduction.html
## Answer by Mustard Tiger (score 2)
https://quant.stackexchange.com/a/26090
The API's found on this site http://developer.edgar-online.com/docs allow you to acces historical SEC filings for most securities.
## Answer by Atul Agarawal (score 2)
https://quant.stackexchange.com/a/26108
If you are using R then try Quandl package to download the data, there you can find almost every kind of report.
## Answer by Matthew Gunn (score 2)
https://quant.stackexchange.com/a/35272
#### 1. Another data provider: Compustat
The canonical source in academic research for the accounting data disclosed in 10-K filings is the Compustat annual database. The Compustat quarterly database contains information listed in 10-Qs and 10-Ks (and can be a bit trickier to work with).
I have no idea on the cost of those products.
#### 2. Downloading 10-Ks and extracting data yourself...
I would NOT recommend this. It's a huge programming project. (I did it myself once to extract share repurchase data from the HTML of 10-Q, 10-K filings and it took months and months of testing.)
## Answer by Jordan-M-Young (score 0)
https://quant.stackexchange.com/a/55333
I know this is an old question but I actually wrote a script that downloads 10-K docs off of the SEC's EDGAR database website.
https://github.com/Jordan-M-Young/pyQuarry
Check the '$E(' folder.
Hope this helps!
## Answer by Ik32 (score 0)
https://quant.stackexchange.com/a/80106
you can extract sensitive sections like 1A,7,7A for your regression analysis using extract API . Here the github link The extracts is in text format. By calling the API you are not calling directly SEC for most volume sensitive requests.
## Answer by John F (score 0)
https://quant.stackexchange.com/a/80738
You can use the open-source datamule python package. Should take 30 minutes or so for SP 500.
```
pip install datamule
```
```
import datamule as dm
downloader = dm.Downloader()
downloader.download(form='10-K',ticker=[TSLA,META,...])
```Shown in full with attribution under the source's licence. Licence: CC BY-SA 4.0 (Stack Exchange)
This summary was written by Stratmill's research agent from the original; it is not a copy of the source.