Skip to content
All library documents

Mapping SEC XBRL Facts to Standardized Financial Statements

Article Quant Q&A · Author: Jan Felix

Summary

The document describes the challenge of turning SEC filings into readable balance sheets and income statements. Its author retrieves reported facts for thousands of companies, using filing identifiers and fields such as reporting date, accounting tag, period duration, and value. The aim is to group many company-specific entries into familiar categories such as working capital and debt, while supporting comparisons across firms and time.

The text frames this as a data-modeling question and asks how to design the classification and normalization process. It provides no proposed mapping method, implementation results, or discussion of validation. The source facts can be numerous and vary across issuers, so a usable standardized statement would need to handle differing tags and reporting practices; those issues remain open in the document.

Key ideas

  • SEC filing facts can be retrieved across many companies using filing identifiers.
  • Reported tags and period fields provide inputs for constructing financial statements.
  • Standardized categories such as working capital and debt would make raw facts easier to analyze.
  • The document raises the normalization problem but does not supply a solution or validation approach.

Tags

Full text
# From SEC Comprehensive Data Set to Clean Balance Sheets/Income Statements


# From SEC Comprehensive Data Set to Clean Balance Sheets/Income Statements












I am looking for advice on how to smartly get from unstructured financial data to a clean summary of balance sheets. I think of this not as a coding problem but of a question how to approach the task - and whether anyone has some information that might help.

I use the following SEC Data base to the 10-K and all 10-Q report of thousands of companies. It is relatively simple to filter each quaterly datafile by adsh (unique ID for all reports) to retrieve something akin to the below (Apple 20210 10-K).

In this picture the "tag" column reports the balance sheet item, "ddate" marks the end of the reporting period, "qtrs"=0 implies a point in time account (aka balance sheet item) and "value" is the USD amount in the balance sheet.

Amazingly, I can loop this process over all companies in the SEC data base to get the data for all companies. However, each 10-K for larger companies has hundreds of entries. My goal would be to have a relatively complete balance sheet in a readable format, that clusters balance sheet items into relevant groups (i.e. Working capital, debt, etc.) as you would know from standard financial models and is ideally flexible enough to apply this to a variety of companies through time.

My question is: How would you approach this?

Please see here for my notebook where I try this out, note that the process requires the user to download the files first and set up the folder structure.

Shown in full with attribution under the source's licence. Licence: CC BY-SA 4.0 (Stack Exchange)

This summary was written by Stratmill's research agent from the original; it is not a copy of the source.