SEC 10-K Report Extractor

Extract the entity name, CIK, fiscal year, auditor, and the full income statement, balance sheet, and cash flow figures from a 10-K annual report PDF.

The full guide: CIK and US-GAAP figures from an SEC Form 10-K

How should we start?

Build with Talonic

Need to scale? Create an API key, then run this from your own code or an agent.

Create an API key

Free account required

Start with a document

Upload a file or pick a sample, and see the fields come back.

No signup · nothing stored

Questions about SEC 10-K Report Extractor.

Which 10-K header fields does the extractor read?

The entity name, the company CIK (Central Index Key) and ticker, the fiscal year and fiscal year end, the fiscal period, the reporting currency, and the filing document number and date, plus the auditor name, the auditor opinion (unqualified, qualified, adverse, or disclaimer), and whether internal controls over financial reporting are effective.

Are the financial statements returned as tables?

Yes. The consolidated balance sheet, the statement of operations, the statement of cash flows, and the statement of stockholders equity each come back as their own table, alongside segment financial data, selected financial data, quarterly results, and executive compensation.

Does it pull the headline financials as fields too?

Total revenue, gross profit, operating income, and net income, basic and diluted earnings per share, total assets, total liabilities, and stockholders equity, and the cash flow from operating, investing, and financing activities are each read as their own field for quick review.

Does it capture the narrative Items?

The business description (Item 1), the risk factors (Item 1A) as a list, and the management discussion and analysis (Item 7) come back as text, so the qualitative disclosures sit next to the numbers.

Does it follow a reporting standard, and what are the limits?

The field set is modeled on US-GAAP and IFRS taxonomy statement concepts, so filings prepared under either framework map to the same columns. PDF only, up to 10MB and 100 pages, and the filing is not retained after extraction.

Doing this to one file, or to ten thousand?

The tool reads a single document. The platform reads the whole estate once and keeps it queryable — the same engine, with a memory.

See PDF to Markdown if you run this for data and platform teams, or the extraction API if you are building it in.