Extract the company, fiscal year, revenue, net income, total assets, auditor opinion, and the financial statement tables from an annual report PDF.
The full guide: Companies House number and audit opinion, UK annual report
Need to scale? Create an API key, then run this from your own code or an agent.
Create an API keyFree account required
Upload a file or pick a sample, and see the fields come back.
No signup · nothing stored
The document number (accession number), filing date, and fiscal year end, the company legal name, registration number, jurisdiction, address, and industry sector, and the headline figures: total revenue, net income, total assets, shareholders equity, currency, and employee count, plus capital expenditures, cash flow from operations, long-term debt, current liabilities, research and development expense, and the effective tax rate.
Yes. The consolidated statement of income (line item, current-period and prior-period amount), the balance sheet, the cash flow statement, a segment revenue table, and a financial ratios table each come back as their own table.
The external auditor name and the audit opinion (unqualified or clean, qualified, adverse, or disclaimer) are read as fields, and the governance structure, a board of directors table (member, position, appointment date, independence status), and an executive officers table are captured.
The business description, key risks (with a risk factors table of category, description, and mitigation strategy), accounting policies, segment performance, dividend information, related party transactions, and executive compensation each come back as text so the qualitative disclosures sit next to the numbers.
The field set captures the figures common to annual reports prepared under IFRS or US-GAAP, and the company jurisdiction field records the framework that applies, so filings from different countries map to the same fields. PDF only, up to 10MB and 100 pages, and the report is not retained after extraction.
The tool reads a single document. The platform reads the whole estate once and keeps it queryable — the same engine, with a memory.
See PDF to Markdown if you run this for data and platform teams, or the extraction API if you are building it in.