Pro Forma Invoice Extractor

Extract the supplier, buyer, validity date, line items, VAT breakdown, and totals from a pro forma invoice PDF. Built on the EN 16931 e-invoicing model.

The full guide: Incoterms basis and validity window on a pro forma invoice

How should we start?

Build with Talonic

Need to scale? Create an API key, then run this from your own code or an agent.

Create an API key

Free account required

Start with a document

Upload a file or pick a sample, and see the fields come back.

No signup · nothing stored

Questions about Pro Forma Invoice Extractor.

What is a pro forma invoice and what does this tool extract from it?

A pro forma invoice is a preliminary invoice issued before goods ship or a final invoice is raised, often to confirm price and terms. The extractor reads the document number and date, the validity or expiration date, the supplier and buyer with their VAT IDs and postal addresses, the currency, and the subtotal, total without VAT, tax amount, and total.

Are the line items and VAT returned as tables?

Yes. The line items table returns each item with its name and description, quantity and unit code, net and gross unit price, line net amount, and VAT category code and rate, and the VAT breakdown table returns the taxable amount, tax amount, category code, rate, and any exemption reason per VAT category.

Does it follow an e-invoicing standard?

The field set follows the EN 16931 semantic data model, with Business Terms such as the buyer reference, contract and purchase order references, and the payment means code (UNTDID 4461), so a pro forma invoice maps to the same fields as a compliant e-invoice.

Which payment and reference fields come out?

Payment terms, due date, the payment IBAN and BIC, the remittance information, the delivery terms, and the invoicing period start and end, alongside the buyer accounting reference for booking the document.

What formats can I export, and what are the limits?

Download the fields and both tables as CSV, XLSX, or JSON. Uploads are PDF only, up to 10MB and 100 pages, and the document is not retained after extraction.

Doing this to one file, or to ten thousand?

The tool reads a single document. The platform reads the whole estate once and keeps it queryable — the same engine, with a memory.

See invoice data extraction if you run this for finance and AP teams, or the extraction API if you are building it in.