Cargo Manifest Extractor

Extract the vessel and voyage, ports of loading and discharge, exporter and consignee EORI, commodity codes, and line items from a cargo manifest PDF.

The full guide: Advance cargo declaration data from a cargo manifest

How should we start?

Build with Talonic

Need to scale? Create an API key, then run this from your own code or an agent.

Create an API key

Free account required

Start with a document

Upload a file or pick a sample, and see the fields come back.

No signup · nothing stored

Questions about Cargo Manifest Extractor.

What does the cargo manifest extractor read?

The manifest number and date, the declaration type (import, export, transit, or customs warehouse), the exporter or shipper and its EORI, the consignee and its EORI, the vessel name and IMO number, the voyage number, the port of loading and port of discharge, and the incoterms.

Are the goods returned as line items?

Yes. The line items table returns each line with its line number, commodity code (Combined Nomenclature or HS), description, quantity, unit of measure, net mass, gross mass, item price amount, country of origin, and preference; a containers table (container number, type, seal number, seal type, status) and a charges table also come back.

Which weights, countries, and customs fields come out?

The total weight and weight unit, total volume, net mass, total packages and package type, the country of origin, country of destination, and country of dispatch or export, the valuation method, the transport-charges method of payment, the location of goods, the supporting documents, and the manifest status (filed, cleared, released, or held).

Which data model does it follow?

The field labels follow the EU Customs Data Model (EUCDM) data-element catalogue, an open EU TAXUD proxy for the WCO Data Model, so manifests routed through different customs systems map to the same columns.

What are the file limits and privacy terms?

PDF only, up to 10MB and 100 pages. The manifest is processed via the Talonic API for extraction, is not retained for training, and is not shared.

Doing this to one file, or to ten thousand?

The tool reads a single document. The platform reads the whole estate once and keeps it queryable — the same engine, with a memory.

See PDF to Markdown if you run this for data and platform teams, or the extraction API if you are building it in.