Catalogue and Price List Extractor

Extract the supplier, currency, products, SKUs, unit prices, tier pricing, and stock status from a catalogue or price list PDF into a spreadsheet.

The full guide: SKUs, GTINs, and quantity price breaks in a catalogue

How should we start?

Build with Talonic

Need to scale? Create an API key, then run this from your own code or an agent.

Create an API key

Free account required

Start with a document

Upload a file or pick a sample, and see the fields come back.

No signup · nothing stored

Questions about Catalogue and Price List Extractor.

What does the catalogue and price list extractor read?

The document number and date, the valid-from and valid-to dates, the supplier name, address, contact email, and phone, the currency (ISO 4217), and the payment and delivery terms.

Which product identifiers and pricing fields come out?

The GTIN and SKU, the product name, description, category, brand, and manufacturer, the country of origin, the unit price, the unit of measure (enum, such as pieces, kg, liters, or meters), the minimum order quantity, the discount percentage, and the availability status (enum, such as in stock, out of stock, discontinued, or on order).

Are the catalogue lines and tier pricing returned as tables?

Yes. A line items table returns each SKU, GTIN, product name, unit price, currency, unit of measure, minimum order quantity, discount, and availability; a pricing tiers table returns each SKU with its quantity-from, quantity-to, unit price, and discount; and a product attributes table returns each SKU with its attribute name and value.

Can it turn a long catalogue into a spreadsheet?

Yes. Because the line items and pricing tiers come back as tables, a multi-page price list becomes rows you can export to CSV or a spreadsheet, with each product price, discount tier, and stock status in its own column.

What are the file limits and privacy terms?

PDF only, up to 10MB and 100 pages. The catalogue is processed via the Talonic API for extraction only, is not retained for training, and is not shared.

Doing this to one file, or to ten thousand?

The tool reads a single document. The platform reads the whole estate once and keeps it queryable — the same engine, with a memory.

See document data extraction if you run this for freight and customs teams, or the extraction API if you are building it in.