KYC and AML Extractor

Extract the subject, identification, PEP status, verification method, risk level, and watchlist screening results from a KYC or AML documentation PDF.

The full guide: Beneficial owners and PEP status in a KYC file

How should we start?

Build with Talonic

Need to scale? Create an API key, then run this from your own code or an agent.

Create an API key

Free account required

Start with a document

Upload a file or pick a sample, and see the fields come back.

No signup · nothing stored

Questions about KYC and AML Extractor.

What does the KYC and AML extractor read about the subject?

The document number and date, the subject full name and type (enum, such as individual or entity), date of birth, identification type (enum) and number, country of origin, residential address, and occupation, and whether the subject is a politically exposed person (PEP).

Which verification and risk fields come out?

The verification status (enum), date, and method (enum), the risk level (enum), the risk factors (array), whether a watchlist screening match was found (boolean) and the screening source, the AML report reference, and whether enhanced due diligence is required. The tool reports these fields as written and gives no compliance judgment.

Are the screening and identification records returned as tables?

Yes. The identification documents table returns each document type, number, issuing country, issue and expiration dates, and verification status; the screening matches table returns each screening date, source, match name and type, score, and action taken; and separate beneficial owners, risk assessment history, and compliance reviews tables come back too.

Does it capture the review and transaction limits?

The reviewer name and review date, the review notes, the transaction limit and currency, the source document, and the governing law each come back as fields, so the outcome of the review sits next to the identity data.

The document has personal and identity data. What happens to it?

The document is processed via the Talonic API for extraction only, is not retained for training, and is not shared. Redact identification numbers and dates of birth where you can. PDF only, up to 10MB and 100 pages.

Doing this to one file, or to ten thousand?

The tool reads a single document. The platform reads the whole estate once and keeps it queryable — the same engine, with a memory.

See document data extraction if you run this for legal and commercial teams, or the extraction API if you are building it in.