EEO-1 Report Extractor

Extract the employer, EIN, NAICS code, and workforce counts by job category, race/ethnicity, and gender from an EEO-1 workforce demographic report PDF.

The full guide: EEO-1 headcounts by job category, race, and sex

How should we start?

Build with Talonic

Need to scale? Create an API key, then run this from your own code or an agent.

Create an API key

Free account required

Start with a document

Upload a file or pick a sample, and see the fields come back.

No signup · nothing stored

Questions about EEO-1 Report Extractor.

What does the EEO-1 report extractor read?

The report number and date, the employer legal name, EIN, and address, the reporting period, the NAICS industry code, the total employee count, the parent company EIN, the EEO form type (initial, annual, or amended), and the certifying officer name and certification date.

How is the workforce demographic data captured?

The workforce demographics table returns each job category, race/ethnicity, and gender cell with its male, female, and total counts, and the job category, race/ethnicity, gender, male count, and female count also come back as fields.

Are protected-class and establishment breakdowns returned as tables?

Yes. A protected class status table returns each job category with veteran status, disability status, and male and female counts, and an establishment breakdown table returns each establishment name, address, NAICS code, and total employees.

Which totals and filing fields come out?

The total male employees and total female employees across all categories, the veteran status and disability status enums, and the filing deadline (the EEOC due date) come back, so the aggregate headcounts and the submission cadence are captured.

The report contains workforce demographics. What happens to it?

The report carries aggregated demographic counts rather than named individuals. It is processed via the Talonic API for extraction only, is not retained for training, and is not shared. PDF only, up to 10MB and 100 pages.

Doing this to one file, or to ten thousand?

The tool reads a single document. The platform reads the whole estate once and keeps it queryable — the same engine, with a memory.

See PDF to Markdown if you run this for data and platform teams, or the extraction API if you are building it in.