Penetration Test Report Extractor

Extract the client, tester, scope, risk rating, severity counts, vulnerabilities with CVSS scores, and remediation actions from a penetration test report PDF.

The full guide: CVSS scores and findings in a penetration test report

How should we start?

Build with Talonic

Need to scale? Create an API key, then run this from your own code or an agent.

Create an API key

Free account required

Start with a document

Upload a file or pick a sample, and see the fields come back.

No signup · nothing stored

Questions about Penetration Test Report Extractor.

What does the penetration test report extractor read?

The report number and date, the test period start and end dates, the client organization and contact person, the testing firm, the scope description and exclusions, the testing type (External, Internal, White-box, Black-box, Gray-box), and the assessment scope type (Infrastructure, Web Application, API, Network, or Cloud).

Which risk-rating and count fields come out?

The overall risk rating (Critical, High, Medium, Low), the critical, high, medium, low, and total findings counts, the recurring finding count, the data compromise risk, the authenticated versus unauthenticated findings breakdown, and the executive summary.

Are the vulnerabilities returned as tables?

Yes. A vulnerabilities table returns each id, title, description, severity, CVSS score, affected system, remediation, and authentication-required flag; a findings table returns each category, risk level, evidence, and impact; and remediation actions, tested systems, and risk summary tables also come back.

Does it capture the methodology and compliance standards?

The methodology (OWASP Testing Guide, PTES, NIST SP 800-115, CVSS) and the compliance standards assessed (PCI-DSS, HIPAA, GDPR, ISO 27001, SOC 2) come back as fields, alongside the remediation deadline. The tool extracts the findings as written and does not perform its own security assessment or advise on fixes.

What are the file limits and privacy terms?

PDF only, up to 10MB and 100 pages. The report is processed via the Talonic API for extraction, is not retained for training, and is not shared.

Doing this to one file, or to ten thousand?

The tool reads a single document. The platform reads the whole estate once and keeps it queryable — the same engine, with a memory.

See PDF to Markdown if you run this for data and platform teams, or the extraction API if you are building it in.