Benchmark Results
Get per-document benchmark results showing matched and diverged fields, or compare two benchmark runs side by side to track extraction accuracy trends.
Benchmark results are the per-document accuracy records of a completed benchmark run: for every document, they show which fields matched the ground truth and which diverged. Each result includes the extracted value, the expected value, and a correct flag per field. A companion compare endpoint puts two runs side by side to quantify accuracy change.
Use the results endpoint after a run reaches complete to find where extraction fails: scan field_results for match: false entries and inspect the expected-versus-actual pairs. This tells you whether a low accuracy_overall comes from one systematically wrong field or from scattered document-level issues.
Each result row scores one dataset entry: accuracy is the fraction of compared fields that matched for that document, and field_results is an array with one object per expected_data key — field, expected, actual (null when nothing was extracted), a boolean match, and the extraction confidence. Matching is forgiving where it should be: numeric-typed fields match within a 0.01 tolerance and strings compare case-insensitively, so formatting noise does not masquerade as an extraction error.
A document with no completed extraction under the run's schema still produces a result row: it scores accuracy: 0 with every actual null, so coverage gaps are visible rather than silently skipped. Group match: false entries by field across all rows to build a per-field failure table — one systematically failing field (a date format, a currency symbol) usually explains most of a low overall score and points at a single schema or instruction fix.
Result rows are returned in full, oldest-first, without pagination, and are filtered by document visibility: rows for documents hidden from your API key's minting user are omitted, while the run's aggregate accuracy_overall still reflects every scored document. The same rows are embedded as results on GET /v1/quality/benchmarks/:id, so one call can serve both status polling and drill-down.
queued or running status has no results to return yet; poll GET /v1/quality/benchmarks/:id until status is complete./v1/quality/benchmarks/:id/resultsRequest (Results)
curl https://api.talonic.com/v1/quality/benchmarks/c3d4e5f6-a7b8-9012-cdef-123456789012/results \
-H "Authorization: Bearer tlnc_..."Results are returned in full, oldest-first, without pagination — one result row per dataset entry. Like ground-truth entries, results are filtered by document visibility: rows for documents hidden from your key's minting user are omitted, so a restricted key can see an accuracy_overall computed over more documents than its visible result rows. The same rows are also embedded as results on GET /v1/quality/benchmarks/:id, so you rarely need both calls.
Response (Results)
Response fields
Response
{
"data": [
{
"id": "d4e5f6a7-b8c9-0123-defa-234567890123",
"document_id": "doc_abc123",
"ground_truth_entry_id": "b2c3d4e5-f6a7-8901-bcde-f12345678901",
"accuracy": 0.67,
"field_results": [
{ "field": "vendor_name", "expected": "Acme Corp", "actual": "Acme Corp", "match": true, "confidence": 0.98 },
{ "field": "total_amount", "expected": 14250.00, "actual": "14000.00", "match": false, "confidence": 0.91 },
{ "field": "invoice_number", "expected": "INV-2024-0847", "actual": "INV-2024-0847", "match": true, "confidence": 0.97 }
],
"created_at": "2024-09-25T12:00:04.200Z"
}
]
}To track accuracy trends over time, compare two benchmark runs side by side. The accuracy_delta shows the difference in overall accuracy between the two runs.
/v1/quality/benchmarks/compareQuery parameters
Response (Compare)
Response fields
Response (Compare)
{
"run_a": {
"id": "c3d4e5f6-a7b8-9012-cdef-123456789012",
"name": "Benchmark 2024-09-25",
"status": "completed",
"accuracy_overall": 0.93,
"documents_total": 50,
"created_at": "2024-09-25T12:00:00.000Z"
},
"run_b": {
"id": "d4e5f6a7-b8c9-0123-defa-234567890123",
"name": "Benchmark 2024-10-01",
"status": "completed",
"accuracy_overall": 0.96,
"documents_total": 50,
"created_at": "2024-10-01T09:00:00.000Z"
},
"accuracy_delta": -0.03
}Errors
Error responses