Skip to main content

Benchmark Results

Get per-document benchmark results showing matched and diverged fields, or compare two benchmark runs side by side to track extraction accuracy trends.

Benchmark results are the per-document accuracy records of a completed benchmark run: for every document, they show which fields matched the ground truth and which diverged. Each result includes the extracted value, the expected value, and a correct flag per field. A companion compare endpoint puts two runs side by side to quantify accuracy change.

Use the results endpoint after a run reaches complete to find where extraction fails: scan field_results for match: false entries and inspect the expected-versus-actual pairs. This tells you whether a low accuracy_overall comes from one systematically wrong field or from scattered document-level issues.

Each result row scores one dataset entry: accuracy is the fraction of compared fields that matched for that document, and field_results is an array with one object per expected_data key — field, expected, actual (null when nothing was extracted), a boolean match, and the extraction confidence. Matching is forgiving where it should be: numeric-typed fields match within a 0.01 tolerance and strings compare case-insensitively, so formatting noise does not masquerade as an extraction error.

A document with no completed extraction under the run's schema still produces a result row: it scores accuracy: 0 with every actual null, so coverage gaps are visible rather than silently skipped. Group match: false entries by field across all rows to build a per-field failure table — one systematically failing field (a date format, a currency symbol) usually explains most of a low overall score and points at a single schema or instruction fix.

Result rows are returned in full, oldest-first, without pagination, and are filtered by document visibility: rows for documents hidden from your API key's minting user are omitted, while the run's aggregate accuracy_overall still reflects every scored document. The same rows are embedded as results on GET /v1/quality/benchmarks/:id, so one call can serve both status polling and drill-down.

Results are only meaningful for finished runs. A run still in queued or running status has no results to return yet; poll GET /v1/quality/benchmarks/:id until status is complete.
GET/v1/quality/benchmarks/:id/results

Request (Results)

curl https://api.talonic.com/v1/quality/benchmarks/c3d4e5f6-a7b8-9012-cdef-123456789012/results \
  -H "Authorization: Bearer tlnc_..."

Results are returned in full, oldest-first, without pagination — one result row per dataset entry. Like ground-truth entries, results are filtered by document visibility: rows for documents hidden from your key's minting user are omitted, so a restricted key can see an accuracy_overall computed over more documents than its visible result rows. The same rows are also embedded as results on GET /v1/quality/benchmarks/:id, so you rarely need both calls.

Response (Results)

Response fields

dataarrayArray of per-document result objects.
data[].idstringResult UUID.
data[].document_idstringDocument evaluated.
data[].ground_truth_entry_idstringGround truth entry compared against.
data[].accuracynumberAccuracy score for this document (0–1): matched fields / compared fields.
data[].field_resultsarrayOne object per expected field: `field`, `expected`, `actual` (null if not extracted), `match` (boolean), `confidence`.
data[].created_atstringISO 8601 timestamp.

Response

{
  "data": [
    {
      "id": "d4e5f6a7-b8c9-0123-defa-234567890123",
      "document_id": "doc_abc123",
      "ground_truth_entry_id": "b2c3d4e5-f6a7-8901-bcde-f12345678901",
      "accuracy": 0.67,
      "field_results": [
        { "field": "vendor_name", "expected": "Acme Corp", "actual": "Acme Corp", "match": true, "confidence": 0.98 },
        { "field": "total_amount", "expected": 14250.00, "actual": "14000.00", "match": false, "confidence": 0.91 },
        { "field": "invoice_number", "expected": "INV-2024-0847", "actual": "INV-2024-0847", "match": true, "confidence": 0.97 }
      ],
      "created_at": "2024-09-25T12:00:04.200Z"
    }
  ]
}

To track accuracy trends over time, compare two benchmark runs side by side. The accuracy_delta shows the difference in overall accuracy between the two runs.

GET/v1/quality/benchmarks/compare

Query parameters

run_a*stringFirst benchmark run ID.
run_b*stringSecond benchmark run ID to compare against.

Response (Compare)

Response fields

run_aobjectFull benchmark object for the first run.
run_bobjectFull benchmark object for the second run.
accuracy_deltanumber | nullDifference in overall accuracy (run_a minus run_b). Null if either run has no accuracy score yet.

Response (Compare)

{
  "run_a": {
    "id": "c3d4e5f6-a7b8-9012-cdef-123456789012",
    "name": "Benchmark 2024-09-25",
    "status": "completed",
    "accuracy_overall": 0.93,
    "documents_total": 50,
    "created_at": "2024-09-25T12:00:00.000Z"
  },
  "run_b": {
    "id": "d4e5f6a7-b8c9-0123-defa-234567890123",
    "name": "Benchmark 2024-10-01",
    "status": "completed",
    "accuracy_overall": 0.96,
    "documents_total": 50,
    "created_at": "2024-10-01T09:00:00.000Z"
  },
  "accuracy_delta": -0.03
}

Errors

Error responses

400bad_requestBoth run_a and run_b query parameters are required for the compare endpoint.
401unauthorizedMissing or invalid API key.
404not_foundOne or both benchmark run IDs not found for your workspace.
429rate_limitedToo many requests. Retry after the period indicated in the Retry-After header.

Frequently asked questions

How is field accuracy calculated?+
Every key in the entry's expected_data is compared against the extracted value for that field: exact equality, with a 0.01 tolerance for numeric-typed fields and case-insensitive string matching. Per-document accuracy is matched/compared fields; accuracy_overall is the matched share of all field comparisons across the run.
What does a negative accuracy_delta mean?+
A negative delta means run_a has lower accuracy than run_b. For example, -0.03 means run_a is 3 percentage points less accurate. Use chronological ordering (older run as run_a) to see improvement as a positive delta.
Can I compare runs from different datasets?+
Yes, but the comparison only shows overall accuracy differences. Per-field comparisons are most meaningful when both runs use the same ground truth dataset.
Why do I see fewer result rows than dataset entries?+
Document-visibility filtering: rows for documents hidden from your key's minting user by source IAM rules are omitted from the response, while still counting toward the aggregate score. Entries without a completed extraction are not skipped — they appear as accuracy 0 rows with null actual values.