Validation Results
Get per-field validation results with match type, similarity score, and LLM judge verdict for every expected-versus-extracted value in a validation run.
Validation results are the per-field comparison records of a validation run: for every document-field pair in the golden sample, they show the expected value, the actual extracted value, a match type, a similarity score for near-misses, and an optional LLM judge verdict. They are the benchmark evidence behind a run's accuracy, not in-pipeline check outcomes, and they are the only place the public API exposes granular scores.
The match_type vocabulary distinguishes six outcomes — use it to separate real extraction errors (mismatch) from coverage gaps (missing_actual, where the run extracted nothing for a field the ground truth expects, and missing_expected, where the ground truth has no value to compare against). both_empty marks pairs where neither side has a value; most integrations treat exact and often partial as correct.
- exact — The extracted value matches the expected value character-for-character (after normalization).
- partial — The values are similar but not identical (e.g. formatting differences);
similarity_scorequantifies how close. - mismatch — The extracted value does not match the expected value.
- missing_actual — The ground truth expects a value but the job run extracted none.
- missing_expected — The job run extracted a value but the ground truth has none recorded.
- both_empty — Neither side has a value; usually excluded from accuracy math.
The judge_verdict field (correct, incorrect, missing, or uncertain) is written by a separate LLM judge pass that semantically reviews comparisons — for example a partial that is really the same invoice number with different punctuation. The judge runs as its own step from the platform's Benchmarks surface, so judge_verdict stays null until that pass has been executed for the run; deterministic match_type values are always present.
Results are returned in full, oldest-first, with no pagination, and rows are filtered by document visibility: comparisons for documents hidden from your key's minting user by source IAM rules are omitted. The judged_only query parameter is accepted for forward compatibility but currently does not filter — the endpoint returns all rows either way, so filter client-side on judge_verdict !== null when you only want judged comparisons.
completed status (poll GET /v1/validation/runs/{id}). Rows with match_type of mismatch or missing_actual are the fastest way to find systematically failing fields — group them by field_name client-side./v1/validation/runs/{id}/resultsQuery parameters
Request
curl https://api.talonic.com/v1/validation/runs/a1b2c3d4-e5f6-7890-abcd-ef1234567890/results \
-H "Authorization: Bearer tlnc_..."Response
Response fields
Response
{
"data": [
{
"id": "e5f6a7b8-c9d0-1234-efab-345678901234",
"validation_run_id": "a1b2c3d4-e5f6-7890-abcd-ef1234567890",
"document_id": "d4e5f6a7-b8c9-0123-defa-234567890123",
"field_name": "invoice_number",
"expected_value": "INV-2024-0042",
"actual_value": "INV 2024-0042",
"match_type": "partial",
"similarity_score": 0.93,
"confidence": 0.97,
"judge_verdict": "correct",
"created_at": "2024-09-14T10:35:00.000Z"
}
]
}Errors
Error responses