Skip to main content

Validation Results

Get per-field validation results with match type, similarity score, and LLM judge verdict for every expected-versus-extracted value in a validation run.

Validation results are the per-field comparison records of a validation run: for every document-field pair in the golden sample, they show the expected value, the actual extracted value, a match type, a similarity score for near-misses, and an optional LLM judge verdict. They are the benchmark evidence behind a run's accuracy, not in-pipeline check outcomes, and they are the only place the public API exposes granular scores.

The match_type vocabulary distinguishes six outcomes — use it to separate real extraction errors (mismatch) from coverage gaps (missing_actual, where the run extracted nothing for a field the ground truth expects, and missing_expected, where the ground truth has no value to compare against). both_empty marks pairs where neither side has a value; most integrations treat exact and often partial as correct.

  • exact — The extracted value matches the expected value character-for-character (after normalization).
  • partial — The values are similar but not identical (e.g. formatting differences); similarity_score quantifies how close.
  • mismatch — The extracted value does not match the expected value.
  • missing_actual — The ground truth expects a value but the job run extracted none.
  • missing_expected — The job run extracted a value but the ground truth has none recorded.
  • both_empty — Neither side has a value; usually excluded from accuracy math.

The judge_verdict field (correct, incorrect, missing, or uncertain) is written by a separate LLM judge pass that semantically reviews comparisons — for example a partial that is really the same invoice number with different punctuation. The judge runs as its own step from the platform's Benchmarks surface, so judge_verdict stays null until that pass has been executed for the run; deterministic match_type values are always present.

Results are returned in full, oldest-first, with no pagination, and rows are filtered by document visibility: comparisons for documents hidden from your key's minting user by source IAM rules are omitted. The judged_only query parameter is accepted for forward compatibility but currently does not filter — the endpoint returns all rows either way, so filter client-side on judge_verdict !== null when you only want judged comparisons.

Fetch results only after the run reaches completed status (poll GET /v1/validation/runs/{id}). Rows with match_type of mismatch or missing_actual are the fastest way to find systematically failing fields — group them by field_name client-side.
GET/v1/validation/runs/{id}/results

Query parameters

judged_onlystringAccepted for forward compatibility. Currently returns all rows regardless of value — filter on judge_verdict client-side.

Request

curl https://api.talonic.com/v1/validation/runs/a1b2c3d4-e5f6-7890-abcd-ef1234567890/results \
  -H "Authorization: Bearer tlnc_..."

Response

Response fields

dataarrayArray of per-field validation result objects, oldest first. Filtered by document visibility for your key.
data[].idstringResult UUID.
data[].validation_run_idstringParent validation run UUID.
data[].document_idstringDocument UUID.
data[].field_namestringField key that was compared.
data[].expected_valuestring | nullGround-truth expected value.
data[].actual_valuestring | nullExtracted value from the job run.
data[].match_typestring | nullComparison outcome: exact, partial, mismatch, missing_actual, missing_expected, or both_empty.
data[].similarity_scorenumber | nullString similarity (0–1) between expected and actual, populated for partial matches.
data[].confidencenumber | nullExtraction confidence for the actual value.
data[].judge_verdictstring | nullLLM judge verdict: correct, incorrect, missing, or uncertain. Null until the judge pass runs.
data[].created_atstringISO 8601 timestamp.

Response

{
  "data": [
    {
      "id": "e5f6a7b8-c9d0-1234-efab-345678901234",
      "validation_run_id": "a1b2c3d4-e5f6-7890-abcd-ef1234567890",
      "document_id": "d4e5f6a7-b8c9-0123-defa-234567890123",
      "field_name": "invoice_number",
      "expected_value": "INV-2024-0042",
      "actual_value": "INV 2024-0042",
      "match_type": "partial",
      "similarity_score": 0.93,
      "confidence": 0.97,
      "judge_verdict": "correct",
      "created_at": "2024-09-14T10:35:00.000Z"
    }
  ]
}

Errors

Error responses

401unauthorizedMissing or invalid API key.
404not_foundValidation run not found or does not belong to your organization.
429rate_limitedToo many requests. Retry after the period indicated in the Retry-After header.

Frequently asked questions

What is the difference between match_type and judge_verdict?+
match_type is a deterministic comparison outcome (exact, partial, mismatch, missing_actual, missing_expected, both_empty) computed for every row. judge_verdict (correct, incorrect, missing, uncertain) is an LLM-based semantic assessment written by a separate judge pass, and stays null until that pass runs.
When is the LLM judge invoked?+
The judge is a distinct pass launched from the platform's Benchmarks surface, typically to review partial matches and mismatches semantically. It is not triggered automatically by creating a run through the API, so expect null verdicts on rows until a judge pass has been executed.
Does judged_only=true filter the results?+
Not currently — the parameter is accepted but the endpoint returns all rows either way. Filter client-side on judge_verdict !== null to isolate judged comparisons.
How do I compute an accuracy score from results?+
Count rows by match_type: exact / (all rows minus both_empty) is the strict score, and (exact + partial) / (all rows minus both_empty) is a lenient one. The run object's own accuracy field is currently always null through the public API, so this client-side computation is the reliable path.