Skip to main content

Create Validation Run

Start a validation run with POST /v1/validation/runs, registering a benchmark of one job run against a golden sample for per-field extraction accuracy.

POST /v1/validation/runs registers a validation run: a benchmark that compares the output of one job run (Structuring Run) against a golden sample of verified expected values. When the run executes, the engine classifies every document-field pair as exact, partial, or mismatch (plus missing-value categories), and an optional LLM judge pass adds semantic verdicts. This measures accuracy after extraction, unlike the in-pipeline checks that gate results before delivery.

The endpoint validates both references before inserting: dataspace_run_id is checked first (404 "Job run ... not found."), then golden_sample_id (404 "Golden sample ... not found."); both must be UUID v4 or the request fails with 400 VALIDATION_ERROR. The job run must be completed (a 400 validation_error otherwise) and must use the same schema as the golden sample. On success the comparison runs in the request and the run is returned already completed, with accuracy, total_comparisons and field_accuracy populated; if the engine hits an error the run is returned as failed with error_message.

This is the same engine the platform's Benchmarks surface (Review → Benchmarks) runs, so scores from the API and the dashboard agree. A golden sample with a few hundred values compares in well under a second; the request returns when the comparison is done, so there is nothing to poll.

For a meaningful benchmark, validate a completed Structuring Run whose schema matches the golden sample's user_schema_id — the engine compares field names defined by that schema, so a schema mismatch produces mostly missing comparisons rather than an error at creation time. If name is omitted, the run is named Validation YYYY-MM-DD from the creation date; use explicit names when creating several runs a day, since the list endpoint has no filters.

Runs are cheap and independent. Create one per job run you want scored and compare them over time with [GET /v1/validation/runs](list-validation-runs).
POST/v1/validation/runs

Body parameters

golden_sample_id*uuidGolden sample (Ground Truth dataset) to validate against.
dataspace_run_id*uuidJob run (Structuring Run) to validate.
namestringOptional human-readable name. Defaults to "Validation YYYY-MM-DD".

Request body

{
  "dataspace_run_id": "b2c3d4e5-f6a7-8901-bcde-f12345678901",
  "golden_sample_id": "c3d4e5f6-a7b8-9012-cdef-123456789012",
  "name": "Q1 Invoice accuracy check"
}

Request

curl -X POST https://api.talonic.com/v1/validation/runs \
  -H "Authorization: Bearer tlnc_..." \
  -H "Content-Type: application/json" \
  -d '{
    "dataspace_run_id": "b2c3d4e5-f6a7-8901-bcde-f12345678901",
    "golden_sample_id": "c3d4e5f6-a7b8-9012-cdef-123456789012",
    "name": "Q1 Invoice accuracy check"
  }'

Response

Response fields (201 Created)

idstringValidation run UUID.
namestringRun name.
statusstringcompleted, or failed when the engine hit an error (see error_message).
dataspace_run_idstringUUID of the job run being validated.
golden_sample_idstringUUID of the golden sample dataset.
accuracynumber | nullWeighted accuracy across all fields and documents (0-1); null on a failed run.
total_comparisonsinteger | nullNumber of document-field pairs compared; null on a failed run.
field_accuracyobject | nullPer-field breakdown with match-type counts and a per-field score.
error_messagestring | nullWhy a failed run failed.
created_atstringISO 8601 creation timestamp.
completed_atstring | nullNull until the run completes.
linksobjectRelated resource URLs (self, results).

Response (201 Created)

{
  "id": "a1b2c3d4-e5f6-7890-abcd-ef1234567890",
  "name": "Q1 Invoice accuracy check",
  "status": "completed",
  "dataspace_run_id": "b2c3d4e5-f6a7-8901-bcde-f12345678901",
  "golden_sample_id": "c3d4e5f6-a7b8-9012-cdef-123456789012",
  "accuracy": 0.9412,
  "total_comparisons": 68,
  "matched_documents": 17,
  "unmatched_golden": 0,
  "unmatched_dataset": 2,
  "field_accuracy": { "invoice_number": { "total": 17, "exact": 17, "partial": 0, "mismatch": 0, "missing_actual": 0, "missing_expected": 0, "both_empty": 0, "accuracy": 1 } },
  "error_message": null,
  "created_at": "2024-09-14T10:32:00.000Z",
  "completed_at": "2024-09-14T10:32:01.000Z",
  "links": {
    "self": "/v1/validation/runs/a1b2c3d4-e5f6-7890-abcd-ef1234567890",
    "results": "/v1/validation/runs/a1b2c3d4-e5f6-7890-abcd-ef1234567890/results"
  }
}

Errors

Error responses

400validation_errordataspace_run_id or golden_sample_id is missing or not a UUID v4.
401unauthorizedMissing or invalid API key.
404not_foundJob run or golden sample not found, or they do not belong to your organization.
429rate_limitedToo many requests. Retry after the period indicated in the Retry-After header.

Frequently asked questions

Does the POST wait for the comparison?+
Yes. The comparison runs inside the request and the response is the finished run (completed, or failed with error_message). There is no pending phase to poll through.
Can I run validation against the same dataset multiple times?+
Yes. You can create multiple validation runs against the same golden sample with different job runs to track accuracy improvements over time. Runs reference the dataset without consuming it.
What is the difference between a validation run and a benchmark run?+
Both are Benchmarks surfaces. A validation run (POST /v1/validation/runs) scores one specific job run against a golden sample of per-field expected values. A benchmark run (POST /v1/quality/benchmarks) scores extraction output against a dataset of per-document expected_data objects and supports run-to-run comparison.