Skip to main content

List Validation Runs

List validation runs with GET /v1/validation/runs. Each run scores a job run against a golden sample dataset; poll status and drill into per-field results.

A validation run is a benchmark: it compares the structured output of one job run (a Structuring Run in the platform) against a golden sample dataset and produces per-field comparison records. GET /v1/validation/runs lists the newest 100 validation runs for your organization, most recent first. This is Benchmarks — accuracy measurement after the fact — not the in-pipeline validation checks that gate results before delivery.

Each run carries a status from the lifecycle pending → queued → processing → completed (or failed). Runs created through the public API execute in the request and come back already completed or failed; only runs created from older clients can still sit in pending. There is no running status.

accuracy is the run's weighted score across all fields and documents (0-1) and total_comparisons is the number of document-field pairs compared; both are null until the run completes. field_accuracy carries the per-field breakdown (match-type counts and a per-field score), and matched_documents / unmatched_golden / unmatched_dataset show how many documents were paired. For row-level detail, fetch [GET /v1/validation/runs/{id}/results](get-validation-results).

The list includes every validation run in the organization, not only API-created ones: benchmarks launched from the platform (golden-sample comparisons, CSV comparisons, unified and LLM-judge runs) appear here too, serialized to the same shape. Use dataspace_run_id to correlate a run with the job it scored and golden_sample_id to correlate it with its dataset; both come back null on run types that do not use them (a CSV comparison has no job run, for example).

The list returns up to 100 runs ordered by created_at descending, with no pagination. Delete obsolete runs with DELETE /v1/validation/runs/{id} if older runs you still need fall outside the window.
GET/v1/validation/runs

Request

curl https://api.talonic.com/v1/validation/runs \
  -H "Authorization: Bearer tlnc_..."

Response

Response fields

dataarrayArray of validation run objects (up to 100, ordered by created_at descending).
data[].idstringValidation run UUID.
data[].namestringHuman-readable run name.
data[].statusstringRun status: pending, queued, processing, completed, or failed.
data[].dataspace_run_idstring | nullUUID of the job run (Structuring Run) being validated.
data[].golden_sample_idstring | nullUUID of the golden sample dataset.
data[].accuracynumber | nullWeighted accuracy across all fields and documents (0-1). Null until the run completes.
data[].total_comparisonsinteger | nullNumber of document-field pairs compared. Null until the run completes.
data[].matched_documentsintegerDocuments paired between the job run and the golden sample.
data[].unmatched_goldenintegerGolden sample documents with no corresponding job-run document.
data[].unmatched_datasetintegerJob-run documents with no corresponding golden sample document.
data[].field_accuracyobject | nullPer-field breakdown: total, exact, partial, mismatch, missing_actual, missing_expected, both_empty, accuracy.
data[].error_messagestring | nullWhy a failed run failed.
data[].created_atstringISO 8601 creation timestamp.
data[].completed_atstring | nullISO 8601 completion timestamp.
data[].linksobjectRelated resource URLs (self, results).

Response

{
  "data": [
    {
      "id": "a1b2c3d4-e5f6-7890-abcd-ef1234567890",
      "name": "Q1 Invoice accuracy check",
      "status": "completed",
      "dataspace_run_id": "b2c3d4e5-f6a7-8901-bcde-f12345678901",
      "golden_sample_id": "c3d4e5f6-a7b8-9012-cdef-123456789012",
      "accuracy": 0.9412,
      "total_comparisons": 68,
      "matched_documents": 17,
      "unmatched_golden": 0,
      "unmatched_dataset": 2,
      "field_accuracy": { "invoice_number": { "total": 17, "exact": 17, "partial": 0, "mismatch": 0, "missing_actual": 0, "missing_expected": 0, "both_empty": 0, "accuracy": 1 } },
      "error_message": null,
      "created_at": "2024-09-14T10:32:00.000Z",
      "completed_at": "2024-09-14T10:35:00.000Z",
      "links": {
        "self": "/v1/validation/runs/a1b2c3d4-e5f6-7890-abcd-ef1234567890",
        "results": "/v1/validation/runs/a1b2c3d4-e5f6-7890-abcd-ef1234567890/results"
      }
    }
  ]
}

Errors

Error responses

401unauthorizedMissing or invalid API key.
429rate_limitedToo many requests. Retry after the period indicated in the Retry-After header.

Frequently asked questions

How many validation runs are returned?+
Up to 100 runs are returned, ordered by created_at descending, in a single response without pagination. Older runs beyond the newest 100 are not listed.
Why is accuracy null on a run?+
The run has not completed. accuracy and total_comparisons are filled in when the comparison finishes; a failed run keeps them null and explains itself in error_message.
What is dataspace_run_id in the response?+
It is the UUID of the job run (Structuring Run) whose output was validated. golden_sample_id identifies the golden sample dataset the output was compared against.
What statuses can a validation run have?+
pending, queued and processing while the engine is comparing, then completed or failed. Runs created through the public API execute in the request and return already completed or failed; there is no running status.
Can I filter the run list by job or dataset?+
No — the endpoint takes no query parameters. Fetch the newest 100 runs and filter client-side on dataspace_run_id or golden_sample_id. If runs you need regularly fall outside the 100-run window, delete superseded runs to keep the window useful.