Skip to main content

Create Benchmark

Start a benchmark run with POST /v1/quality/benchmarks to score extraction output against a ground truth dataset. Returns per-field and overall accuracy.

POST /v1/quality/benchmarks starts a benchmark run that evaluates your current extraction output against a ground truth dataset. The benchmark compares each document in the dataset entry-by-entry and field-by-field, producing an overall accuracy score and per-field breakdowns.

The typical workflow is: create a benchmark after making extraction changes, read the scores off the response, then drill into per-document results. Run multiple benchmarks against the same dataset over time to track accuracy trends.

The evaluation runs inside the request. The response is the finished benchmark: status: complete, accuracy_overall and accuracy_by_field populated, documents_processed equal to documents_total (the dataset's entry count), and duration_ms recording the evaluation time. A 50-entry dataset typically finishes in a few seconds.

A benchmark does not re-extract anything: at execution the engine takes each dataset entry, finds the latest completed extraction result for that document under the given user_schema_id, and compares the extracted values field-by-field against expected_data. Entries whose documents have no completed result for that schema simply cannot match — run your extraction (Spec pipeline or Structuring Run) over the benchmark documents before benchmarking, or accuracy will understate real quality.

Multiple benchmarks can run in parallel against different datasets. Use GET /v1/quality/benchmarks/compare after completion to compare two runs side by side. The dataset_id is fixed at creation — to benchmark against a different dataset, create a new run.

This is the same engine the platform's Benchmarks surface runs, so API and dashboard scores agree. Because the evaluation happens in the request, a large dataset means a longer request; keep your HTTP client timeout generous for datasets in the hundreds of entries.
POST/v1/quality/benchmarks

Body parameters

dataset_id*stringGround truth dataset to benchmark against. Must be a valid UUID.
user_schema_id*stringUser schema that defines the fields to evaluate. Must be a valid UUID.
namestringHuman-readable name for this run. Defaults to "Benchmark YYYY-MM-DD".

Request

curl -X POST https://api.talonic.com/v1/quality/benchmarks \
  -H "Authorization: Bearer tlnc_..." \
  -H "Content-Type: application/json" \
  -d '{
    "dataset_id": "a1b2c3d4-e5f6-7890-abcd-ef1234567890",
    "user_schema_id": "5e6f7a8b-9c0d-1234-abcd-ef0123456789",
    "name": "Post-schema-change check"
  }'

dataset_id is existence-checked at creation (an unknown id returns 404 not_found), and documents_total is snapshotted from the dataset's live entry count at that moment — entries added afterwards do not join an already-created run. user_schema_id is validated as a UUID only; make sure it names a real schema in your workspace, since the run scores fields against that schema's definition. If name is omitted, the run is named Benchmark YYYY-MM-DD from the creation date.

Response

Response fields (201 Created)

idstringBenchmark run UUID.
namestringBenchmark run name.
dataset_idstringGround truth dataset ID.
user_schema_idstringUser schema whose fields are evaluated.
statusstringcomplete — the evaluation ran inside the request.
accuracy_overallnumberOverall accuracy (0-1) across every expected field.
accuracy_by_fieldobjectPer-field accuracy (0-1).
documents_processedintegerEntries evaluated; equals documents_total on a complete run.
documents_totalintegerTotal entries in the dataset.
duration_msintegerEvaluation wall-clock time.
created_atstringISO 8601 creation timestamp.
completed_atstringISO 8601 completion timestamp.
links.selfstringURL to this benchmark run.
links.resultsstringURL to the per-document results.

Response (201 Created)

{
  "id": "c3d4e5f6-a7b8-9012-cdef-123456789012",
  "name": "Benchmark 2024-09-25",
  "dataset_id": "a1b2c3d4-e5f6-7890-abcd-ef1234567890",
  "user_schema_id": "5e6f7a8b-9c0d-1234-abcd-ef0123456789",
  "status": "complete",
  "accuracy_overall": 0.943,
  "accuracy_by_field": { "invoice_number": 1, "total_amount": 0.92, "vendor_name": 0.91 },
  "documents_processed": 50,
  "documents_total": 50,
  "duration_ms": 2380,
  "accuracy_delta": null,
  "compared_to_run_id": null,
  "created_at": "2024-09-25T12:00:00.000Z",
  "completed_at": null,
  "links": {
    "self": "/v1/quality/benchmarks/c3d4e5f6-a7b8-9012-cdef-123456789012",
    "results": "/v1/quality/benchmarks/c3d4e5f6-a7b8-9012-cdef-123456789012/results"
  }
}

Errors

Error responses

400validation_errorMissing or invalid dataset_id or user_schema_id (both must be UUIDs).
401unauthorizedMissing or invalid API key.
404not_foundThe specified dataset_id does not exist for your workspace.
429rate_limitedToo many requests. Retry after the period indicated in the Retry-After header.

Frequently asked questions

Can I run multiple benchmarks simultaneously?+
Yes. Benchmark runs are independent and can execute in parallel against different datasets or the same dataset.
Does the POST wait for the evaluation?+
Yes. The evaluation runs inside the request and the response is the finished run with status complete and the accuracy fields populated. duration_ms records the exact evaluation time; a 50-entry dataset typically finishes in a few seconds.
Does a benchmark re-extract my documents?+
No. It compares expected_data against the latest completed extraction result each document already has under the benchmark's user_schema_id. Make sure the documents have been processed with that schema before benchmarking.
Is user_schema_id required when creating a benchmark?+
Yes. Both dataset_id and user_schema_id are required UUIDs. The schema defines which fields the benchmark evaluates, so extracted values and ground truth entries are compared on the same field set.