Create Benchmark
Start a benchmark run with POST /v1/quality/benchmarks to score extraction output against a ground truth dataset. Returns per-field and overall accuracy.
POST /v1/quality/benchmarks starts a benchmark run that evaluates your current extraction output against a ground truth dataset. The benchmark compares each document in the dataset entry-by-entry and field-by-field, producing an overall accuracy score and per-field breakdowns.
The typical workflow is: create a benchmark after making extraction changes, read the scores off the response, then drill into per-document results. Run multiple benchmarks against the same dataset over time to track accuracy trends.
The evaluation runs inside the request. The response is the finished benchmark: status: complete, accuracy_overall and accuracy_by_field populated, documents_processed equal to documents_total (the dataset's entry count), and duration_ms recording the evaluation time. A 50-entry dataset typically finishes in a few seconds.
A benchmark does not re-extract anything: at execution the engine takes each dataset entry, finds the latest completed extraction result for that document under the given user_schema_id, and compares the extracted values field-by-field against expected_data. Entries whose documents have no completed result for that schema simply cannot match — run your extraction (Spec pipeline or Structuring Run) over the benchmark documents before benchmarking, or accuracy will understate real quality.
Multiple benchmarks can run in parallel against different datasets. Use GET /v1/quality/benchmarks/compare after completion to compare two runs side by side. The dataset_id is fixed at creation — to benchmark against a different dataset, create a new run.
/v1/quality/benchmarksBody parameters
Request
curl -X POST https://api.talonic.com/v1/quality/benchmarks \
-H "Authorization: Bearer tlnc_..." \
-H "Content-Type: application/json" \
-d '{
"dataset_id": "a1b2c3d4-e5f6-7890-abcd-ef1234567890",
"user_schema_id": "5e6f7a8b-9c0d-1234-abcd-ef0123456789",
"name": "Post-schema-change check"
}'dataset_id is existence-checked at creation (an unknown id returns 404 not_found), and documents_total is snapshotted from the dataset's live entry count at that moment — entries added afterwards do not join an already-created run. user_schema_id is validated as a UUID only; make sure it names a real schema in your workspace, since the run scores fields against that schema's definition. If name is omitted, the run is named Benchmark YYYY-MM-DD from the creation date.
Response
Response fields (201 Created)
Response (201 Created)
{
"id": "c3d4e5f6-a7b8-9012-cdef-123456789012",
"name": "Benchmark 2024-09-25",
"dataset_id": "a1b2c3d4-e5f6-7890-abcd-ef1234567890",
"user_schema_id": "5e6f7a8b-9c0d-1234-abcd-ef0123456789",
"status": "complete",
"accuracy_overall": 0.943,
"accuracy_by_field": { "invoice_number": 1, "total_amount": 0.92, "vendor_name": 0.91 },
"documents_processed": 50,
"documents_total": 50,
"duration_ms": 2380,
"accuracy_delta": null,
"compared_to_run_id": null,
"created_at": "2024-09-25T12:00:00.000Z",
"completed_at": null,
"links": {
"self": "/v1/quality/benchmarks/c3d4e5f6-a7b8-9012-cdef-123456789012",
"results": "/v1/quality/benchmarks/c3d4e5f6-a7b8-9012-cdef-123456789012/results"
}
}Errors
Error responses