Skip to main content

List Benchmark Datasets

List every benchmark dataset in your workspace with GET /v1/quality/ground-truth. Each dataset holds verified field values that score extraction accuracy.

A benchmark dataset is a collection of manually verified field values that serves as the gold standard for benchmarking extraction accuracy. These datasets and benchmark runs together make up Benchmarks (the /v1/quality namespace). GET /v1/quality/ground-truth lists every dataset in your workspace so you can pick one to benchmark against. This is offline accuracy measurement, not the inline validation checks that gate individual results before delivery.

Do not confuse this resource with the golden samples under [/v1/validation/ground-truth](list-ground-truth): both hold verified values, but they feed different engines. A /v1/quality dataset is schema-scoped, stores one expected_data object per document, and drives repeatable benchmark runs with per-field accuracy scores. A /v1/validation golden sample stores per-field expected values and scores one specific Structuring Run.

Use this endpoint to see all available datasets before creating a benchmark run. A typical workflow is to list datasets, select the one covering the document type you want to evaluate, then pass its id to POST /v1/quality/benchmarks to start a run.

Each dataset includes a name, optional description, user_schema_id (the schema it is scoped to), a document_count, and a links.self URL for the detail endpoint. Datasets are returned in descending creation order with cursor-based pagination: the cursor is keyset-based over (created_at, id), so pages stay stable even while new datasets are being created between requests.

Treat document_count as advisory rather than authoritative: it is maintained by the platform's ground-truth workflow and does not increment when you add entries through POST /v1/quality/ground-truth/:datasetId/entries. To count entries reliably, list them via the [entries endpoint](quality-entries) or read the samples array on the [dataset detail](get-quality-dataset). Benchmark runs are unaffected — documents_total is computed from the live entry count at run creation.

Create separate datasets for different document types or schema versions to track accuracy independently. Pair with the benchmark endpoints to measure extraction accuracy over time: run benchmarks after schema or extraction changes to detect regressions.

Every ground truth dataset is scoped to one schema via user_schema_id, so its field names line up with the extraction output the benchmark compares against.
  • Each dataset contains verified entries mapping documents to expected field values
  • Datasets can be scoped to a specific user schema via user_schema_id
  • Use datasets as inputs to benchmark runs for per-field accuracy measurement
GET/v1/quality/ground-truth

Query parameters

limitintegerMaximum number of results to return (1–100). Default: 20
cursorstringPagination cursor from a previous response.
orderstringSort direction by created_at. Default: desc

Request

curl "https://api.talonic.com/v1/quality/ground-truth?limit=20&order=desc" \
  -H "Authorization: Bearer tlnc_..."

Response

Response fields

dataarrayArray of ground truth dataset objects.
data[].idstringDataset UUID.
data[].namestringDataset name.
data[].descriptionstring | nullOptional description.
data[].user_schema_idstring | nullAssociated user schema ID, if any.
data[].document_countintegerEntry counter maintained by the platform workflow. Not incremented by API-added entries — use the entries endpoint for an authoritative count.
data[].created_atstringISO 8601 creation timestamp.
data[].links.selfstringURL to this dataset.
pagination.totalintegerTotal number of datasets.
pagination.limitintegerMaximum results per page.
pagination.has_morebooleanWhether more results exist beyond this page.
pagination.next_cursorstring | nullCursor to fetch the next page.

Response

{
  "data": [
    {
      "id": "a1b2c3d4-e5f6-7890-abcd-ef1234567890",
      "name": "Invoice Accuracy Set",
      "description": "Manually verified invoices for Q3 2024",
      "user_schema_id": null,
      "document_count": 50,
      "created_at": "2024-09-01T10:00:00.000Z",
      "links": {
        "self": "/v1/quality/ground-truth/a1b2c3d4-e5f6-7890-abcd-ef1234567890"
      }
    }
  ],
  "pagination": {
    "total": 3,
    "limit": 20,
    "has_more": false,
    "next_cursor": null
  }
}

Errors

Error responses

401unauthorizedMissing or invalid API key.
429rate_limitedToo many requests. Retry after the period indicated in the Retry-After header.

Frequently asked questions

How many ground truth datasets can I create?+
There is no hard limit on the number of datasets. Create separate datasets for different document types or schema versions to track accuracy independently.
What is the recommended number of entries per dataset?+
For statistically meaningful accuracy scores, aim for at least 30-50 entries per dataset. Smaller datasets may produce volatile accuracy metrics.
What is the difference between a ground truth dataset and validation Ground Truth?+
A ground truth dataset under `/v1/quality` stores per-document `expected_data` objects and feeds benchmark runs with per-field accuracy scores. Validation Ground Truth under `/v1/validation` (the API resource is still named golden sample) stores per-field expected values and feeds validation runs that score one specific Job.