Skip to main content

Summary

Get the aggregate N-Shot summary for a job run: additional shot count, green/yellow/red agreement breakdown, override count, and the overall agreement rate.

N-Shot measures extraction stability by re-running extraction multiple times ("shots") and comparing the results field by field: when independent shots agree, the value is trustworthy; when they diverge, the field needs attention. Shot 1 is always the run's original extraction. When N-Shot is enabled on a run's validation configuration, the platform executes the additional shots in parallel after the run completes — each shot is a fresh single-pass extraction per document against the run's schema snapshot, deliberately independent of the original pipeline.

Comparison is type-aware: dates and numbers are compared with exact matching after normalization, short text uses fuzzy Jaro-Winkler similarity, and long text is scored by a semantic LLM comparison. Each document-field cell becomes one comparison row with an agreement status and score. The N-Shot endpoints expose these per-cell comparisons for a job run, plus overrides and judge decisions that record which value is correct. All routes are nested under /v1/jobs/runs/{runId}/nshot/....

  • Green — every shot produced the same value (unanimous agreement)
  • Yellow — more than half of the shots agree on the majority value, but not all
  • Red — half or fewer of the shots agree (no reliable majority)

The summary endpoint returns aggregate statistics over every comparison in the run: total comparisons, the green/yellow/red breakdown, how many cells carry an override, and the overall agreement_rate (green comparisons over total, rounded to four decimal places). Use it as the first read after a run finishes — a single number tells you whether a schema change improved or degraded extraction stability before you drill into individual comparisons.

A practical regression workflow: run the same document set through two schema variants with N-Shot enabled, then diff the two summaries. A falling agreement_rate or a growing red count localizes instability introduced by the change. Because the response is a handful of integers, it is cheap to poll and easy to store as a time series alongside your own CI metadata.

shot_count counts the ADDITIONAL re-extraction shots only — the original extraction is shot 1 and is not included. A run configured for 3 total shots therefore reports shot_count: 2, and each comparison's values array has 3 entries.
GET/v1/jobs/runs/{runId}/nshot/summary

Request

curl https://api.talonic.com/v1/jobs/runs/a1b2c3d4-e5f6-7890-abcd-ef1234567890/nshot/summary \
  -H "Authorization: Bearer tlnc_your_api_key"

Response

Response fields

run_idstringThe job run UUID.
shot_countintegerNumber of additional re-extraction shots recorded for this run. The original extraction (shot 1) is not counted.
total_comparisonsintegerTotal number of document-field comparisons.
greenintegerComparisons where every shot agreed (unanimous).
yellowintegerComparisons where more than half of the shots agree, but not all.
redintegerComparisons where half or fewer of the shots agree.
overriddenintegerNumber of comparisons carrying an override (manual or from an accepted judge decision).
agreement_ratenumber | nullgreen / total_comparisons, rounded to 4 decimal places. Null when the run has no comparisons.

Response

{
  "run_id": "a1b2c3d4-e5f6-7890-abcd-ef1234567890",
  "shot_count": 2,
  "total_comparisons": 420,
  "green": 374,
  "yellow": 32,
  "red": 14,
  "overridden": 6,
  "agreement_rate": 0.8905
}

Errors

Error responses

401unauthorizedMissing or invalid API key.
404not_foundNo job run with this ID exists for your organization.
429rate_limitedDaily request quota for your tier reached. The counter resets at midnight UTC.

Frequently asked questions

How do comparisons come to exist for a run?+
N-Shot is part of the run's validation configuration (enabled with a shot count, optionally alongside the LLM judge). When the run finishes extracting, the platform executes the additional shots in parallel and writes one comparison per document-field cell. Runs without N-Shot enabled report zero comparisons and a null agreement_rate.
What is a good agreement_rate?+
An agreement rate above 0.90 indicates stable extraction. Rates between 0.75-0.90 suggest the schema needs tuning — inspect the yellow and red comparisons to see which fields diverge. Below 0.75 typically indicates structural issues with the schema or inconsistent source documents.
Why does shot_count show one less than I configured?+
shot_count counts only the additional re-extraction shots. The original extraction is shot 1 and lives in the run itself, so a run configured for 3 total shots reports shot_count: 2 while each comparison's values array still has 3 entries.
Does the summary update as I submit judge decisions?+
The overridden count increments with each accepted judge decision and each manual override. The agreement breakdown (green/yellow/red) and agreement_rate reflect the original shot outcomes and never change when overrides are applied — they measure extraction stability, not review progress.