Summary
Get the aggregate N-Shot summary for a job run: additional shot count, green/yellow/red agreement breakdown, override count, and the overall agreement rate.
N-Shot measures extraction stability by re-running extraction multiple times ("shots") and comparing the results field by field: when independent shots agree, the value is trustworthy; when they diverge, the field needs attention. Shot 1 is always the run's original extraction. When N-Shot is enabled on a run's validation configuration, the platform executes the additional shots in parallel after the run completes — each shot is a fresh single-pass extraction per document against the run's schema snapshot, deliberately independent of the original pipeline.
Comparison is type-aware: dates and numbers are compared with exact matching after normalization, short text uses fuzzy Jaro-Winkler similarity, and long text is scored by a semantic LLM comparison. Each document-field cell becomes one comparison row with an agreement status and score. The N-Shot endpoints expose these per-cell comparisons for a job run, plus overrides and judge decisions that record which value is correct. All routes are nested under /v1/jobs/runs/{runId}/nshot/....
- Green — every shot produced the same value (unanimous agreement)
- Yellow — more than half of the shots agree on the majority value, but not all
- Red — half or fewer of the shots agree (no reliable majority)
The summary endpoint returns aggregate statistics over every comparison in the run: total comparisons, the green/yellow/red breakdown, how many cells carry an override, and the overall agreement_rate (green comparisons over total, rounded to four decimal places). Use it as the first read after a run finishes — a single number tells you whether a schema change improved or degraded extraction stability before you drill into individual comparisons.
A practical regression workflow: run the same document set through two schema variants with N-Shot enabled, then diff the two summaries. A falling agreement_rate or a growing red count localizes instability introduced by the change. Because the response is a handful of integers, it is cheap to poll and easy to store as a time series alongside your own CI metadata.
/v1/jobs/runs/{runId}/nshot/summaryRequest
curl https://api.talonic.com/v1/jobs/runs/a1b2c3d4-e5f6-7890-abcd-ef1234567890/nshot/summary \
-H "Authorization: Bearer tlnc_your_api_key"Response
Response fields
Response
{
"run_id": "a1b2c3d4-e5f6-7890-abcd-ef1234567890",
"shot_count": 2,
"total_comparisons": 420,
"green": 374,
"yellow": 32,
"red": 14,
"overridden": 6,
"agreement_rate": 0.8905
}Errors
Error responses