Files, Spans & Tags
Stream a document's original file, turn a provenance quote into bounding boxes on the page (scans included), and tag documents with batch_id and metadata.
Three document routes connect structured output back to the original file. GET /v1/documents/{id}/file streams the bytes you uploaded. GET /v1/documents/{id}/span-geometry turns a provenance quote into rectangles on the original page, so a reviewer can see exactly where a value came from, even on a scanned PDF or a photo. PATCH /v1/documents/{id} sets the caller-owned batch_id and metadata tags after upload. Together with the provenance entries that results endpoints return, they let you build an evidence viewer without storing a second copy of every document.
Stream the original file
The file route answers with the stored MIME type (or application/octet-stream), Content-Disposition: inline with the original filename (ASCII plus an RFC 5987 filename* for non-ASCII names), and Accept-Ranges: bytes. A Range request that can be satisfied returns 206 with Content-Range, which lets PDF viewers load large files page by page. Content-Length is sent when the size is known. For a document that was deduplicated against an earlier upload, the route streams the canonical bytes but keeps this upload's own filename. A missing document, one hidden from your key by access rules, or one with no stored bytes is a 404.
/v1/documents/{id}/fileDownload a document, or its first 64 KB
curl -s -o invoice.pdf https://api.talonic.com/v1/documents/c3d4e5f6-a7b8-9012-cdef-123456789012/file \
-H "Authorization: Bearer $TALONIC_API_KEY"
curl -s -o head.bin -H "Range: bytes=0-65535" \
https://api.talonic.com/v1/documents/c3d4e5f6-a7b8-9012-cdef-123456789012/file \
-H "Authorization: Bearer $TALONIC_API_KEY"Locate a quote on the page
span-geometry takes a provenance entry's text (the verbatim quote, passed exactly once as the text query parameter) and an optional 0-based page hint (the entry's page_index), and returns rectangles on the original page. Rectangles are page fractions in [0, 1] with a top-left origin, one per text line: render the page at any resolution and multiply by its width and height. The route uses the document's own text layer when it has one; for scanned, image-only PDFs and image uploads it OCRs the page for word boxes and aligns the quote to them. Results are cached per document.
Gate your highlight on method. pdf_text_layer and ocr_word_boxes are the exact span; section is the enclosing line, a region hint; llm_vision is a model-located box verified against the page's own words; page means the page was found but no rectangle aligned; none means nothing was found and page_index is null. An empty result is an honest miss, so show no highlight rather than a wrong one. This is an interactive, per-click route: a cold page can take several seconds of OCR, one call per organization runs at a time (a concurrent call gets 429 with Retry-After: 5), and quotes longer than 500 characters are truncated before matching.
/v1/documents/{id}/span-geometryQuery parameters
Response
{
"document_id": "c3d4e5f6-a7b8-9012-cdef-123456789012",
"page_index": 0,
"rects": [
{ "left": 0.6120, "top": 0.0815, "width": 0.2210, "height": 0.0142 }
],
"method": "ocr_word_boxes"
}Provenance fields to feed it
Request include=provenance on GET /v1/pipelines/{id}/results or GET /v1/run/{id}/results to get one provenance entry per field. Every entry carries kind (span, assembly_override, auto_adjudication, human, derived, legacy, or matcher_decision), match (exact, approximate, or none, the grade of the quote against the source), confidence (nullable), and path, the field path as an array of keys and indexes. Kinds that carry a quote add text, source_document_id, page_index (0-based), and char_start/char_end, UTF-16 offsets into the document's served markdown. Pass text and page_index straight to span-geometry, and use match to decide how much to trust the highlight.
Tag a document after upload
PATCH /v1/documents/{id} updates the caller-owned tags that POST /v1/run sets at upload. The body must include at least one of batch_id or metadata. A string batch_id sets it (trimmed), null clears it, and a blank string is a 400. metadata is merged into the existing bag: patch keys win, other keys are kept, and null is stored as a value, never treated as a deletion. To retire a key, set it to a value your systems treat as empty. A merged bag that exceeds the metadata caps is rejected whole with a 400. The response is { id, updated: true, batch_id, metadata }, echoing the full merged bag; each tag appears only when set.
Move a document to another batch and add a tag
curl -X PATCH https://api.talonic.com/v1/documents/c3d4e5f6-a7b8-9012-cdef-123456789012 \
-H "Authorization: Bearer $TALONIC_API_KEY" \
-H "Content-Type: application/json" \
-d '{ "batch_id": "ERP-2026-10-06-002", "metadata": { "reviewed_by": "ap-team", "priority": null } }'
# -> { "id": "c3d4…", "updated": true, "batch_id": "ERP-2026-10-06-002",
# "metadata": { "source_system": "sap", "reviewed_by": "ap-team", "priority": null } }null is stored as a value. GET /v1/documents does not filter by batch_id; to read a batch back, keep your own list of document ids from the POST /v1/run response.