Ingest a Document
Upload a document (up to 500 MB) into a source with POST /v1/sources/:id/ingest, tag it with batch_id and metadata, and list a source's documents via the API.
The ingest endpoint uploads one document (up to 500 MB) into a source and queues it for extraction on a worker — the response returns immediately with status: "queued" and a document_id to track. POST /v1/sources/:id/ingest and POST /v1/sources/:id/documents are the same operation under two paths; use whichever reads better in your integration. Both require an API key with the extract scope.
Two optional form fields tag the document for correlation with your own systems: batch_id, an opaque grouping key (max 200 characters) stamped on the document, and metadata, a JSON string encoding a flat object of string | number | boolean | null values (at most 50 keys, keys up to 128 characters, string values up to 1,024). Nested objects or arrays are rejected with 400 invalid_metadata. Both tags are echoed back on document reads, so you can attribute platform output to your ERP job or ticket without a lookup table.
Uploads are content-deduplicated: a byte-identical file returns status: "duplicate" instead of queued, carrying both ids — document_id is a thin per-upload link row created for this call, and existing_document_id is the canonical document that already holds the content. The canonical keeps its original tags from its first submission; only the link row carries the new call's batch_id/metadata.
processing_mode chooses the cost/latency trade-off: realtime (default) processes immediately, while batch defers extraction through provider batch APIs at a 50% credit discount with results within 48 hours. Batch submissions accumulate into your workspace's pending batch and are submitted once the batch threshold is met. When batch infrastructure is not configured for your deployment, a batch request silently falls back to realtime — read the echoed processing_mode in the response for the mode that actually applied.
/v1/sources/:id/ingestMultipart form fields
realtimeRequest
curl -X POST https://api.talonic.com/v1/sources/a1b2c3d4-e5f6-7890-abcd-ef1234567890/ingest \
-H "Authorization: Bearer $TALONIC_API_KEY" \
-F "file=@contract-2024.pdf" \
-F "batch_id=ERP-2026-08-29-001" \
-F 'metadata={"source_system":"sap","priority":1}'Response
Response fields
Response (queued)
{
"document_id": "d4e5f6a7-b8c9-0123-defa-234567890123",
"filename": "contract-2024.pdf",
"size_bytes": 204800,
"status": "queued",
"processing_mode": "realtime",
"source_id": "a1b2c3d4-e5f6-7890-abcd-ef1234567890",
"links": {
"document": "/v1/documents/d4e5f6a7-b8c9-0123-defa-234567890123",
"source": "/v1/sources/a1b2c3d4-e5f6-7890-abcd-ef1234567890"
}
}Response (duplicate)
{
"status": "duplicate",
"message": "File already exists.",
"document_id": "0f8e7d6c-5b4a-3921-8765-43210fedcba9",
"existing_document_id": "d4e5f6a7-b8c9-0123-defa-234567890123",
"filename": "contract-2024.pdf",
"size_bytes": 204800
}Errors
Error responses
List a source's documents
The read counterpart, GET /v1/sources/:id/documents, pages through every document the source has ingested, newest first by default. It uses cursor pagination — follow pagination.next_cursor until has_more is false — and is filtered by Sources-IAM visibility, evaluated as the API key's minting user, so a key sees exactly the documents its creator is admitted to.
/v1/sources/:id/documentsQuery parameters
20descResponse
{
"data": [
{
"id": "d4e5f6a7-b8c9-0123-defa-234567890123",
"filename": "contract-2024.pdf",
"status": "completed",
"size_bytes": 204800,
"type_detected": "Contract",
"created_at": "2026-08-14T12:30:00.000Z",
"links": { "self": "/v1/documents/d4e5f6a7-b8c9-0123-defa-234567890123" }
}
],
"pagination": {
"total": 234,
"limit": 20,
"has_more": true,
"next_cursor": "ZDRlNWY2YTd8MjAyNi0wOC0xNFQxMjozMDowMC4wMDBa"
}
}