Skip to main content

Ingest a Document

Upload a document (up to 500 MB) into a source with POST /v1/sources/:id/ingest, tag it with batch_id and metadata, and list a source's documents via the API.

The ingest endpoint uploads one document (up to 500 MB) into a source and queues it for extraction on a worker — the response returns immediately with status: "queued" and a document_id to track. POST /v1/sources/:id/ingest and POST /v1/sources/:id/documents are the same operation under two paths; use whichever reads better in your integration. Both require an API key with the extract scope.

Two optional form fields tag the document for correlation with your own systems: batch_id, an opaque grouping key (max 200 characters) stamped on the document, and metadata, a JSON string encoding a flat object of string | number | boolean | null values (at most 50 keys, keys up to 128 characters, string values up to 1,024). Nested objects or arrays are rejected with 400 invalid_metadata. Both tags are echoed back on document reads, so you can attribute platform output to your ERP job or ticket without a lookup table.

Uploads are content-deduplicated: a byte-identical file returns status: "duplicate" instead of queued, carrying both ids — document_id is a thin per-upload link row created for this call, and existing_document_id is the canonical document that already holds the content. The canonical keeps its original tags from its first submission; only the link row carries the new call's batch_id/metadata.

processing_mode chooses the cost/latency trade-off: realtime (default) processes immediately, while batch defers extraction through provider batch APIs at a 50% credit discount with results within 48 hours. Batch submissions accumulate into your workspace's pending batch and are submitted once the batch threshold is met. When batch infrastructure is not configured for your deployment, a batch request silently falls back to realtime — read the echoed processing_mode in the response for the mode that actually applied.

The oversize check runs at the edge: a file beyond your tier's upload cap is rejected 413 before any document record is created, so a failed upload never leaves a half-ingested document behind.
POST/v1/sources/:id/ingest

Multipart form fields

file*fileThe document to ingest (multipart, max 500 MB; your tier cap may be lower).
processing_modestringrealtime or batch. Batch defers extraction at a 50% discount (results within 48 hours); falls back to realtime when batch is not configured. Default: realtime
batch_idstringOptional opaque grouping key (max 200 chars), stamped on the document and echoed on document reads.
metadatastringOptional JSON string of a FLAT object { key: string | number | boolean | null }. Nested values are rejected 400. Limits: ≤50 keys, key ≤128 chars, string value ≤1024 chars.

Request

curl -X POST https://api.talonic.com/v1/sources/a1b2c3d4-e5f6-7890-abcd-ef1234567890/ingest \
  -H "Authorization: Bearer $TALONIC_API_KEY" \
  -F "file=@contract-2024.pdf" \
  -F "batch_id=ERP-2026-08-29-001" \
  -F 'metadata={"source_system":"sap","priority":1}'

Response

Response fields

document_idstringThe created document UUID. On a duplicate: the thin link row created for this upload.
filenamestringThe uploaded filename.
size_bytesinteger | nullUploaded file size in bytes.
statusstring`queued` when accepted for processing, or `duplicate` when the content already exists.
processing_modestringThe mode the document was actually queued under: realtime or batch.
source_idstringThe source it was ingested into. Present on queued responses.
messagestringHuman-readable note, present when status is duplicate.
existing_document_idstring | nullOn a duplicate: the canonical document that already holds this content.
linksobjectRelated resource URLs (document, source). Present on queued responses.

Response (queued)

{
  "document_id": "d4e5f6a7-b8c9-0123-defa-234567890123",
  "filename": "contract-2024.pdf",
  "size_bytes": 204800,
  "status": "queued",
  "processing_mode": "realtime",
  "source_id": "a1b2c3d4-e5f6-7890-abcd-ef1234567890",
  "links": {
    "document": "/v1/documents/d4e5f6a7-b8c9-0123-defa-234567890123",
    "source": "/v1/sources/a1b2c3d4-e5f6-7890-abcd-ef1234567890"
  }
}

Response (duplicate)

{
  "status": "duplicate",
  "message": "File already exists.",
  "document_id": "0f8e7d6c-5b4a-3921-8765-43210fedcba9",
  "existing_document_id": "d4e5f6a7-b8c9-0123-defa-234567890123",
  "filename": "contract-2024.pdf",
  "size_bytes": 204800
}

Errors

Error responses

400missing_document / invalid_metadataNo file was provided, or metadata is not a valid flat JSON object within the caps.
401unauthorizedMissing or invalid API key, or the key lacks the extract scope.
404not_foundNo source with this ID exists for your organization.
413file_too_largeThe file exceeds your plan's maximum upload size (at most 500 MB).
429rate_limitedToo many requests. Retry after the period indicated in the Retry-After header.

List a source's documents

The read counterpart, GET /v1/sources/:id/documents, pages through every document the source has ingested, newest first by default. It uses cursor pagination — follow pagination.next_cursor until has_more is false — and is filtered by Sources-IAM visibility, evaluated as the API key's minting user, so a key sees exactly the documents its creator is admitted to.

GET/v1/sources/:id/documents

Query parameters

limitintegerMaximum number of documents to return (1-100). Default: 20
cursorstringOpaque pagination cursor from a previous response's pagination.next_cursor.
orderstringSort order by creation date: asc or desc. Default: desc

Response

{
  "data": [
    {
      "id": "d4e5f6a7-b8c9-0123-defa-234567890123",
      "filename": "contract-2024.pdf",
      "status": "completed",
      "size_bytes": 204800,
      "type_detected": "Contract",
      "created_at": "2026-08-14T12:30:00.000Z",
      "links": { "self": "/v1/documents/d4e5f6a7-b8c9-0123-defa-234567890123" }
    }
  ],
  "pagination": {
    "total": 234,
    "limit": 20,
    "has_more": true,
    "next_cursor": "ZDRlNWY2YTd8MjAyNi0wOC0xNFQxMjozMDowMC4wMDBa"
  }
}
pagination.total counts the source's visible documents before the cursor slice, so it is stable across pages — safe to render as "N documents" while paging.

Frequently asked questions

How is /v1/sources/:id/ingest different from POST /v1/sources/:id/documents?+
They are the same operation under two paths — same fields, same responses, same scope requirement. Ingest exists as a clearer verb for submitting a single document; pick one and stay consistent.
What happens if I upload a duplicate file?+
Content-identical uploads return status "duplicate" with both ids: existing_document_id is the canonical document that already holds the content, and document_id is a thin link row representing this specific upload. The canonical keeps its original batch_id/metadata; only the link row carries the new call's tags.
What does batch processing_mode do?+
It defers extraction through the provider batch APIs at roughly half the credit cost, delivered within 48 hours instead of seconds. Batch submissions accumulate toward a submission threshold. If batch infrastructure is not configured, the request silently runs realtime — check the echoed processing_mode.
How do I track a document after ingesting it?+
Use the document_id from the response with the documents API (the links.document URL) to check processing status, and GET /v1/sources/:id/documents to watch the whole stream. Documents progress pending → processing → completed (or error).
What file formats are supported?+
Talonic supports 25+ formats including PDF, DOCX, XLSX, CSV, PPTX, MSG, EML, PNG, JPG, HTML, XML, and JSON. See the supported file types documentation for the full list.