Skip to main content

Export Filter Results

Pull every document matching filter conditions out of Talonic: page through POST /v1/documents/filter with a stable sort and project field values to CSV/JSONL.

The document filter is pageable to the full result set, which makes it the building block for exporting a filtered slice of your corpus: page through POST /v1/documents/filter with your conditions, project the rows you need, and write them to CSV, JSONL, or your warehouse. limit goes up to 500 per page, and total tells you up front how many pages the export needs.

Keep the export deterministic by fixing a sort — sorting by a date field descending is the usual choice — so pages do not shift under you while documents are being ingested concurrently. Include every field you want to project as a condition (an is_not_empty condition works as a pure projection trigger for fields you do not want to constrain), because field_values inlines only the fields the conditions reference.

For counts-only reporting, skip the paging entirely and use the [document counts pattern](document-counts). For schema-shaped tabular exports of pipeline output — rows with named columns rather than document-level field maps — prefer the pipeline results endpoint or a delivery binding: this filter export is the right tool when the unit of interest is the document, filtered by what was extracted from it.

The same mechanics give you incremental exports: filter on a date field with gte set to your last run's high-water mark (or between the last run and now) and only the documents that arrived since are returned. Persist the newest value you exported as the next run's lower bound, and the export becomes an idempotent sync job instead of a full re-pull — the count pattern with the same condition tells you up front whether the run has anything to do.

The filter is IAM-visibility-scoped to the API key's minting user, so an export contains exactly what that key is admitted to see — totals and pages are consistent with each other, but a different key can legitimately export a different set.

Export all matching documents as JSONL

PAGE=1
while :; do
  RESP=$(curl -s -X POST https://api.talonic.com/v1/documents/filter \
    -H "Authorization: Bearer $TALONIC_API_KEY" \
    -H "Content-Type: application/json" \
    -d '{
      "conditions": [
        { "fieldId": "6f1e9a2b-4c3d-4e5f-8a7b-9c0d1e2f3a4b", "operator": "eq", "value": "Acme Corp" },
        { "fieldId": "8a2b3c4d-5e6f-4a1b-9c8d-7e6f5a4b3c2d", "operator": "is_not_empty" }
      ],
      "sort": { "fieldId": "8a2b3c4d-5e6f-4a1b-9c8d-7e6f5a4b3c2d", "direction": "desc" },
      "page": '"$PAGE"',
      "limit": 500
    }')
  echo "$RESP" | jq -c '.data[]' >> export.jsonl
  COUNT=$(echo "$RESP" | jq '.data | length')
  [ "$COUNT" -lt 500 ] && break
  PAGE=$((PAGE + 1))
done
wc -l export.jsonl

Project to CSV (id, filename, tested value)

jq -r '[.id, .filename,
  (.field_values["6f1e9a2b-4c3d-4e5f-8a7b-9c0d1e2f3a4b"].value_text // "")]
  | @csv' export.jsonl > export.csv

Practical notes

  • Page size: 500 is the maximum limit; a larger value is clamped, not rejected.
  • Termination: stop when a page returns fewer than limit rows; total gives the expected row count for a consistency check.
  • Projection: field_values entries are typed — read value_text, value_number, value_date, or value_boolean per the field's data type.
  • Full records: for fields beyond your conditions, follow each id into the documents/extractions endpoints, or export a pipeline's results instead.

Frequently asked questions

Is there a dedicated export endpoint for filter results?+
No — the filter itself is the export surface: page through POST /v1/documents/filter at limit 500 and project the rows. Dedicated tabular exports exist for pipeline output (pipeline results, delivery bindings, data product exports); use those when you want schema-shaped rows rather than filtered documents.
How do I keep a long export consistent while documents are being ingested?+
Fix a sort on a stable field (a date field descending) so new arrivals land on early pages you have already passed rather than reshuffling later ones, and compare your exported row count against the response total when done.
How do I export field values for fields I am not filtering on?+
Add an is_not_empty condition per extra field — it constrains only to documents that have the field and pulls its value into field_values. For genuinely optional fields, fetch the full record per document via the documents API instead.