Export Filter Results
Pull every document matching filter conditions out of Talonic: page through POST /v1/documents/filter with a stable sort and project field values to CSV/JSONL.
The document filter is pageable to the full result set, which makes it the building block for exporting a filtered slice of your corpus: page through POST /v1/documents/filter with your conditions, project the rows you need, and write them to CSV, JSONL, or your warehouse. limit goes up to 500 per page, and total tells you up front how many pages the export needs.
Keep the export deterministic by fixing a sort — sorting by a date field descending is the usual choice — so pages do not shift under you while documents are being ingested concurrently. Include every field you want to project as a condition (an is_not_empty condition works as a pure projection trigger for fields you do not want to constrain), because field_values inlines only the fields the conditions reference.
For counts-only reporting, skip the paging entirely and use the [document counts pattern](document-counts). For schema-shaped tabular exports of pipeline output — rows with named columns rather than document-level field maps — prefer the pipeline results endpoint or a delivery binding: this filter export is the right tool when the unit of interest is the document, filtered by what was extracted from it.
The same mechanics give you incremental exports: filter on a date field with gte set to your last run's high-water mark (or between the last run and now) and only the documents that arrived since are returned. Persist the newest value you exported as the next run's lower bound, and the export becomes an idempotent sync job instead of a full re-pull — the count pattern with the same condition tells you up front whether the run has anything to do.
Export all matching documents as JSONL
PAGE=1
while :; do
RESP=$(curl -s -X POST https://api.talonic.com/v1/documents/filter \
-H "Authorization: Bearer $TALONIC_API_KEY" \
-H "Content-Type: application/json" \
-d '{
"conditions": [
{ "fieldId": "6f1e9a2b-4c3d-4e5f-8a7b-9c0d1e2f3a4b", "operator": "eq", "value": "Acme Corp" },
{ "fieldId": "8a2b3c4d-5e6f-4a1b-9c8d-7e6f5a4b3c2d", "operator": "is_not_empty" }
],
"sort": { "fieldId": "8a2b3c4d-5e6f-4a1b-9c8d-7e6f5a4b3c2d", "direction": "desc" },
"page": '"$PAGE"',
"limit": 500
}')
echo "$RESP" | jq -c '.data[]' >> export.jsonl
COUNT=$(echo "$RESP" | jq '.data | length')
[ "$COUNT" -lt 500 ] && break
PAGE=$((PAGE + 1))
done
wc -l export.jsonlProject to CSV (id, filename, tested value)
jq -r '[.id, .filename,
(.field_values["6f1e9a2b-4c3d-4e5f-8a7b-9c0d1e2f3a4b"].value_text // "")]
| @csv' export.jsonl > export.csvPractical notes
- Page size: 500 is the maximum
limit; a larger value is clamped, not rejected. - Termination: stop when a page returns fewer than
limitrows;totalgives the expected row count for a consistency check. - Projection:
field_valuesentries are typed — readvalue_text,value_number,value_date, orvalue_booleanper the field's data type. - Full records: for fields beyond your conditions, follow each
idinto the documents/extractions endpoints, or export a pipeline's results instead.