Skip to main content

Create Matching Config

Create a matching configuration with field mappings, match types (exact, fuzzy_string, date_range, numeric_range), and per-field weights that sum to 1.0.

Create a matching configuration that defines how documents are compared against a reference dataset. Each field mapping specifies an extracted_field (from extracted documents), a reference_field (a column in the reference data), a match_type, and a relative weight.

The typical workflow is: upload reference data via POST /v1/matching/reference-data, create a config with field mappings, then trigger a run via POST /v1/matching/configs/:id/run. For complex datasets, use POST /v1/matching/strategies/generate first to get AI-recommended mappings and weights.

The response returns the config with the saved field_mappings, threshold (defaults to 0.85), and links.runs URL for triggering runs. The reference_data_id is fixed at creation — to match against a different dataset, create a new config.

Choose match types carefully: use exact for standardized codes and IDs, fuzzy_string for names with potential typos, date_range for dates with tolerance, and numeric_range for amounts with rounding differences. Weights must sum to 1.0 — fields with higher weights have more influence on the overall confidence score.

Field weights should sum to 1.0. The overall confidence score for a match is the weighted sum of per-field scores. Use the generate strategy endpoint to get AI-recommended mappings if you are unsure which fields and weights to use.
  • exact — case-insensitive string equality. Best for codes, IDs, and standardized values.
  • fuzzy_string — Levenshtein similarity with a per-mapping threshold (default 0.8). Handles name variations and minor typos.
  • date_range — date proximity within a configurable tolerance_days window.
  • numeric_range — numeric proximity within a configurable tolerance_pct. Handles rounding differences.
  • best_of — scores several alternatives (each its own field pair and match type) and keeps the best. sequence compares list fields element-wise via sequence_config.
POST/v1/matching/configs

Body parameters

name*stringHuman-readable name for the config.
reference_data_id*stringID of the uploaded reference dataset to match against.
field_mappings*arrayField mapping array (at least one entry). Each entry has `extracted_field`, `reference_field`, `match_type`, and `weight`.
field_mappings[].extracted_field*stringField name in the extracted document.
field_mappings[].reference_field*stringColumn name in the reference dataset.
field_mappings[].match_type*stringComparison type: `exact`, `fuzzy_string`, `date_range`, `numeric_range`, `best_of`, or `sequence`.
field_mappings[].weight*numberRelative weight (0–1).
field_mappings[].thresholdnumberfuzzy_string similarity threshold. Defaults to 0.8.
field_mappings[].tolerance_daysnumberdate_range tolerance window in days.
field_mappings[].tolerance_pctnumbernumeric_range tolerance as a percentage.
thresholdnumberAuto-accept confidence threshold (0–1). Defaults to 0.85. Default: 0.85
target_typestringTarget scope: `run` (default), `schema`, or `document_filter`. Default: run
target_valueobjectTarget identifier payload (e.g. run_id, schema_id).

Request body

{
  "name": "Vendor Invoice Match",
  "reference_data_id": "b2c3d4e5-f6a7-8901-bcde-f12345678901",
  "field_mappings": [
    { "extracted_field": "vendor_name", "reference_field": "name", "match_type": "fuzzy_string", "weight": 0.4 },
    { "extracted_field": "invoice_date", "reference_field": "date", "match_type": "date_range", "tolerance_days": 5, "weight": 0.3 },
    { "extracted_field": "amount", "reference_field": "total", "match_type": "numeric_range", "tolerance_pct": 1, "weight": 0.3 }
  ],
  "threshold": 0.85
}

Response

Response fields (201 Created)

idstringConfiguration UUID.
namestringConfig name.
reference_data_idstringReference dataset ID.
target_typestringTargeting mode.
target_valueobjectTarget-specific values.
field_mappingsarraySaved field mapping array.
thresholdnumberConfidence threshold.
created_atstringISO 8601 creation timestamp.
updated_atstringISO 8601 last update timestamp.
links.selfstringURL to this config.
links.runsstringURL to trigger a run.

Response (201 Created)

{
  "id": "a1b2c3d4-e5f6-7890-abcd-ef1234567890",
  "name": "Vendor Invoice Match",
  "reference_data_id": "b2c3d4e5-f6a7-8901-bcde-f12345678901",
  "target_type": "run",
  "target_value": {},
  "field_mappings": [
    { "extracted_field": "vendor_name", "reference_field": "name", "match_type": "fuzzy_string", "weight": 0.4 },
    { "extracted_field": "invoice_date", "reference_field": "date", "match_type": "date_range", "tolerance_days": 5, "weight": 0.3 },
    { "extracted_field": "amount", "reference_field": "total", "match_type": "numeric_range", "tolerance_pct": 1, "weight": 0.3 }
  ],
  "threshold": 0.85,
  "created_at": "2024-10-01T08:00:00.000Z",
  "updated_at": "2024-10-01T08:00:00.000Z",
  "links": {
    "self": "/v1/matching/configs/a1b2c3d4-e5f6-7890-abcd-ef1234567890",
    "runs": "/v1/matching/configs/a1b2c3d4-e5f6-7890-abcd-ef1234567890/run"
  }
}

Errors

Error responses

400validation_errorMissing required fields (name, reference_data_id, field_mappings), or invalid field mapping format.
401unauthorizedMissing or invalid API key.
404not_foundThe specified reference_data_id does not exist for your workspace.
429rate_limitedToo many requests. Retry after the period indicated in the Retry-After header.

Frequently asked questions

What match types are available for field matching?+
Four core types: exact (case-insensitive equality), fuzzy_string (Levenshtein similarity with a per-mapping threshold, default 0.8), date_range (date proximity within tolerance_days), and numeric_range (numeric proximity within tolerance_pct). Two composite types — best_of and sequence — cover alternatives and list fields.
Do field weights need to sum to exactly 1.0?+
Weights should sum to 1.0 for meaningful confidence scores. If they do not sum to 1.0, the system normalizes them internally, but explicitly setting weights to sum to 1.0 gives you predictable confidence values.
Can I use the same reference dataset column in multiple mappings?+
Yes. A single reference_field can appear in multiple field mappings with different extracted fields and match types, which is useful when multiple document fields might correspond to the same reference column.