Skip to main content
Bridgeway

BRIDGEWAY · LOGISTICS · UNITED STATES · CASE IN PRODUCTION

Two thousand three hundred loads a day.Four documents each.One person opening every one of them.

Bridgeway is a freight brokerage in Pittsburgh with an asset based trucking arm of its own, operating across three locations.

The software it had returned two things per document: what type it is, and which load it belongs to. We returned thirty to forty structured fields, matched each document against the seven to ten thousand loads open at that moment, and abstained when we were not sure. Then we found out the hardest part was not reading the documents. It was building something true enough to measure against.

batch received · 3 documentstalonic · match
  • BOL

    load
    U6191364
    rate
    $1,842.50
    carrier
    MERIDIAN LINES
    consignee
    FAYETTEVILLE, GA

    8,4124061

    MATCHED · AUTOMATErate matched to the penny · 4 of 4 high weight fields agree
  • CARRIER INVOICE

    load
    none found
    rate
    $2,310.00
    carrier
    NORTHFIELD CARTAGE
    ref
    PO 88-41207

    8,4124062

    NO CONFIDENT MATCH · BILLER QUEUEno load number on the document · reference matched two open loads · abstained
  • RATE CONFIRMATION

    load
    K2288104
    load
    T9034551
    rate
    $1,190.00
    batch
    2 load numbers

    8,4124020

    REFUSED · SAFETY RULEtwo load numbers in one batch · never automated

matching against 8,412 open loads

Illustrative trace. Load numbers, carriers and rates are invented. On your documents the decisions come from your loads and your rules. The pool narrows by embedding nearest neighbours, then by weighted scoring, then by an exact rate check, at about two seconds per thousand loads.

~550,000 LOADS BILLED A YEAR5 TO 5.5 MILLION PAGES7,000 TO 10,000 LOADS OPEN AT ANY MOMENT23 AGENT BRANDSA 250 GB TMS DATABASE WITH TABLES BACK TO 2003

Loads billed, not documents. The page count is the volume the engagement was scoped on. The transport management database is the customer’s own.

Bridgeway

Logistics · USA

Industry
Freight logistics
Function
Billing operations
Solution
Matching and auto-billing

IN PRODUCTION · AUTO BILLING IN SHADOW RUN SINCE 12 AUGUST 2026

96% vs 78%

same documents, same labels, same scoring

Note: The same hand labelled documents, two fields, both systems scored the same way. Replacing a six figure incumbent.

THE PLATFORM RUN

Five stages, the same five on every case page. The figures under them are this engagement’s own.

  1. READ ONCE

    60,000+

    documents read in production

    The existing paperwork stream is duplicated to Talonic rather than rerouted. Four to five documents arrive for every load.

    As of 13 August 2026. Stated by both sides.

  2. STRUCTURE

    2 → 30 to 40

    values per document

    The incumbent returned two: document type and load number. Talonic returns twenty named fields, and the two benchmark fields are the first two of them.

    Thirty to forty values per document, tuned deliberately: past a certain field count accuracy falls.

  3. RESOLVE

    2 sec

    per 1,000 open loads

    Every carrier document is scored against the seven to ten thousand loads open at that moment, then narrowed to one load or to none.

    Matching speed measured April 2026 on the embedding matcher.

  4. DELIVER

    RUNNING

    auto bill decision emitted per load, in shadow

    Every load still goes to a biller. The auto bill decision is emitted beside it and the audit table records what a human changed.

    Auto billing has run in shadow alongside the team since 12 August 2026. Coverage and precision are in section 05.

  5. QUERY FOREVER

    ONCE

    four surfaces read the record, none opens the document again

    The biller queue, the load record, the customer invoice and the automation decision all read the registry.

    The four surfaces are the four named above. ONCE describes the read.

The fifth stage prints ONCE on purpose: it is the claim every page in this family exists to receipt.

INGEST ONCE. QUERY FOREVER.

01 / the ruler came first

The ruler came first.

The software it had returned two things per document. We returned thirty to forty structured fields, and abstained when we were not sure. The hardest part was building something true enough to measure against.

On the same hand labelled documents, scored the same way, Talonic read the document type and the load number correctly 96 percent of the time. The legacy OCR system read them correctly 78 percent of the time.

Before we could beat the incumbent we had to build the ruler. When we tried to score ourselves, Bridgeway’s own records turned out to be about 92 percent right, so we built ground truth three times over, and then the model started auditing the labels.

Where the system cannot decide it says so instead of guessing, and the abstentions are printed in the account rather than folded into the accuracy figure.

Auto billing has run in shadow alongside the team since 12 August 2026, producing the same decisions in parallel and being compared, before it decides alone.

96

TALONIC

78

LEGACY OCR

percent of documents with both fields correct

April 2026. Hand labelled documents, two fields: document type and load number. Both systems scored against the same labels, the same way. The customer’s own prior analysis produced the same figure for the incumbent.

Talonic’s own measurement against ground truth Talonic built, taken before the incumbent cutover.

The 92 percent is a different figure: the accuracy of the customer’s own labels, measured against the ground truth Talonic hand built, on the same two fields. It is why those records could not be used as the benchmark.

Everything below this band is the receipt for it: the work a biller does today, the head to head, the ruler it was scored with, the decision to abstain, and the shadow run’s own numbers.

02 / the work

The work was never reading the document. It was checking it.

A biller opens each document in a side panel and hunts field by field: find the weight, find the rate, find the load number, compare each one against the system, decide. Then the next document, then the next load.

The rules that govern whether a load can be billed are not written down anywhere. They live in the billing team’s heads and they exist in no system. And the founding insight of the engagement is Bridgeway’s own: what a biller checks in order to release a load, and what the customer needs in order to pay, are two different lists. Nobody had written the second one down. That gap is the whole opportunity.

What a biller does today: three values hunted on a document and typed into a form in a different orderTHE DOCUMENT, IN A SIDE PANELTHE SYSTEM, IN THE OTHER HALF OF THE SCREENshipperconsigneecarrierload numberweightrateweightrateload number× 4 to 5 documents · × 2,300 loads a day
The order the values appear on the document is not the order the form asks for them. Four to five documents per load, about 2,300 loads a day.
Lone semi truck silhouetted against an interstate sunsetload_number · U6191364Stand in imagery, not a photograph of the customer.Stand in imagery. Not the customer.

03 / the contest

Same documents. Scored the same way.

The incumbent was a legacy indexing system built on commodity OCR and entity recognition, returning exactly two things per document. In Bridgeway’s own operating experience that ran at roughly 85 percent on document type, roughly 85 percent on the load number, and roughly 75 percent with both correct at once.

In April 2026 we hand labelled a set of documents and scored both systems against the same labels, the same way. The customer had run their own analysis before ours and their figure matched.

April 2026. Hand labelled documents, two fields: document type and load number. Both systems scored against the same labels, the same way.

WHAT WE FOUND AFTERWARDS, AND TOLD THEM

The incumbent could take a single load number found at high confidence somewhere in a batch and apply it to every document in that batch, which lifts the score on batched paperwork without reading anything new. We raised it rather than bank it.

The customer confirmed the same post processing could be applied to our output too. The comparison got harder for us and more useful for them.

Two separate measurements, shown separately: the hand labelled benchmark, and the incumbent in operationTHE SAME DOCUMENTS, APRIL 2026Talonic96Legacy OCR78050100percent of documents with both fields correctTHE INCUMBENT IN OPERATION, AS DESCRIBED BYTHE CUSTOMERdocument type~85load number~85both together~75050100different population, different measurement
April 2026. Hand labelled documents, two fields: document type and load number. Both systems scored against the same labels the same way. The customer’s own prior analysis produced the same figure for the incumbent.

The incumbent’s own operating accuracy as described by the customer: document type alone, load number alone, and both correct together. A different measurement from the hand labelled benchmark, on a different population.

04 / the ruler

You cannot score a system against records that are 92 percent right.

01

We asked for ground truth. There was not any.

Bridgeway’s own load numbers and document types were the obvious yardstick, and they turned out to be about 92 to 93 percent accurate, with errors that were systematic rather than random. Customer reference numbers recorded in the load number field. Documents deliberately labelled as something they were not, because a label was the only way to make a document route where a biller needed it to go. The customer told us plainly to treat their own labels as signals, not as truth.

Measured against ground truth Talonic hand built, on the same two fields as the contest above.

02

So we built the ruler. Three times.

Fifty documents, three people, one full day, in March. A set of documents hand labelled on two fields in April, which is the benchmark the contest above was scored on. And in July, one shared versioned benchmark, one version in one place, after we stopped letting it exist in five copies at once, because a benchmark that exists in five versions is not a benchmark.

03

Then the machine started auditing the labels.

On the thirteenth run against our own ground truth we were at 93 percent, and when we asked why not higher, the answer was that the remaining gap was the ground truth, not the model. It listed what was wrong with our labels. The rule we took from it: where several independent models agree on a value that disagrees with the label, the label is probably the thing that is wrong.

04

That cuts both ways and we are going to say so.

The figures in the low 90s rest on labels a human wrote. The figures in the high 90s rest partly on labels a language model produced. They are not the same kind of number and we do not present them as one. The last gate we set ourselves was a sample nobody had been allowed to iterate against, and on several of those documents the system refused a label the humans had assigned. The billing team, asked to adjudicate, sided with the system.

SAID IN AN INTERNAL REVIEW

Anytime you have these kind of questions, you have to benchmark. I’ve met with many people who talk about it theoretically. In the end, they all come down to: but now I need to empirically test.

Talonic engineering, internal, 17 July 2026
Ground truth built three times, and the run that audited the labelsMARCH50 documents · 3 people · 1 dayAPRILhand labelled set · 2 fieldsthis is the benchmarkin section 03JULYone shared versioned benchmark · one channel · one versionRUN 13the model disagrees with the labels · the labels were wrongthe audit points back up the ladder
Fifty documents with three people in March, a hand labelled set on two fields in April, and a shared versioned benchmark in July.

05 / the decision

Coverage and precision. Never blended into one number.

Most document systems are graded on one number, so most document systems guess. A guess that is wrong costs more than a blank, because a biller who has caught one wrong value now has to check every value, and you have bought nothing.

So the system declines. Where confidence is low it abstains, and we report two numbers instead of one: how much of the volume it answered on, and how often it was right when it answered. When it declines, nothing breaks. The document routes to the biller queue it was already going to. The floor of this system is the process that exists today.

As of 13 August 2026, on about two thousand production documents, matching covered 75 to 78 percent of documents at about 95 percent precision, which is about 71 percent end to end.

Talonic’s own measurement, 13 August 2026, on about two thousand production documents, against ground truth Talonic built, taken before the incumbent cutover.

THE THREE RULES UNDER THE DECISION

The rate matches to the penny, or it is not a match.

Names and addresses can be fuzzy. Weight is ignored on purpose: the number written on a carrier’s paperwork is not a measurement, and a system that insists on it rejects correct matches all day.

Field weight is earned from the data, not assigned by hand.

Each field’s weight is how discriminating a match on it actually is, computed across about twenty fields. If a column does not hold unique values, a good match on it scores zero.

Two load numbers in one batch is never automated.

Two load numbers of Bridgeway’s format in one batch means the batch is not automatable. No scoring, no confidence, no cleverness. The failure it prevents is one customer receiving another customer’s paperwork, and that failure is not worth any amount of coverage.

Production documents split into answered and correct, answered and wrong, and declinedprecision ~95% of what it answeredcoverage 75 to 78%share of production volumeanswered, correct · ~72answered, wrong · ~4declined · ~24~10% of the corpus is structurally unextractable ·nothing reads what is not there
Talonic’s own measurement, 13 August 2026, on about two thousand production documents, against ground truth Talonic built, taken before the incumbent cutover. The unextractable tail is an internal estimate of the same date.

About one document in ten cannot be extracted by anything.

Photographs of the side of a truck. Half a page, taken at an angle, of a document whose other half was never captured. There is no model that reads what is not there.

That tail caps achievable automation somewhere near 90 percent, and any vendor quoting you a higher ceiling on freight paperwork has not looked at the corpus.

Internal estimate, 13 August 2026.

06 / the shadow run

Nothing was rerouted. The stream was duplicated.

Paperwork arrives by the load documents inbox or by a per load QR portal. The existing indexing system keeps running, untouched. A second forwarder duplicates the same stream to Talonic.

Talonic receives the document with a batch id and a webhook url, returns a job id synchronously, then posts back per document with the type, the load number, the extracted fields, a duplicate flag, a confidence and an automation decision. The customer stores it and attaches the document to the load.

The ingest topology: the existing stream is duplicated to Talonic, not reroutedcarrier andcustomer paperworkload docs inbox /QR portalthe stream forks hereUNCHANGEDexisting indexingsystemthe process as itruns todayDUPLICATED, NOT REROUTEDPOST document + batchid + webhook urlTalonic: type, loadid, 30 to 40 fieldsmatch against the readonly snapshotwebhook back: type, load,fields, duplicate flag,confidence, automationdecisionattached to the loadautomation true → auto invoiceautomation false → billerqueueaudit table: every field a billerchanges
The upper branch is the process as it runs today and it is not touched. The lower branch is the duplicate. The floor of the system is the upper branch.

FIVE DESIGN DECISIONS, AND WHY

Webhooks, not polling.

At thousands of documents an hour, nobody wants their servers polled every two seconds.

A read only snapshot.

No reference data is passed in. The load row always exists before the paperwork arrives, because there is at least a day between booking and shipping, so a periodically refreshed snapshot is safe and nothing needs live access to the system of record.

One shot extraction at 150 DPI.

The raw PDF goes to a visual language model against the production extraction schema. Accuracy peaks at 150 DPI and falls above it, because raising it adds pixels rather than information. OCR runs in parallel and is deliberately non blocking, which saves 6 to 9 seconds per document.

Arrays, not nested objects.

Multi value fields come back as arrays, and sums are computed in post processing rather than by the model, because models are unreliable at arithmetic.

Never mutate a live spec.

Mid delivery a required schema change could not be made safely in place, so the running spec was frozen and the new one built alongside it, and it cleared full benchmarking and load testing again before anything switched over. The swap is one step in the ingest path, with an instant rollback.

Keep production unfazed.

WHAT IS ACTUALLY RUNNING

25,000

pages · ~4,000 runs

Customer side count, 30 July 2026.

2,000

documents per minute

Internal load test, 23 July 2026. About 30 seconds average per document, peak batch under ten minutes.

5 hours

~7,000 documents · 0 errors

Internal soak test, 23 July 2026.

FIELD STABILITY

50 of 57 fields returned the same value across six runs of the same documents. The seven that moved were arrays plus three fields that had just been introduced.

Internal, 17 July 2026.

07 / the loop

If the biller changed nothing, the machine could have billed it.

This was the customer’s design, and it is the smartest deployment idea in the engagement. The system emits an automation decision. The load goes through manual billing anyway. An audit table records every field a biller changes. A load that went through untouched is a load that could have been billed automatically, and that is a measurement rather than an opinion.

Cohorts where billers consistently change nothing get flipped to automatic, one slice at a time. Where they do change something, the change itself is the finding: the address had to be corrected before the load could go out, so was that address actually required, and if it was, it goes into the schema. The undocumented billing rules get read out of the audit log rather than out of interviews. At 2,300 loads a day, the dataset builds itself in days.

THE AUTOMATION JOURNEY

On one round of loads the automatable share went from 12 percent to 30 percent to 76 percent.

Both jumps have a cause and neither is a model improvement: the first was filtering to the loads that had actually reached the billing stage, the second was fixing a prefix bug in a reference field on the customer’s side.

Measured on the loads supplied for that round in May 2026, roughly 150 of them.

Two discrete corrections raised the automatable share on one round of loads12%30%76%filtered to loads that hadreached the billing stagea reference field prefixbug fixed on the customersidethe extractable ceiling, section 050100
Measured on the loads supplied for that round in May 2026, roughly 150 of them. Both jumps have a stated cause and neither is a model improvement.

08 / next

Send us the documents you think will break it.

Send a representative sample of your operational paperwork. We run it through Talonic and send back your extracted data within five business days, with field level coverage, confidence and provenance on every value. Send the photographs of the side of a truck too. We would rather show you where the ceiling is than average it away.

SIX CASE STUDIES: INDUSTRIAL ENERGY · GETEC / PHARMA WHOLESALE · PHOENIX PHARMAHANDEL / LOGISTICS · BRIDGEWAY / AUTOMOTIVE · MARUTI SUZUKI / RESEARCH · WZB / INDUSTRIAL HYDRAULICS · BOSCH REXROTH

THE LEDGER

Everything above this line, with its receipt.

The argument ends here. What follows is the working: every population, every date, every denominator, and the places our own account needs checking. Nothing below is needed to understand the case. All of it is needed to check it.

L1 / what we got wrong

The remaining gap was the ground truth, not the model.

We asked for ground truth and there was not any. Bridgeway’s own load numbers and document types were the obvious yardstick, and they turned out to be about 92 to 93 percent accurate, with errors that were systematic rather than random. So we built the ruler ourselves, three times, and until July it existed in five copies at once. A benchmark that exists in five versions is not a benchmark.

On the thirteenth run against our own ground truth we were at 93 percent, and when we asked why not higher, the answer was that the remaining gap was the ground truth, not the model. It listed what was wrong with our labels. The labels we had been scoring ourselves against were themselves wrong.

For part of that work we let a language model resolve the discrepancies itself and build the ground truth, and we did not check every item by hand. So the figures in the low 90s rest on labels a human wrote, and the figures in the high 90s rest partly on labels a language model produced. They are not the same kind of number and we do not present them as one.

Every accuracy figure on this page is a Talonic measurement against ground truth Talonic built, taken before the incumbent cutover.

The rule we took from it: where several independent models agree on a value that disagrees with the label, the label is probably the thing that is wrong.

L2 / the receipts

Every number on this page, and how it was measured.

THE SHAPE OF THE ACCOUNT

CUSTOMER
Bridgeway · Pittsburgh
SECTOR
Freight brokerage with an asset based trucking arm
SCALE
~550,000 loads billed a year · 5 to 5.5 million pages
OPERATING SHAPE
23 agent brands · three locations · 7,000 to 10,000 loads open at any moment
SYSTEM OF RECORD
TMW · Microsoft SQL · 250 GB · no archiving · tables back to 2003
PER DOCUMENT
Two values before. 30 to 40 fields, a load match and an automation decision after.
INTEGRATION
Push in, webhook back. Read only snapshot. No on premise deployment.
MODE
Shadow run: the stream is duplicated, not rerouted
ENGAGEMENT
January to August 2026 · contract to production in about seven weeks

Every accuracy figure on this page is a Talonic measurement against ground truth Talonic built, taken before the incumbent cutover.

THE PRODUCTION EXTRACTION SCHEMA

20 named fields

The load number is always one letter followed by exactly seven digits. That format rule is what makes the refusal in section 05 possible.

The schema was built out of what held still: identical documents were run repeatedly and every field classified as consistent, high variance, critical or promising.

Twenty fields as defined in the production extraction schema. The delivered set runs to thirty to forty values per document, because several of these fields are arrays.

SHOW THE FIELD NAMES
  • document type
  • load number
  • reference numbers
  • weight (lbs)
  • piece count, handling units
  • shipper name
  • shipper address
  • shipper city and state
  • consignee name
  • consignee address
  • consignee city and state
  • carrier name
  • carrier MC number
  • carrier USDOT number
  • carrier invoice number
  • total payable to carrier
  • total charged to customer
  • factoring company name
  • factoring company address
  • billing party

THE MEASUREMENT LEDGER

Two rules govern this table. A figure’s grade is a property of how it was measured, not of how good it sounds. And where a benchmark rests on labels a language model produced rather than labels a person wrote, the table says so: the matching benchmark, the full schema extraction figures and the document type figure all do.

OPEN THE FULL MEASUREMENT TABLE · 27 ROWS
FIGUREWHAT IT MEASURESDATESAMPLE AND DENOMINATORWHO MEASURED ITGRADE
96% vs 78%Head to head against the incumbent, document type and load number both correctApril 2026Hand labelled documents, two fields. Both systems scored against the same labels the same way.Talonic. The customer’s own prior analysis produced the same figure for the incumbent.MEASURED
92 to 93%Human biller accuracy on the same task2026Measured against the ground truth Talonic hand built, on the same two fields. This is why the customer’s own records could not be used as the benchmark.TalonicMEASURED
~550,000Loads billed per yearStated in January and restated in May 2026Loads billed, not documents.The customer, cross checked across two conversationsMEASURED
2,200 to 2,400Loads billed per day2026The daily rate behind the annual figure above.The customerMEASURED
5 to 5.5 millionPages per yearMay 2026Pages, and the volume the engagement was scoped on.The customer, re baselined during scopingMEASURED
7,000 to 10,000Loads open at any moment2026The pool a document has to find its load in. Cross checked two ways.The customerMEASURED
4 to 5Documents per load2026Operational description, not a corpus count.The customerMEASURED
2 → 30 to 40Data points returned per document, incumbent against TalonicJuly 2026, in productionThirty to forty values per document, tuned deliberately: past a certain field count accuracy falls.Talonic, against the production extraction schemaMEASURED
20Named fields in the production extraction schemaJuly 2026Twenty fields as defined in the production extraction schema. The delivered set runs to thirty to forty values per document, because several of these fields are arrays.TalonicMEASURED
60,000+Documents processed in production13 August 2026Cumulative production volume through the shadow run. Stated by both sides.Stated by both sidesMEASURED
25,000 pages · ~4,000 runsProduction traffic at the end of July30 July 2026Customer side count.The customerMEASURED
2,000 documents / minuteThroughput23 July 2026Internal load test.TalonicMEASURED
~30 secondsAverage time per document23 July 2026Internal load test. Peak batch completed in under ten minutes.TalonicMEASURED
5 hours · ~7,000 docs · 0 errorsSoak test23 July 2026Internal soak test.TalonicMEASURED
2 seconds per 1,000 loadsMatching speedApril 2026Measured on the embedding matcher.TalonicMEASURED
12% → 30% → 76%Automation journey on one round of supplied loadsMay 2026Measured on the loads supplied for that round, roughly 150 of them. Both jumps have a stated cause and neither is a model improvement.TalonicMEASURED
150 DPIRender setting findingJuly 2026Accuracy peaks at 150 DPI and falls above it.TalonicMEASURED
6 to 9 secondsSaved per document by making OCR non blockingJuly 2026Pipeline timing on the production ingest path.TalonicMEASURED
3Times ground truth was rebuiltMarch to July 2026Fifty documents with three people in March, a hand labelled set on two fields in April, and a shared versioned benchmark in July.TalonicMEASURED
250 GB · tables to 2003The system of record2026The customer’s transport management database. No archiving.The customerMEASURED
23Agent brands carried in the brand enum2026The semi autonomous brands the operation runs through.Talonic, from the production configurationMEASURED
75 to 78% coverage at ~95% precision, ~71% end to endCurrent matching performance13 August 2026About two thousand production documents, against ground truth Talonic built, taken before the incumbent cutover.TalonicINTERNAL
~94% accuracy, ~97% precision on the non abstained setMatching benchmark4 August 2026560 real documents. The 97 percent figure was the best of several cases measured, not the average.TalonicINTERNAL
96.4% / 92.2%Full schema extraction20 July 202655 files, across the extraction platform and the resolution layer. The higher figure is the customer specific schema, the lower is overall.TalonicINTERNAL
~99%Document type accuracy20 July 2026Same 55 file run.TalonicINTERNAL
50 of 57 fieldsField stability across repeated runs17 July 2026The same documents run six times. The seven that varied were arrays plus three newly introduced fields.TalonicINTERNAL
~10%The structurally unextractable tail13 August 2026Internal estimate. Photographs of the side of a truck, partial captures. It caps achievable automation near 90 percent.TalonicINTERNAL

Note: Prints in full, open or closed.

I found multiple cases where Talonic did not think the document was a BOL. The human said it was. And when I looked at it, I didn’t think it was. And I had a couple of the billing department look at it, and they’re like, no, this is not a BOL. Talonic was correct, I think, in many of those cases.

Nate Bloom, Chief of Transformation, Bridgeway
~550,000 LOADS5.5M PAGES10,000 OPEN LOADS2 → 40 FIELDS60,000+ DOCUMENTS2,000 DOCS / MINUTE

L3 / withheld

Two boundaries this page keeps by design.

  1. 01

    Coverage and precision are two numbers, never one.

    A single percentage hides whether the system answered on everything or declined on the hard third. Every accuracy figure here is a Talonic measurement against ground truth Talonic built, and commercial figures belong to the customer.

  2. 02

    The ceiling is near 90, and it is printed.

    About one document in ten in this corpus cannot be read by anything, and that tail is stated rather than averaged away.

SUMMARY

What this case shows

FOUR CLAIMS

Four things this engagement demonstrates, each with the figure that backs it. Every one of them is checkable against the ledger above.

  1. 01

    Same documents, same ruler, different result

    The same hand-labelled documents, two fields each, both systems scored identically. The comparison is measured, not argued.

    96% vs 78% · same documents, same scoring
  2. 02

    Abstaining is part of the score

    Where the system cannot decide it says so instead of guessing, and the abstentions are printed in the account rather than folded into the accuracy figure.

    Abstentions printed, not hidden
  3. 03

    Proof before replacement

    The incumbent is a six-figure line item. The case for replacing it is a head-to-head on its own benchmark, with the ruler published.

    Six-figure incumbent · measured head-to-head
  4. 04

    Live means shadow first

    Auto billing has run in shadow alongside the team since 12 August 2026, producing the same decisions in parallel and being compared, before it decides alone.

    Shadow run since 12 Aug 2026
Document testresponse within 1 business day · your data back within 5

Response within 1 business day.