Every number on this page, and how it was measured.
The shape of the account
- Customer
- Bridgeway · Pittsburgh
- Sector
- Freight brokerage with an asset based trucking arm
- Scale
- ~550,000 loads billed a year · 5 to 5.5 million pages
- Operating shape
- 23 agent brands, three locations, 7,000 to 10,000 loads open at any moment
- System of record
- TMW, Microsoft SQL, 250 GB, no archiving, tables back to 2003
- Per document
- Two values before. 30 to 40 fields, a load match and an automation decision after.
- Integration
- Push in, webhook back. Read only snapshot. No on premise deployment.
- Mode
- Shadow run: the stream is duplicated, not rerouted
- Engagement
- January to August 2026 · contract to production in about seven weeks
The production extraction schema
20 named fields
The load number is always one letter followed by exactly seven digits. That format rule is what makes the refusal in section 05 possible.
The schema was built out of what held still: identical documents were run repeatedly and every field classified as consistent, high variance, critical or promising.
Twenty fields as defined in the production extraction schema. The delivered set runs to thirty to forty values per document, because several of these fields are arrays.
SHOW THE FIELD NAMES
- document type
- load number
- reference numbers
- weight (lbs)
- piece count, handling units
- shipper name
- shipper address
- shipper city and state
- consignee name
- consignee address
- consignee city and state
- carrier name
- carrier MC number
- carrier USDOT number
- carrier invoice number
- total payable to carrier
- total charged to customer
- factoring company name
- factoring company address
- billing party
The measurement ledger
Two rules govern this table. A figure’s grade is a property of how it was measured, not of how good it sounds. And where a benchmark rests on labels a language model produced rather than labels a person wrote, the table says so: the matching benchmark, the full schema extraction figures and the document type figure all do.
OPEN THE FULL MEASUREMENT TABLE · 27 ROWS
| FIGURE | WHAT IT MEASURES | DATE | SAMPLE AND DENOMINATOR | WHO MEASURED IT | GRADE |
|---|
| 96% vs 78% | Head to head against the incumbent, document type and load number both correct | April 2026 | Hand labelled documents, two fields. Both systems scored against the same labels the same way. | Talonic. The customer’s own prior analysis produced the same figure for the incumbent. | MEASURED |
|---|
| 92 to 93% | Human biller accuracy on the same task | 2026 | Measured against the ground truth Talonic hand built, on the same two fields. This is why the customer’s own records could not be used as the benchmark. | Talonic | MEASURED |
|---|
| ~550,000 | Loads billed per year | Stated in January and restated in May 2026 | Loads billed, not documents. | The customer, cross checked across two conversations | MEASURED |
|---|
| 2,200 to 2,400 | Loads billed per day | 2026 | The daily rate behind the annual figure above. | The customer | MEASURED |
|---|
| 5 to 5.5 million | Pages per year | May 2026 | Pages, and the volume the engagement was scoped on. | The customer, re baselined during scoping | MEASURED |
|---|
| 7,000 to 10,000 | Loads open at any moment | 2026 | The pool a document has to find its load in. Cross checked two ways. | The customer | MEASURED |
|---|
| 4 to 5 | Documents per load | 2026 | Operational description, not a corpus count. | The customer | MEASURED |
|---|
| 2 → 30 to 40 | Data points returned per document, incumbent against Talonic | July 2026, in production | Thirty to forty values per document, tuned deliberately: past a certain field count accuracy falls. | Talonic, against the production extraction schema | MEASURED |
|---|
| 20 | Named fields in the production extraction schema | July 2026 | Twenty fields as defined in the production extraction schema. The delivered set runs to thirty to forty values per document, because several of these fields are arrays. | Talonic | MEASURED |
|---|
| 60,000+ | Documents processed in production | 13 August 2026 | Cumulative production volume through the shadow run. Stated by both sides. | Stated by both sides | MEASURED |
|---|
| 25,000 pages · ~4,000 runs | Production traffic at the end of July | 30 July 2026 | Customer side count. | The customer | MEASURED |
|---|
| 2,000 documents / minute | Throughput | 23 July 2026 | Internal load test. | Talonic | MEASURED |
|---|
| ~30 seconds | Average time per document | 23 July 2026 | Internal load test. Peak batch completed in under ten minutes. | Talonic | MEASURED |
|---|
| 5 hours, about 7,000 docs, 0 errors | Soak test | 23 July 2026 | Internal soak test. | Talonic | MEASURED |
|---|
| 2 seconds per 1,000 loads | Matching speed | April 2026 | Measured on the embedding matcher. | Talonic | MEASURED |
|---|
| 12% → 30% → 76% | Automation journey on one round of supplied loads | May 2026 | Measured on the loads supplied for that round, roughly 150 of them. Both jumps have a stated cause and neither is a model improvement. | Talonic | MEASURED |
|---|
| 150 DPI | Render setting finding | July 2026 | Accuracy peaks at 150 DPI and falls above it. | Talonic | MEASURED |
|---|
| 6 to 9 seconds | Saved per document by making OCR non blocking | July 2026 | Pipeline timing on the production ingest path. | Talonic | MEASURED |
|---|
| 3 | Times ground truth was rebuilt | March to July 2026 | Fifty documents with three people in March, a hand labelled set on two fields in April, and a shared versioned benchmark in July. | Talonic | MEASURED |
|---|
| 250 GB · tables to 2003 | The system of record | 2026 | The customer’s transport management database. No archiving. | The customer | MEASURED |
|---|
| 23 | Agent brands carried in the brand enum | 2026 | The semi autonomous brands the operation runs through. | Talonic, from the production configuration | MEASURED |
|---|
| 75 to 78% coverage at ~95% precision, ~71% end to end | Current matching performance | 13 August 2026 | About two thousand production documents, against ground truth Talonic built, taken before the incumbent cutover. | Talonic | INTERNAL |
|---|
| ~94% accuracy, ~97% precision on the non abstained set | Matching benchmark | 4 August 2026 | 560 real documents. The 97 percent figure was the best of several cases measured, not the average. | Talonic | INTERNAL |
|---|
| 96.4% / 92.2% | Full schema extraction | 20 July 2026 | 55 files, across the extraction platform and the resolution layer. The higher figure is the customer specific schema, the lower is overall. | Talonic | INTERNAL |
|---|
| ~99% | Document type accuracy | 20 July 2026 | Same 55 file run. | Talonic | INTERNAL |
|---|
| 50 of 57 fields | Field stability across repeated runs | 17 July 2026 | The same documents run six times. The seven that varied were arrays plus three newly introduced fields. | Talonic | INTERNAL |
|---|
| ~10% | The structurally unextractable tail | 13 August 2026 | Internal estimate. Photographs of the side of a truck, partial captures. It caps achievable automation near 90 percent. | Talonic | INTERNAL |
|---|
Note: Prints in full, open or closed.
I found multiple cases where Talonic did not think the document was a BOL. The human said it was. And when I looked at it, I didn’t think it was. And I had a couple of the billing department look at it, and they’re like, no, this is not a BOL. Talonic was correct, I think, in many of those cases.
Nate Bloom, Chief of Transformation, Bridgeway
~550,000 LOADS5.5M PAGES10,000 OPEN LOADS2 → 40 FIELDS60,000+ DOCUMENTS2,000 DOCS / MINUTE~550,000 LOADS5.5M PAGES10,000 OPEN LOADS2 → 40 FIELDS60,000+ DOCUMENTS2,000 DOCS / MINUTE