One contract is six documents.
GETEC has fifty thousand of them.
Read once. Queried by every department since.
GETEC is a German energy services group, headquartered in Magdeburg. It builds and operates energy generation plants for industrial and municipal customers, across a group of 62+ legal entities that includes acquired and renamed companies, all of which have to be identified correctly.
Roughly 55,000 contracts, each one spread across a main agreement, its amendments, its technical annexes and its attachments. That is the full contract estate. Phase one covers 8,500. A department typed the fields into Microsoft Dynamics by hand, and kept typing as new contracts arrived. Talonic reads the estate once, assembles each contract into one dataset with field-level provenance, and delivers it into Dynamics through SAP BTP.
SIX DOCUMENTS · ONE ANCHOR · ONE DATASET · ↓ SCROLL TO INSPECT
talonic · assembly traceREPRESENTATIVE BUNDLE
THE BUNDLE
THE CONSEQUENCE
1 contract dataset
one subject, one record, provenance on every value
DELIVERED · 75 FIELDS
the governed core, per contract dataset
Microsoft Dynamics via SAP BTP
the contracted destination
HELD IN THE REGISTRY
everything else the read captured, waiting for whoever asks next
Point at a department to see which lane answers it, and which slice of that lane.
Representative bundle shape. Document types are the customer’s own vocabulary, given in English. Field counts are measured on the Milestone 1 batch, July 2026: an average of 123 fields captured per document, median 70, against a delivered schema of 75 fields per contract dataset. Per document and per contract are different denominators.
GETEC
Energy · Germany
Industry
Industrial energy
Function
Contract management
Solution
Contract estate structuring
LIVE TENANT · 26 AUGUST 2026
1M+
source observations, live tenant
Note: The migration needed 75 fields per contract dataset. One read of the estate holds 1,042,665 source observations and 52,937 reusable concepts, live tenant, 26 August 2026.
THE PLATFORM RUN
READ ONCE
5,977
documents in the searchable library, read from the live tenant on 26 August 2026, against about 50,000 documents staged for phase one. A dated count of what had been read by then.
STRUCTURE
52,937
reusable concepts in the registry, resolved from 1,042,665 source observations on the same reading, and held for whoever asks next.
RESOLVE
1
anchor per contract, the main agreement, with the customer’s own document-type vocabulary around it.
DELIVER
75
fields per contract dataset, in the shape Dynamics imports, destined for Microsoft Dynamics via SAP BTP.
QUERY FOREVER
ONCE
read. 0 further reads since. Every department that arrived after the first one was served from that read.
The fifth stage prints ONCE on purpose: it is the claim every page in this family exists to receipt.
Five stages counted in five different units, so nothing across this band is a rate, a ratio or a total. 5,977 is documents in the searchable library of the live tenant on 26 August 2026, against about 50,000 documents staged for phase one. It is a dated count. 52,937 is reusable concepts across that same reading, resolved from 1,042,665 source observations; the concept and theme labels are machine generated and are not the customer’s own vocabulary. 75 fields are per contract dataset, Schema V3, frozen at Milestone 1. The delivery leg is the contracted destination, and its milestone is dated in the receipts. The registry export of 13 July 2026 and the Milestone 1 batch are earlier populations and appear in the receipts, not in this band.
INGEST ONCE. QUERY FOREVER.
01 / what one read holds
What one read holds
Seventy-five fields were what the migration needed.
The first five thousand documents held more than a million source observations.
Nobody read the estate twice.
Seventy-five fields per contract dataset is what the digitalisation and the system migration needed, and that set is governed, frozen and delivered. The read was never scoped to it. Read from the live tenant on 26 August 2026, a searchable library of 5,977 documents holds 1,042,665 source observations, which the registry organises on its own into 52,937 reusable concepts.
That is the floor and not the ceiling. Phase one alone stages about 50,000 documents, the full contract estate is about 55,000 contracts, and every document read after today lands in the same registry.
In a project scoped to the seventy-five, everything else in those documents is read once and then thrown away. A digitalisation project keeps what its schema asked for. An extraction project keeps what it was paid to extract. Everything the estate also said goes back to being unreadable at scale.
Here it is an asset instead, and an asset compounds. The next department is served from it, so are the AI initiatives behind them, and so are agents like Money Found, which is a standing question the estate can now be asked rather than a study somebody commissions.
Released, and captured
RELEASED
75
delivered fields, Schema V3
per contract dataset
frozen at Milestone 1
CAPTURED · LIVE TENANT · 26 AUGUST 2026
The seventy-five-field figure is the delivered schema, frozen at Milestone 1, counted per contract dataset. The 1,042,665 figure is source observations across the live tenant on 26 August 2026. They are different counters over different populations, and the comparison is one of scale, not a ratio. The two halves are not drawn to a common scale and the numbers do not divide into each other. The compression figure of 19.7 times is the tenant’s own, and it relates only the two figures on the right.
Three populations appear on this page, from three dates. The registry export of 13 July 2026 covers 4,967 ingested documents. The Milestone 1 batch of July 2026 covers 3,465 documents. The live tenant reading of 26 August 2026 covers a library of 5,977. They are not the same set and the numbers are not interchangeable.
The seventy-five-field figure is the delivered schema, frozen at Milestone 1, counted per contract dataset, and it is the same seventy-five the migration needed. The 1,042,665 figure is source observations across the live tenant on 26 August 2026. They are different counters over different populations, and the comparison is one of scale, not a ratio. The two halves are not drawn to a common scale and the numbers do not divide into each other.
"The first five thousand documents" is that library stated as a floor: 5,977 documents in the live tenant on 26 August 2026. "More than a million" is 1,042,665 source observations on the same reading, stated the same way.
The concept and theme labels behind the registry counts are machine generated and are not the customer’s own vocabulary.
The Money Found findings are in CFO review at GETEC. What is published is what the estate can now be asked; the amounts stay with the customer.
Everything below this band is the receipt for it. What the estate is, how it was read, how six documents become one contract, and why delivery is selective.
02 / the estate
The hard part was never reading a contract. It was knowing which contract you were reading.
Start with the identifier. The cost carrier number is what ties a document to a plant and therefore to a contract. One cost carrier number is one plant is one project. As a rule it is not printed in the document at all, and contracts written in the late 2010s and early 2020s carry cost carrier numbers under a numbering scheme that no longer exists. About 50,000 documents staged for phase one, and no reliable way to say which six of them are one contract.
The metadata cannot carry the routing either. Document type was entered by hand over many years and is routinely wrong, and one category functions as a catch-all for anything nobody classified. It was ruled out as a routing signal before a single document was read. The terms themselves are scattered: obligations recur across separate paragraphs rather than sitting in one clause, ownership boundaries are frequently expressed only in the technical drawings that arrive as a technical annex, and the same field can carry different values in the main agreement, in its amendment and in its attachment. Something has to decide which document wins.
The recurring cost is the lookup, not the read. Reading a contract once is a bounded cost. Answering a question about a contract, every time somebody downstream needs the answer, is not. That is the cost that compounds, and it is the one this engagement was built to remove.
An estate is not a corpus of one document type. It is a long tail with a small core.
Two populations appear below, from two different dates. The registry export of 13 July 2026 covers 4,967 ingested documents. The Milestone 1 batch of July 2026 covers 3,465 documents. They are not the same set and their numbers are not interchangeable.
What the estate is made of
Field registry export, GETEC tenant, 13 July 2026.
13 July 2026 registry export, 139 classified document types in total. Occurrence rows are pruned by lifecycle retention, so these counts are a floor rather than a census. The export records 670 as unclassified at that date, reproduced here as recorded.
Milestone 1 batch · July 2026
3,465documents in the batch
99.1%completeness
Completeness, not accuracy: 3,435 of 3,465 documents produced a full extraction. Accuracy is measured separately, against the customer’s gates.
avg 123fields captured per document
Median 70. Per document, not per contract dataset.
Handled in the Milestone 1 batch. Two counts from the ledger, printed as they are recorded and not added together.
Milestone 1 batch, July 2026. A different and earlier population from the registry export above, and the two sets of numbers are not interchangeable.
cost_carrier_number · one plant, one contractStand-in imagery, not a GETEC site · Photo: A.Savin, FALStand-in imagery. Not a GETEC site.
03 / the read
The read is not scoped to the schema. It is scoped to the document.
01
One shot, not two passes.
The raw PDF goes into the model together with the schema prompt, and markdown is produced in parallel for everything downstream. There is no layout pass in front of the read and no second pass behind it. Around that single read sit the things that make an archive tractable: classification at ingest, a generated document summary, byte-identical deduplication that links a duplicate to its original rather than blocking it, and a merge step for documents that arrived split. Both sides of this engagement arrived at the one-shot argument independently, which is worth saying because it is the decision everything else rests on.
02
Routing runs on the customer’s own ontology, not on the archive’s metadata.
Rules are driven by document type, and the customer’s own type ontology is auto-matched to Talonic’s at ingest. So the type that routes a document is the one the read assigns it, on both ontologies at once. The document type hand-entered into the source system over many years was ruled out as a routing signal before a document was read, for the reasons in section 02, and nothing downstream was allowed to depend on it.
03
The read captures more than the schema asks for, and that is a measurement rather than a claim.
In the Milestone 1 batch, capture averaged 123 fields per document, median 70, against a delivered schema of 75 fields per contract dataset. The surplus is not rhetorical. It is the average.
Average captured per document in the Milestone 1 batch, July 2026, against a delivered schema of 75 fields per contract dataset. Per document and per contract are different denominators, and a contract is about six documents. Those two numbers do not divide into each other.
04
Open capture is a decision, not a model behaviour.
Nothing about a model makes it reach past the brief on its own. The fields nobody asked for are captured because somebody decided they would be, and they are stored in a parallel structure rather than in the governed set. That means a field can be captured specifically because a later question is foreseeable but unaskable now. Operating electricity, the field that records who carries the operating-electricity cost, is the clearest example: a system given no instruction is unlikely to find it, so it was named and added to the capture set.
DOCUMENT TYPES · THE CUSTOMER’S WORDS
The customer’s vocabulary. Their word wins.
Talonic classifies every document against its own 529-type ontology across ten categories, so that a German employment contract and an English one resolve to one canonical type. But the ontology the reviewer works in has to be the customer’s own. In this estate Talonic’s document type became the customer’s own type list, "Other" became technical annexes and attachments, and the anchor document type is the main agreement. Every document is classified on both axes at once, so reporting works in Talonic’s canonical types and the review screen works in the reviewer’s own words. One pass, two labels, no translation layer in between.
Main agreement
the anchor
the document the dataset is built on
Amendment
overrides
parties, dates, commercial terms
Technical annex
annexes
technical scope, often as drawings
Attachment
annexes
attached, and never overriding
Correspondence
ignored
in the bundle, out of the dataset
Cost carrier number
the join key
one plant, one project, rarely printed
Document types are the customer’s own vocabulary, given here in English. The roles are the override gates above, stated in words.
THE MEASURED EFFECT
Declaring the customer’s own field names is not cosmetic. In a separate evaluation on 20 documents in German, English and Chinese, with hand-authored ground truth run through the production pipeline, capture rose from 8 of 81 fields to 61 of 81 when the customer’s ontology was declared. A lift of 65 points, with dual-axis classification correct on 19 of 19.
That evaluation ran on its own 20-document corpus, separate from GETEC’s estate. It measures the mechanism itself.
Classifying on both axes at once is what lets a later department slice the same read in its own vocabulary without a re-ingest. The words change. The read does not happen again.
04 / assembly
Assembly is not a merge. It is an argument about which document is the contract.
01
One of them is the anchor.
A contract is a main agreement, two to four amendments, one or two technical annexes and one or two attachments. About six documents, written over years, sometimes by different parties, sometimes contradicting each other. One of them is the anchor. Everything else inherits from it, overrides it, or annexes it. Get the anchor right and the rest of the bundle resolves. Get it wrong and every field in the dataset is wrong in a way that looks correct.
02
The archive gave us the grouping. It did not give us the anchor.
Every document in the source system carries exactly one primary key, shared by all documents belonging to one contract partner, and the customer exported it as an accompanying table rather than making us reconstruct the grouping. That removed most of the hard problem. It did not remove the anchor problem, and where a bundle has two candidate main-contract documents, the markdown of both goes to a small fast model, which decides whether they are one dataset or two. That decision is cheap, and it is the only part of assembly a model is allowed to arbitrate.
03
Everything else is deterministic, and every value keeps its receipt.
Document-type-level override gates state which type may override which field. They are configured per document type and per estate, and they are deterministic once configured, so the model cannot violate them. Every value that survives assembly carries the page and line it was read from, its scan region, the model’s confidence and reasoning, the human decision that approved, corrected or overrode it, and the delivery event that pushed it. AI values are labelled as AI values so they are never mistaken for the record, and a human decision outranks a machine one.
If the anchor is unclear, assembly does not fail loudly. It fails quietly. So much of the dataset is built on top of the anchor that any ambiguity cascades, and cascading errors are silent ones. Extraction was solved early. Composition is the part that stayed hard.
The anchor, and what may override it
Override gates are configured per document type and per estate, and they are deterministic once configured. The gate configuration shown is illustrative of the mechanism: the one arbitration a model is permitted is the dashed edge, where two documents both look like the main contract.
05 / the argument
Extract now. Release later.
01
The one-shot requirement came from both sides.
Extracting to one department’s schema means extracting again when the next department arrives, and paying twice. So the read is exhaustive and the delivery is selective. Talonic wanted it because re-reading is waste, and the customer wanted it because management funds one read, not four.
02
A second schema is what stops the first department’s contract from blocking everyone else.
The business acceptance owner’s field set is the governed core that goes into Dynamics, and every change to it is a negotiation. If every new department had to widen that set, expansion would stall. So a second schema holds the non-governed fields in a parallel structure, and a new department adds what it needs without renegotiating the first department’s contract. It is an organisational decision expressed as a data model, and it is the most interesting decision on this page.
03
The proof that it works.
The second department in was asset management, and it did not ask for more fields of the same kind. It asked for images, schematics, technical diagrams, delivery and performance boundaries, measurement concepts, and free-text search across the estate. The Milestone 1 batch had already handled 1,514 schematic drawings and 1,116 technical diagrams, so the request was substantially already in the estate, because the first read had not been scoped to the first schema.
Technical drawings sit outside the 75-field delivered schema by design: they were descoped from the first step and held in the read.
Four field counts, four denominators
Capture is exhaustive and delivery is selective, so the delivered set is the smallest of four concentric things. It is drawn that way here, and it is never drawn as a line going up.
22,826core fields in the registryper estate · 13 July 2026 export
avg 123fields captured, median 70per document · Milestone 1 batch
82fields extracted by the pipelineper pipeline · 5 August 2026
123 is per document. 75 is per contract dataset. A contract is about six documents. Those two numbers do not divide into each other.
Four different denominators. Per estate, per document, per pipeline, per contract dataset. This is containment, not growth over time. Frame sizes are not to scale, and the sequence is not a chronology. The delivered schema of 75 fields was frozen at Milestone 1. It grew from 59 over about thirteen weeks, and freezing it is what made Milestone 1 reviewable.
06 / the registry
The same meaning repeats across the estate. That is what makes it comparable.
Three populations appear on this page, from three dates. The registry export of 13 July 2026 covers 4,967 ingested documents. The Milestone 1 batch of July 2026 covers 3,465 documents. The live tenant reading of 26 August 2026 covers a library of 5,977. They are not the same set and the numbers are not interchangeable.
01
Observations become concepts.
Read from the live tenant on 26 August 2026, 1,042,665 source observations resolve into 52,937 reusable concepts. The mechanism is the product’s own: observations with similar meaning cluster automatically using embeddings, and as the same thing is read out of many documents the system synthesises a master extraction instruction for it, a reusable directive that captures the best way to read it and improves accuracy on every subsequent run. Repetition is what makes cross-document comparison possible. Nothing is held here for its own sake: what is held is what makes a question across a whole estate answerable at all.
02
Three tiers, in the product’s published vocabulary.
Tier 1 Core are universal across many document types and the most reliable. Tier 2 Established were promoted out of Tier 3 after meeting frequency thresholds. Tier 3 Emerging are newly discovered, and are candidates for promotion as more data arrives. 19,169 of the 52,937 reusable concepts are established across tiers 1 and 2, and that subset is the one worth a number, because generated schemas are produced from Tier 1 and Tier 2 registry fields. The thresholds are the product’s own; this page describes the promotion mechanism.
03
The registry is still learning, in public.
The tenant’s own momentum feed records a concept named elevation.313_20 promoted from Tier 3 to Tier 2, four days before this reading, marked CONCEPT STRENGTHENED. That is a machine-generated key out of this estate crossing a threshold. It is the promotion rule firing on real data, in a live tenant, four days before we looked.
The tiers, and one promotion
Read from the live GETEC tenant on 26 August 2026.
Tier definitions are the product’s own. Generated schemas are produced from Tier 1 and Tier 2 registry fields, which is why the established subset is the one with a number against it. The promotion event is reproduced from the live tenant’s own momentum feed, 26 August 2026. The registry key is machine generated. Promotion thresholds are the product’s own. The rows are drawn equal in height; tier sizes are not drawn to scale.
Read from the live GETEC tenant on 26 August 2026. A later and larger population than the 13 July 2026 registry export elsewhere on this page. The two sets of numbers are not interchangeable and do not subtract into each other.
The concept and theme labels behind the registry counts are machine generated and are not the customer’s own vocabulary.
AND EVERY SYSTEM AFTER DYNAMICS
Dynamics is the first destination, not the last. A parser-based migration solves the first destination and re-extracts for every one after it. A registry captures the estate once and projects it into each schema as it appears, which is also what an agent needs: an agent that can ask for every price adjustment formula in the estate and receive them as rows can work the whole estate algorithmically, rather than trying to fit a corpus into a prompt.
Microsoft Dynamics through SAP BTP is the contracted destination, with its go-live milestone dated in the receipts. An operational data lake, a reporting layer and agents are the identified positions after it.
MONEY FOUND
07 / the surplus
Somebody asked a question the original scope never anticipated. The answer was already in the registry.
Talonic extracts everything from every document once, holds the fields nobody has asked for yet, and releases them to each department as that department arrives. Money Found is the sharpest instance of that, and it is not a separate product.
The findings in this section are in CFO review at GETEC. What is published here is what the estate can now be asked; the amounts stay with the customer.
01
Cost allocation that is contested by nature.
Who bears the operating-electricity cost under the contract is a genuinely contested term in energy contracting. It was added to the capture set deliberately, on the reasoning that a system given no instruction is unlikely to find it, which is the same decision section 03 describes as open capture. It is in the read now, so operations and technical can ask it of the estate rather than of a filing cabinet.
02
An entitlement with no system that executes it.
This one is structural rather than per contract. The price adjustment is defined in the contract, and the system where it would be actioned is a different system from the one holding the contract data, so the terms and the billed values have never been in the same place. The registry is the first place they meet. Asset management is where the gap shows. Price management is where the tasks would be built.
03
Escalation entitlement, released from inventory.
Price adjustment clauses carry formulas, parameters, base values and annual index factors. Those fields were already being captured and held back from the delivered dataset a month before anyone framed an entitlement use case. The same fields were then queried and resolved into the full formula. Money Found was not a new capability. It was the release of inventory that already existed.
What that changes is the shape of the work, not the size of a number.
An afternoon, not a project.
The finding was demonstrated in a session, unprompted and free of charge, because the fields had already been captured. The marginal cost of serving department number two is close to zero, because the documents have already been read.
A standing query, not a one-off study.
The incumbent for this class of question is a one-off manual study by an outside consultancy. A registry makes it a question you can ask on the day you want the answer, and ask again the day after.
Three further directions are queued from the same registry.
AND ONE BOUNDARY CASE, WHICH IS ABOUT PROVENANCE
When a value in the registry disagrees with the system of record, one of the two can show you the page and line it was read from. In the customer’s own review of extraction results, some disagreements resolved in the registry’s favour.
That review is a lower bound. The customer’s own verification tool had known failure modes on handwriting and on dates.
Types of finding
OPERATIONS
Cost allocation, contested by nature
IN THE CONTRACTwho bears operating-electricity cost
IN THE REGISTRYa field added because nothing finds it unprompted
ASSET MANAGEMENT
An entitlement with no system that executes it
IN THE CONTRACTthe price adjustment the contract defines
IN THE REGISTRYthe first place the contract terms and the billed values meet
PRICE MANAGEMENT
Escalation entitlement, released from inventory
IN THE CONTRACTformulas, parameters, base values, index factors
IN THE REGISTRYalready captured and held, a month before anyone asked
NEXT IN THE QUEUE
contributions contractually due and not received
the obligation gap, in both directions, between service delivered and service owed
payables reconciled against contract data
Types of finding. The amounts stay with the customer, and the findings are in CFO review at GETEC. The lower row is the next set of questions queued against the registry.
The queue behind the first department
Departments by function only. One lane is live and the rest are arriving, which is the whole argument for holding the fields nobody had asked for yet.
department
what it needs
status
Contract management
The governed core, 75 fields, into Dynamics
liveLive. The original scope
Asset management
Images, schematics, delivery and performance boundaries, measurement concepts, free-text search
second inSecond department in
Price management
Price adjustment parameters, held in a separate finance system
namedIn CFO review. The entitlement findings are with finance
HSEQ
Certifications and employer data
namedFolded into the field table
Operations
Operating-electricity cost responsibility, a genuinely contested term
namedField added
Service delivery
Service level definitions, ticketing, key-user model
openOpen
Legal
Out of scope, and a deliberate boundary
boundaryBoundary noted
IT and integration
Field-name parity across CRM and middleware
openOngoing
08 / next
Want to see this on your estate?
Send a representative sample of contracts, amendments, annexes and attachments. We run them through Talonic and send back your extracted data within five business days: field-level coverage, confidence and provenance. Judge it field by field.
SIX CASE STUDIES: INDUSTRIAL ENERGY · GETEC / PHARMA WHOLESALE · PHOENIX PHARMAHANDEL / LOGISTICS · BRIDGEWAY / AUTOMOTIVE · MARUTI SUZUKI / RESEARCH · WZB / INDUSTRIAL HYDRAULICS · BOSCH REXROTH
THE LEDGER
Everything above this line, with its receipt.
The argument ends here. What follows is the working: every population, every date, every denominator, and the places our own estimates were wrong. Nothing below is needed to understand the case. All of it is needed to check it.
L1 / what we got wrong
We estimated one minute a field. We watched twenty.
The review queue was designed field by field, on an estimate of under a minute per field. Sitting with the reviewers on 30 July 2026, we observed about twenty minutes, because every field meant opening the CRM and then opening the sibling documents. The estimate was not slightly wrong. It was wrong by a factor of twenty.
Our assumptions for the review queue were badly wrong, and not unfortunately. We could not have known better from a spec. We only found it by sitting next to somebody doing the work. The redesign was same-day: review the assembled data product, not the field, with concurrency moved to the data-product level.
A separate internal measurement, taken on 3 July 2026 and four weeks before that observation, put review load at about 20 items per assembly, down from about 56, against a target of 5 to 6. It is a different measurement on a different date, taken before the redesign above, and it is progress rather than a finish line.
Talonic’s own estimate against Talonic’s own observation, sitting with reviewers on 30 July 2026, and Talonic’s own measurement of Talonic’s own interface on 3 July 2026. Two dates, two measurements, both Talonic’s own.
One minute, estimated. Twenty, observed.
Talonic’s own estimate against Talonic’s own observation, sitting with reviewers on 30 July 2026. The estimate bar is drawn to scale against the observation, and it is meant to be almost invisible. The pair on the right is a separate internal measurement dated 3 July 2026, four weeks earlier, taken before the redesign that followed the observation.
Two things we chose not to show, and they belong here.
The verdict that was never wrong.
A verdict that told a reviewer a value was likely fine was deleted, because we could not find a case where it was wrong. A verdict that is never wrong in testing is a verdict nobody will check.
The chip with no count.
The review-count chip shipped with no count on it, because a single large number at the top of a queue reads as a verdict on the person, not on the work.
AND ONE MORE, ABOUT OUR OWN TESTING
A regression fired manual approval prompts on a pipeline that had already completed. Within 24 hours we abandoned DOM-based end-to-end testing for a visual browse agent, because the interface had no stable identifiers for a test to anchor to, and a test suite nobody trusts is worse than none.
A model will occasionally be completely confident that a wrong value is right. That does not go away. What changes is whether the interface pretends otherwise.
L2 / the receipts
What we measure, and who owns the gates.
The gates belong to the customer. This engagement runs against contractual accuracy gates at successive milestone reviews, and at every review the customer has the right to stop the project if the targets are not met. The commercial structure follows from that. No subscription revenue is earned until the customer accepts the first accuracy gate in writing. We do not get paid for accuracy until the customer says it is there.
≥90%first milestone review
≥94%then
≥97%then
These are the thresholds the customer measures the engagement against, at each milestone review.
Accuracy is measured against ground truth authored by the customer’s own subject-matter expert, behind three blocking validation rules. The measurement was designed to be failable. That is the point of it.
anchor completeness
amount consistency
grouping and join coverage
Completeness, and what it measures.
In the Milestone 1 batch, 3,435 of 3,465 documents were fully extracted. That is 99.1 percent completeness.
Completeness is the share of documents that produced a full extraction. Accuracy is a separate measurement, taken against the gates above.
Engagement-level accuracy is shared under NDA. When a gate is formally accepted, the number is published with its denominator, its date and who measured it.
And the gates that have nothing to do with models.
Enterprise AI in Germany clears gates that have nothing to do with models. A supplier security assessment. A data protection assessment owned by the data protection officer. And a works council framework agreement for AI, tracked on this project as a named risk from the start, because co-determination is a real approval path and not a formality.
The deployment that clears those gates is a single-tenant one on Talonic’s own Azure infrastructure, region Frankfurt, with the OCR running inside it. No external AI APIs and no sub-processors, and no direct write path into the system of record: delivery goes through the middleware the customer’s architecture board already mandated.
The deployment model as contracted, and as answered in the supplier security assessment below.
THE SUPPLIER SECURITY ASSESSMENT
The supplier security risk assessment was cleared, with one open condition: the business impact analysis, owned by a business owner. The separate GDPR assessment was deferred.
The assessment covers the deployment model. Encryption specifics, ISO 27001 and SOC 2 sat outside it.
The rest of the engagement, in plain rows.
CUSTOMERGETEC
ESTATE~55,000 contracts in total · phase one covers 8,500 contracts and ~50,000 documents
DOCUMENTS PER CONTRACT~6 · 1 main agreement, 2 to 4 amendments, 1 to 2 technical annexes, 1 to 2 attachments
DELIVERED SCHEMA75 fields · Schema V3, frozen at Milestone 1 · grew from 59 over about 13 weeks
DEPLOYMENTSingle tenant · Azure · Frankfurt · Entra ID and OIDC SSO with MFA · 99.5% availability target
THE MANUAL BASELINE, MODELLED
12,750 to 17,000 hours7 to 9 person-years
A modelled manual baseline: 8,500 contract bundles at about six documents each, at roughly 1.5 to 2 hours per bundle to read, extract 75 fields and consolidate amendments. The assumptions are the model, and it is stated in hours only.
The milestone arc, as contracted
Milestones as contracted, with their contractual dates. The delivery leg from SAP BTP into Dynamics is built by GETEC’s own IT.
~55,000 CONTRACTS~50,000 DOCUMENTS IN SCOPE5,977 DOCUMENTS READ · LIVE TENANT139 DOCUMENT TYPES · 13 JULY 2026 EXPORT75 FIELDS DELIVERED PER CONTRACT DATASETAVG 123 FIELDS CAPTURED PER DOCUMENT62+ LEGAL ENTITIES1 ANCHOR PER CONTRACT~55,000 CONTRACTS~50,000 DOCUMENTS IN SCOPE5,977 DOCUMENTS READ · LIVE TENANT139 DOCUMENT TYPES · 13 JULY 2026 EXPORT75 FIELDS DELIVERED PER CONTRACT DATASETAVG 123 FIELDS CAPTURED PER DOCUMENT62+ LEGAL ENTITIES1 ANCHOR PER CONTRACT
L3 / withheld · WHAT STAYS WITH THE CUSTOMER
Accuracy is measured against gates the customer owns, and engagement-level accuracy is shared under NDA.
Commercial figures belong to the customer: the contract value, the pricing and the order history.
Individual names stay with the customer, and a quotation appears here once it is consented and verified.
Amounts from the customer’s own data, systems and internal processes stay with the customer. The section on what the registry made askable names the types of finding.
Money Found amounts stay with the customer. The findings are in CFO review at GETEC.
SUMMARY
What this case shows
FOUR CLAIMS
Four things this engagement demonstrates, each with the figure that backs it. Every one of them is checkable against the ledger above.
01
Reading once costs less than reading repeatedly
Contract management needed 75 fields. The departments that arrived after it needed different ones: images, schematics, price parameters, certifications, cost responsibility. All of them were answered from a single pass over the documents.
5 departments · 1 read
02
What you capture beyond the brief is what makes it reusable
The read averaged 123 fields per document when only 75 were required. Everything past the brief is what answered each department that turned up later, recorded before anyone knew to ask for it.
avg 123 captured · 75 delivered
03
The hard part is identity, not extraction
Pulling a date off a page is the easy half. Deciding which of six documents states the terms, and which of them is allowed to change it, is where a contract estate is won or lost, and it fails quietly when you get it wrong.
1 anchor per contract · rules per document type
04
The numbers are built for scrutiny
The accuracy gates belong to GETEC and are measured against ground truth its own expert authors. What we got wrong is printed alongside what we got right, and every figure on this page names its measurement date.
3 measurement dates · 0 unsourced figures
Document testresponse within 1 business day · your data back within 5