GR-03
Hours are recorded verbatim.
2/3, +2/3, 1,5 and (4) are kept exactly as printed. Only a dash or a blank becomes 0, and that one recoding was the researchers’ own request.
SOCIAL SCIENCE RESEARCH · GERMANY · ONE BATCH DELIVERED
The WZB Berlin Social Science Center is a social science research institute in Berlin and part of the Leibniz Association. A research team there is building a dataset of every school timetable regulation issued by every German federal state since 1949, so that the amount of teaching time given to each subject can be compared across states and across decades. That data exists only inside legal texts, and until now researchers hand-coded it.
Talonic read one state’s batch. Twenty regulations from Mecklenburg-Vorpommern, 1996 to 2021, twelve of them 300 dpi page images with no text layer at all. They were structured exactly as printed, under a seventeen-rule spec the researchers co-authored across six rounds, into 3,179 rows that each carry a source page, a confidence score and a note on anything ambiguous. The rest of the corpus, more than a thousand documents across sixteen states, is targeted for the bulk run that follows acceptance of this batch.
24 TABLES · 7 FLAGGED · INCLUDING BOTH THE RESEARCHERS FOUND BY HAND · ↓ SCROLL TO INSPECT
talonic · deterministic check
REDRAWN FRAGMENT · § 19 SPECIAL-NEEDS SCHOOL, 1996
| Subject | Prep | 1 | 2 | 3 | 4 | 5 | 6 | 7 | 8 | 9 | Total |
|---|---|---|---|---|---|---|---|---|---|---|---|
| GermanAS PRINTED | 6 | 6 | 5 | 5 | 6 | 6 | 5 | 4 | 4 | 47 | |
| GermanAS RECOVERED | · | · | · | · | · | · | · | · | · | 51 | |
| Mathematics | · | · | · | · | · | · | · | · | · | · | |
| General studies | · | · | · | · | · | · | · | · | · | · | |
| PE | · | · | · | · | · | · | · | · | · | · |
THE SAME SHIFT, ON THE COLUMN TOTALS
A shift of exactly one column. The record transcribes this part of the row and no more, so this part is what is drawn.
CELL-COUNT: FAIL · ROW-SUM: FAIL · ACTION: RE-READ THE PAGE
A printed timetable carries its own totals: a total column per subject and usually a total row per grade. When layout recovery drops the blank cells in a column that looks empty, every value slides one column left and the sums stop matching. On the founding decree of this batch the deterministic checks flagged 7 of the 24 recovered tables, including both of the two the researchers had reported by hand. That is a flag rate, on one document.
WZB
Research · Germany
ONE BATCH DELIVERED · THE CORPUS GATED
3,179 rows
from 18 of one state’s 20 regulations
Note: German school timetable law from 1996 to 2021, transcribed exactly as printed. Every row carries its source page, a confidence score, and a note on anything ambiguous.
THE PLATFORM RUN
READ ONCE
20
documents in one German state’s batch, Mecklenburg-Vorpommern 1996 to 2021, twelve of them 300 dpi page images with no text layer at all. One state’s batch, against a corpus of more than a thousand documents targeted for the bulk run.
STRUCTURE
3,179
rows in the Timetable Data sheets of eighteen files from that batch, at the 10 August 2026 redelivery, under a seventeen-rule extraction spec at version six. Rows. Not documents and not fields.
RESOLVE
0
cross-document reconciliations, by rule. One regulation is never resolved against another, because a reconciled view is not a legal record. Mapping subject names onto research categories is the researchers’ decision and it stays theirs.
DELIVER
1
Excel workbook per document, three sheets: Metadata, Timetable Data, Footnotes, plus a ZIP per batch, delivered as files to the research team.
QUERY FOREVER
ONCE
read. Zero re-reads of the scans since: every one of the six schema rounds was measured by re-running from cached layout recovery, without repeating OCR. The remaining thousand-plus documents follow in the bulk run after acceptance of this batch.
The fifth stage prints ONCE on purpose: it is the claim every page in this family exists to receipt.
Five stages counted in five different units, so nothing across this band is a rate, a ratio or a total. The twenty documents in the first stage and the eighteen files behind the row count are not the same set: the row count was reported on the eighteen files outside the two the researchers reviewed. The zero in the third stage is a design decision. The workbook in the fourth stage is a file.
READING LAW EXACTLY AS PRINTED
01 / the checksum
Twenty regulations, 1996 to 2021, twelve of them 300 dpi page images with no text layer. The document prints its own checksum. So "did we read it right" is arithmetic, not an opinion. On the founding decree of this batch, 7 of 24 tables flagged. A flag rate on one document.
A printed timetable carries its own totals: a total column per subject and usually a total row per grade. When layout recovery drops the blank cells in a column that looks empty, every value slides one column left and the sums stop matching.
Talonic read one state’s batch. Twenty regulations from Mecklenburg-Vorpommern, structured exactly as printed under a seventeen-rule spec the researchers co-authored across six rounds, into 3,179 rows that each carry a source page, a confidence score and a note on anything ambiguous.
The check costs nothing, it runs before any structuring, and it is deterministic. The validator does not repair. It refuses, and sends the page back to be read again.
The rest of the corpus, more than a thousand documents across sixteen states, is targeted for the bulk run that follows acceptance of this batch.
7 of 24
RECOVERED TABLES FLAGGED ON THE FOUNDING DECREE OF THIS BATCH
Flagged by the deterministic cell-count and row-sum checks before any structuring, including both of the two tables the researchers had already reported by hand.
A flag rate, on one document.
3,179 is rows in the Timetable Data sheets of eighteen files from that batch, at the 10 August 2026 redelivery. Rows, not documents and not fields.
Everything below this band is the receipt for it. What arrived and why it resists, the rules the researchers wrote, the arithmetic that holds the reading to the page, and what we got wrong on the way.
02 / the corpus
A school timetable, literally an hours table in the law’s own vocabulary, is a table of subject by grade by weekly hours, printed inside a regulation. The regulations are what a state education ministry publishes in its gazette, and they arrive under the law’s own names: regulation, decree, circular decree, notice, amending regulation, correction notice. The numbers a researcher wants have been public for decades. They have simply been typeset rather than tabulated, and the typesetting is the work.
Reading one is not OCR. German timetables join several adjacent subjects with a brace and print one shared value for the group. The brace is a typographic glyph with no height, so when layout recovery flattens the page the shared value lands on whichever row the engine thought was nearest. Four subjects become one subject with a number and three subjects with nothing, or the value is copied onto all four. Both readings are wrong, and only one of them looks wrong.
Four more traps, one sentence each. A column that is blank on every subject row still exists as a column, because only the total row carries a value there, and dropping it slides every value one place left. Physics / biology or natural sciences is a choice between two subjects and one, not a long subject name, and only the arithmetic can tell those apart. Hours are printed as 2/3, +2/3, 1,5, (4) and 12 to 14, so the field is free text and not an integer. And one timetable in this batch is three sentences of prose telling you to replace a five with a two in one cell.
The print technology is the last of it. Twelve of the twenty delivered documents are pure page images at 300 dpi with no text layer at all, roughly 2,432 by 3,496 pixels. Typewritten curricula from 1956 sit elsewhere in the corpus. Documents printed in blackletter type, which several states’ timetable law reaches back into, were excluded from this project outright, and that exclusion is worth saying out loud: one spec does not read every printing technology these gazettes were ever set in, and this one was never asked to.
Copying the shared value onto each member turns 4 hours into 16. Neither is what the law says.
THE SHAPE OF ONE DOCUMENT
One 19-page founding decree produced about 1,200 rows. One printed page inside it, the section listing nine special-needs school variants by three lines by up to ten grades, produced about 170 rows on its own. And one 2009 gazette issue runs to 76 pages of which about four carry the timetable regulation, the rest being two other regulations and about forty pages of certificate templates.
Rows from single named documents. Row counts vary by an order of magnitude across this corpus, and these figures are per document rather than per batch.
THE PAGES THEMSELVES
Four pages from the delivered corpus, as printed. This is what every row in the delivery traces back to.




REPRODUCED FROM THE MECKLENBURG-VORPOMMERN GAZETTES · PUBLIC REGULATIONS
REDRAWN FRAGMENT · FOUR SUBJECTS, ONE BRACE, ONE PRINTED VALUE · NOT A REPRODUCTION OF A GAZETTE PAGE
row · source page, confidence, noteStand-in imagery, not a photograph taken for this page · Photo: Coenen, CC BY-SA 3.0Stand-in imagery · Coenen, CC BY-SA 3.003 / the reading rule
The contract is two datasets and not one. The first is the timetable as printed: subjects and groups verbatim, combined grade levels such as 5–6 and 7–9 preserved as they are shown, hours exactly as printed, ambiguity indicators included. The second is the legal text around it: titles, footnotes, contextual notes and amendment references, structured so that they can be linked back to rows. There is no automatic redistribution, no normalisation and no semantic reconciliation between the two. One explicit exception exists, at the researchers’ request: a hyphen or an empty cell is recoded to 0.
The spec is the product. Seventeen global rules at version six, co-authored across six rounds of comment and revision. In product terms that spec is a Spec, defined in the platform’s pipeline: platform configuration, not bespoke software written for one corpus. The record calls it not a field list, but a codified reading of how German timetable law is typeset, and that is the honest description of what changed hands here.
The customer’s vocabulary wins, every time. Subject names are never normalised, because the variation is itself historical data: geography alone has three spellings inside one state’s twenty documents, work and technology has six, religion and ethics has five. The group separator had to become a plus sign rather than a slash, because a subject name such as work / economics / technology already contains a slash. That looks cosmetic. It is actually about whether the boundary between group members survives the trip.
Two boundaries are design. Mapping raw subject names onto research categories is a research decision and not an extraction one, so it stays with the researchers. One regulation is never reconciled against another, by rule, which is the zero the band’s RESOLVE stage prints.
This corpus saturated the general-purpose single pass. So the platform was configured for it, rather than the corpus being forced down the default path.
GR-03
2/3, +2/3, 1,5 and (4) are kept exactly as printed. Only a dash or a blank becomes 0, and that one recoding was the researchers’ own request.
GR-01
A figure given for a phase stays on the phase. An early run distributed one across years and invented a school year that state does not have.
GR-07 · GR-11
An absorbed subject gets no row. A subject whose cell is garbled gets a row with 0 hours and a confidence below 40. An omission and a zero must never look the same in the output.
GR-15
When the cells and the printed total disagree, record the cells as printed, drop confidence to between 50 and 70, and state the contradiction in the note.
GR-13
No reconciling against other documents. What comes out is a legal record and not a reconciled view of one.
GR-17
Rows come out in the source’s own subject sequence, table by table and section by section, so a reviewer can read the scan and the sheet top to bottom in step.
04 / the validator
Every printed timetable regulation carries its own checksum. It prints a total column per subject and usually a total row per grade, and that turns "is this table correctly aligned" from a judgment into a sum. The check costs nothing, it runs before any structuring, and it is deterministic. Ground truth is normally the expensive part of this work. On this corpus it is printed on the page.
The checks below are not code written for one corpus. They run as agentic validation in the same pipeline the Spec is defined in, so the rule about how a page should be read and the arithmetic that holds the reading to it are a single configuration.
Cell-count
A data row with fewer cells than its header. It catches any dropped blank cell, including in tables that print no total at all.
Row-sum
Numeric cells that do not add up to the printed total. It catches shifted, duplicated and misread values.
Coverage guard
Output far below the rows estimated from table-block and heading counts. It catches a 76-page complete re-issue producing one row, or a 24-page amendment producing zero.
The brace detector
Label-less value rows left behind by a flattened brace. It appends an automatic note and caps confidence, and it never edits hours.
ON THE FOUNDING DECREE OF THIS BATCH
Twenty-four recovered tables, seven flagged by the deterministic cell-count and row-sum checks before any structuring, including both of the two tables the researchers had already reported by hand.
A flag rate, on one document.
ON THE HEALTHY BATCH
The coverage guard estimated 1,192 rows against 1,197 actually delivered, a divergence of 0.4 percent, while correctly ignoring genuinely table-free documents such as a two-line correction notice.
A divergence between two row counts.
05 / the argument
Repair is delegated to a re-read of the source. Arithmetic cannot tell a dropped duplicate from a dropped blank, so "insert one blank cell here" can be entirely plausible and still be wrong, and a plausible wrong answer is the exact failure class this whole design exists to remove. The validator therefore refuses. It does not tidy.
When a document fails validation on this pipeline the original page images are read again at high resolution, the whole document and never a page range, because tables span pages and because sending the whole document keeps the printed total row and column in the same request as the data they validate. The re-read records the page each row was read from, which is why page is observed rather than inferred. Choosing what does the re-reading is a platform capability: when complex image assets are detected, Talonic routes them to the vision model its own benchmarks rank best for that kind of page. That is a capability of the platform.
Every row carries a confidence between 0 and 100, scored on that row’s own evidence and tied by rule to what its note says, and an extraction_note in plain language. A review threshold means a reviewer reads the flagged rows rather than all of them. The researchers asked for an alert when the system was not sure, and these three fields are the answer to that request.
Three arithmetic contradictions were found during QA in the sources themselves: printed grade values totalling one number, the printed per-subject total column totalling another. They were recorded exactly as printed, with a confidence between 50 and 70 and a note stating the contradiction, and they were raised with the researchers as a question. They were not corrected. This is the paragraph the case exists for.
A pipeline that "fixes" them would be falsifying the legal record; one that ignores them hides information the researcher needs.
06 / next
You already know which folder it is. Send twenty documents from it: scans with no text layer, tables that span pages, footnotes that live three sections away from the table they govern. We run them through Talonic and send back structured data within five business days, with a source page, a confidence score and a note on every value we were not sure about. Judge it row by row against the page it came from.
SIX CASE STUDIES: INDUSTRIAL ENERGY · GETEC / PHARMA WHOLESALE · PHOENIX PHARMAHANDEL / LOGISTICS · BRIDGEWAY / AUTOMOTIVE · MARUTI SUZUKI / RESEARCH · WZB / INDUSTRIAL HYDRAULICS · BOSCH REXROTH
THE LEDGER
The argument ends here. What follows is the working: every population, every date, every denominator, and the places our own account needs checking. Nothing below is needed to understand the case. All of it is needed to check it.
L1 / what we got wrong
01
The first schema asked for one fully resolved row per school form, per class and per subject, with hours distributed across grade ranges, footnotes interpreted into values, and amendments reconciled against the regulations they modify, all in a single structuring pass. It was built. It was tested through several iterations on samples from three states.
02
Asked to extract, calculate and reason about legal scope at the same time, the structuring engine loses track partway through: values drift between runs and reproducibility fails. The researchers’ line-level feedback showed the errors were unsystematic, which is the signature of an extraction step carrying too much logic at once. Reproducibility was the team’s core trust test, and it was the test that failed.
03
An early iteration distributed a phase allocation across the years inside the phase, and in doing so invented a thirteenth school year that Thuringia does not have. The researchers’ correction became rule GR-01, and the rule is one word long: never distribute.
04
A prose preamble that echoed a template string once made a chunk’s structured output unparseable, and 402 rows across three sections were dropped without an error. The parser now tries successive starts, retries a failed chunk once, and aborts rather than returning partial data. Those 402 are rows lost by a bug that was found and fixed. They are not rows missing from a delivery.
05
The researchers reviewed two of the twenty delivered files in July 2026 and raised four defects. The August redelivery recovered a 2009 complete re-issue from one row to 157 rows, added 123 rows to a 2003 amendment’s section 20, and changed 35 individual hours values, each one spot-checked against the scan. Row totals across the eighteen unreviewed files went from 3,056 to 3,179.
06
Two questions were open with the researchers as of 10 August 2026: what to do about the arithmetic contradictions in the sources, and how the correction notice document type should be treated. Both are recorded as open questions with the researchers at that date.
L2 / the receipts
The figures below are engineering and verification measurements, all of them Talonic’s own.
Three of these are row counts on different populations and they are not interchangeable. 3,179 is rows across eighteen files. 3,056 is the same eighteen files before the August fixes. 1,197 is rows delivered on one document. Nothing on this page divides one by another.
The acceptance bar.
A numeric acceptance bar for this batch was agreed with the research team in January 2026. The bar and its measurement stay with the research team.
The research itself.
The study is the researchers’ to publish. This page describes documents and a method; what the dataset shows stays with the research team.
Quotations.
The correspondence in the record is an English rendering of exchanges conducted in German, so no quotation from it runs here.
Names.
No individual is named on either side, and the partner institute in the research collaboration stays unnamed.
The rest of the corpus.
More than a thousand documents across sixteen states are targeted for the bulk run that follows acceptance of this batch.
SUMMARY
FOUR CLAIMS
Four things this engagement demonstrates, each with the figure that backs it. Every one of them is checkable against the ledger above.
The transcription preserves the document’s own wording. Nothing is normalised or interpreted between the page and the dataset.
Transcribed as printed · no interpretationThe document prints its own checksum, so whether it was read correctly is arithmetic rather than opinion.
Checksum · arithmetic, not opinionSource page, confidence score, and a note on anything ambiguous, on every row, so a researcher can cite the value and check it.
Source page + confidence on every rowWhat we got wrong has its own ledger section, protected from later cuts, alongside the receipts and what is withheld.
L1 · what we got wrong