Skip to main content

SOCIAL SCIENCE RESEARCH · GERMANY · ONE BATCH DELIVERED

The WZB Berlin Social Science Center is a social science research institute in Berlin and part of the Leibniz Association. A research team there is building a dataset of every school timetable regulation issued by every German federal state since 1949, so that the amount of teaching time given to each subject can be compared across states and across decades. That data exists only inside legal texts, and until now researchers hand-coded it.

The document prints its own checksum. So "did we read it right" is arithmetic, not an opinion.

Talonic read one state’s batch. Twenty regulations from Mecklenburg-Vorpommern, 1996 to 2021, twelve of them 300 dpi page images with no text layer at all. They were structured exactly as printed, under a seventeen-rule spec the researchers co-authored across six rounds, into 3,179 rows that each carry a source page, a confidence score and a note on anything ambiguous. The rest of the corpus, more than a thousand documents across sixteen states, is targeted for the bulk run that follows acceptance of this batch.

24 TABLES · 7 FLAGGED · INCLUDING BOTH THE RESEARCHERS FOUND BY HAND · ↓ SCROLL TO INSPECT

talonic · deterministic check

REDRAWN FRAGMENT · § 19 SPECIAL-NEEDS SCHOOL, 1996

A redrawn German timetable fragment with an empty preparatory class column, its values slid one column left after layout recovery, and a verdict strip reading cell-count fail, row-sum fail, re-read the page.
SubjectPrep123456789Total
GermanAS PRINTED66556654447
GermanAS RECOVERED·········51
Mathematics··········
General studies··········
PE··········

THE SAME SHIFT, ON THE COLUMN TOTALS

PRINTED TOTAL ROW
1518202223
COLUMN SUMS, AS RECOVERED
1820222325

A shift of exactly one column. The record transcribes this part of the row and no more, so this part is what is drawn.

CELL-COUNT: FAIL · ROW-SUM: FAIL · ACTION: RE-READ THE PAGE

  • header carries 11 columns, the recovered data row carries 10 cells
  • the recovered row was delivered totalling 51, the printed total says 47
  • the validator does not repair. The page is read again.

A printed timetable carries its own totals: a total column per subject and usually a total row per grade. When layout recovery drops the blank cells in a column that looks empty, every value slides one column left and the sums stop matching. On the founding decree of this batch the deterministic checks flagged 7 of the 24 recovered tables, including both of the two the researchers had reported by hand. That is a flag rate, on one document.

WZB

Research · Germany

Industry
Social research
Function
Legal data collection
Solution
Regulation transcription

ONE BATCH DELIVERED · THE CORPUS GATED

3,179 rows

from 18 of one state’s 20 regulations

Note: German school timetable law from 1996 to 2021, transcribed exactly as printed. Every row carries its source page, a confidence score, and a note on anything ambiguous.

THE PLATFORM RUN

  1. READ ONCE

    20

    documents in one German state’s batch, Mecklenburg-Vorpommern 1996 to 2021, twelve of them 300 dpi page images with no text layer at all. One state’s batch, against a corpus of more than a thousand documents targeted for the bulk run.

  2. STRUCTURE

    3,179

    rows in the Timetable Data sheets of eighteen files from that batch, at the 10 August 2026 redelivery, under a seventeen-rule extraction spec at version six. Rows. Not documents and not fields.

  3. RESOLVE

    0

    cross-document reconciliations, by rule. One regulation is never resolved against another, because a reconciled view is not a legal record. Mapping subject names onto research categories is the researchers’ decision and it stays theirs.

  4. DELIVER

    1

    Excel workbook per document, three sheets: Metadata, Timetable Data, Footnotes, plus a ZIP per batch, delivered as files to the research team.

  5. QUERY FOREVER

    ONCE

    read. Zero re-reads of the scans since: every one of the six schema rounds was measured by re-running from cached layout recovery, without repeating OCR. The remaining thousand-plus documents follow in the bulk run after acceptance of this batch.

The fifth stage prints ONCE on purpose: it is the claim every page in this family exists to receipt.

Five stages counted in five different units, so nothing across this band is a rate, a ratio or a total. The twenty documents in the first stage and the eighteen files behind the row count are not the same set: the row count was reported on the eighteen files outside the two the researchers reviewed. The zero in the third stage is a design decision. The workbook in the fourth stage is a file.

READING LAW EXACTLY AS PRINTED

01 / the checksum

The page proves its own reading

Twenty regulations, 1996 to 2021, twelve of them 300 dpi page images with no text layer. The document prints its own checksum. So "did we read it right" is arithmetic, not an opinion. On the founding decree of this batch, 7 of 24 tables flagged. A flag rate on one document.

A printed timetable carries its own totals: a total column per subject and usually a total row per grade. When layout recovery drops the blank cells in a column that looks empty, every value slides one column left and the sums stop matching.

Talonic read one state’s batch. Twenty regulations from Mecklenburg-Vorpommern, structured exactly as printed under a seventeen-rule spec the researchers co-authored across six rounds, into 3,179 rows that each carry a source page, a confidence score and a note on anything ambiguous.

The check costs nothing, it runs before any structuring, and it is deterministic. The validator does not repair. It refuses, and sends the page back to be read again.

The rest of the corpus, more than a thousand documents across sixteen states, is targeted for the bulk run that follows acceptance of this batch.

7 of 24

RECOVERED TABLES FLAGGED ON THE FOUNDING DECREE OF THIS BATCH

Flagged by the deterministic cell-count and row-sum checks before any structuring, including both of the two tables the researchers had already reported by hand.

A flag rate, on one document.

3,179 is rows in the Timetable Data sheets of eighteen files from that batch, at the 10 August 2026 redelivery. Rows, not documents and not fields.

Everything below this band is the receipt for it. What arrived and why it resists, the rules the researchers wrote, the arithmetic that holds the reading to the page, and what we got wrong on the way.

02 / the corpus

The data was never missing. It was typeset.

A school timetable, literally an hours table in the law’s own vocabulary, is a table of subject by grade by weekly hours, printed inside a regulation. The regulations are what a state education ministry publishes in its gazette, and they arrive under the law’s own names: regulation, decree, circular decree, notice, amending regulation, correction notice. The numbers a researcher wants have been public for decades. They have simply been typeset rather than tabulated, and the typesetting is the work.

Reading one is not OCR. German timetables join several adjacent subjects with a brace and print one shared value for the group. The brace is a typographic glyph with no height, so when layout recovery flattens the page the shared value lands on whichever row the engine thought was nearest. Four subjects become one subject with a number and three subjects with nothing, or the value is copied onto all four. Both readings are wrong, and only one of them looks wrong.

Four more traps, one sentence each. A column that is blank on every subject row still exists as a column, because only the total row carries a value there, and dropping it slides every value one place left. Physics / biology or natural sciences is a choice between two subjects and one, not a long subject name, and only the arithmetic can tell those apart. Hours are printed as 2/3, +2/3, 1,5, (4) and 12 to 14, so the field is free text and not an integer. And one timetable in this batch is three sentences of prose telling you to replace a five with a two in one cell.

The print technology is the last of it. Twelve of the twenty delivered documents are pure page images at 300 dpi with no text layer at all, roughly 2,432 by 3,496 pixels. Typewritten curricula from 1956 sit elsewhere in the corpus. Documents printed in blackletter type, which several states’ timetable law reaches back into, were excluded from this project outright, and that exclusion is worth saying out loud: one spec does not read every printing technology these gazettes were ever set in, and this one was never asked to.

Copying the shared value onto each member turns 4 hours into 16. Neither is what the law says.

What arrived

A grid of twenty document cells, twelve drawn as dense page images and eight drawn in outline with text rules.PAGE IMAGES · NO TEXT LAYER1996 TO 2021BORN-DIGITAL · TEXT LAYER PRESENT12 OF 20 · 300 DPI PAGE IMAGES · NO TEXT LAYER
Pure page images at 300 dpi with no text layer at all, roughly 2,432 by 3,496 pixels, in the delivered batch of twenty. Documents printed in blackletter type, which several states’ timetable law reaches back into, were excluded from this project.

THE SHAPE OF ONE DOCUMENT

One 19-page founding decree produced about 1,200 rows. One printed page inside it, the section listing nine special-needs school variants by three lines by up to ten grades, produced about 170 rows on its own. And one 2009 gazette issue runs to 76 pages of which about four carry the timetable regulation, the rest being two other regulations and about forty pages of certificate templates.

Rows from single named documents. Row counts vary by an order of magnitude across this corpus, and these figures are per document rather than per batch.

One value, four subjects

Four subject rows joined by a tall brace to one shared value, beside two readings: the shared value copied onto every member, and the printed per-subject total values that prove that reading wrong.GRADES 7 TO 10PhysicsChemistryBiologyAstronomy4printed once, for the groupTHE READING THAT LOOKS FINEthe shared value copied onto each memberPhysics4Chemistry4Biology4Astronomy4what the group now claims16WHAT THE PAGE ACTUALLY PRINTSthe per-subject total column, which settles itPhysics7Chemistry5Biology9Astronomy1
The brace is a typographic glyph. When the layout is flattened it has no height, so the shared value lands on whichever row the engine thought was nearest.

THE PAGES THEMSELVES

Four pages from the delivered corpus, as printed. This is what every row in the delivery traces back to.

Printed timetable for a sports secondary school: subject rows joined by braces to one shared value per grade, with a per-subject total column
§ 10 · SPORTS SECONDARY SCHOOLThe brace, as printed: physics, chemistry, biology and astronomy share one value per grade column, while the total column carries per-subject totals of 7, 5, 9 and 1.
Full gazette page: two columns of legal prose above the § 20 special-needs school timetable with a two-level header
GAZETTE NO. 19 · 18 OCTOBER 2002Prose paragraphs first, then § 20: a two-level header nests support levels I to III over grade columns, with bracketed (9H) values and separate total columns at the right edge.
Gazette correction page: three numbered sentences amending the timetable regulation, with no table at all
CORRECTION NOTICE · 19 AUGUST 1996An entire timetable amendment in three sentences, one of which replaces a 5 with a 2 in a single cell. The cell it lands in exists only if you read the prose.
Ministry bulletin page: dozens of word-level substitution clauses, ending in a table of subjects by starting grade
FIFTH AMENDMENT · 10 JUNE 2021A quarter century later the form persists: word-level substitutions clause by clause, and then a table. Reading 2021 correctly means having read everything since 1996.

REPRODUCED FROM THE MECKLENBURG-VORPOMMERN GAZETTES · PUBLIC REGULATIONS

An enlarged redrawn fragment of a German timetable, four subject rows joined by a brace.PhysicsChemistryBiologyAstronomy4

REDRAWN FRAGMENT · FOUR SUBJECTS, ONE BRACE, ONE PRINTED VALUE · NOT A REPRODUCTION OF A GAZETTE PAGE

The curved pink and blue facade of the WZB building in Berlin, five storeys of square windows on a drumrow · source page, confidence, noteStand-in imagery, not a photograph taken for this page · Photo: Coenen, CC BY-SA 3.0Stand-in imagery · Coenen, CC BY-SA 3.0

03 / the reading rule

A brace is not a multiplier. Seventeen rules say so, and the researchers wrote most of them.

The contract is two datasets and not one. The first is the timetable as printed: subjects and groups verbatim, combined grade levels such as 5–6 and 7–9 preserved as they are shown, hours exactly as printed, ambiguity indicators included. The second is the legal text around it: titles, footnotes, contextual notes and amendment references, structured so that they can be linked back to rows. There is no automatic redistribution, no normalisation and no semantic reconciliation between the two. One explicit exception exists, at the researchers’ request: a hyphen or an empty cell is recoded to 0.

The spec is the product. Seventeen global rules at version six, co-authored across six rounds of comment and revision. In product terms that spec is a Spec, defined in the platform’s pipeline: platform configuration, not bespoke software written for one corpus. The record calls it not a field list, but a codified reading of how German timetable law is typeset, and that is the honest description of what changed hands here.

The customer’s vocabulary wins, every time. Subject names are never normalised, because the variation is itself historical data: geography alone has three spellings inside one state’s twenty documents, work and technology has six, religion and ethics has five. The group separator had to become a plus sign rather than a slash, because a subject name such as work / economics / technology already contains a slash. That looks cosmetic. It is actually about whether the boundary between group members survives the trip.

Two boundaries are design. Mapping raw subject names onto research categories is a research decision and not an extraction one, so it stays with the researchers. One regulation is never reconciled against another, by rule, which is the zero the band’s RESOLVE stage prints.

This corpus saturated the general-purpose single pass. So the platform was configured for it, rather than the corpus being forced down the default path.

Six of the seventeen rules

GR-03

Hours are recorded verbatim.

2/3, +2/3, 1,5 and (4) are kept exactly as printed. Only a dash or a blank becomes 0, and that one recoding was the researchers’ own request.

GR-01

Never distribute.

A figure given for a phase stays on the phase. An early run distributed one across years and invented a school year that state does not have.

GR-07 · GR-11

Not taught is not the same as unreadable.

An absorbed subject gets no row. A subject whose cell is garbled gets a row with 0 hours and a confidence below 40. An omission and a zero must never look the same in the output.

GR-15

The printed total wins the argument and loses the edit.

When the cells and the printed total disagree, record the cells as printed, drop confidence to between 50 and 70, and state the contradiction in the note.

GR-13

Structure only what this document prints.

No reconciling against other documents. What comes out is a legal record and not a reconciled view of one.

GR-17

Source order is the output order.

Rows come out in the source’s own subject sequence, table by table and section by section, so a reviewer can read the scan and the sheet top to bottom in step.

Seventeen rules at version six, co-authored with the researchers across six rounds. The rules are the product of this engagement: tacit reading knowledge, written down so that any future document and any future engine version is held to it.

04 / the validator

Three guards and one detector.

Every printed timetable regulation carries its own checksum. It prints a total column per subject and usually a total row per grade, and that turns "is this table correctly aligned" from a judgment into a sum. The check costs nothing, it runs before any structuring, and it is deterministic. Ground truth is normally the expensive part of this work. On this corpus it is printed on the page.

The checks below are not code written for one corpus. They run as agentic validation in the same pipeline the Spec is defined in, so the rule about how a page should be read and the arithmetic that holds the reading to it are a single configuration.

Cell-count

A data row with fewer cells than its header. It catches any dropped blank cell, including in tables that print no total at all.

Row-sum

Numeric cells that do not add up to the printed total. It catches shifted, duplicated and misread values.

Coverage guard

Output far below the rows estimated from table-block and heading counts. It catches a 76-page complete re-issue producing one row, or a 24-page amendment producing zero.

The brace detector

Label-less value rows left behind by a flattened brace. It appends an automatic note and caps confidence, and it never edits hours.

ON THE FOUNDING DECREE OF THIS BATCH

Twenty-four recovered tables, seven flagged by the deterministic cell-count and row-sum checks before any structuring, including both of the two tables the researchers had already reported by hand.

A flag rate, on one document.

ON THE HEALTHY BATCH

The coverage guard estimated 1,192 rows against 1,197 actually delivered, a divergence of 0.4 percent, while correctly ignoring genuinely table-free documents such as a two-line correction notice.

A divergence between two row counts.

05 / the argument

The validator does not repair. It refuses, and sends the page back to be read again.

Detect, do not repair.

Repair is delegated to a re-read of the source. Arithmetic cannot tell a dropped duplicate from a dropped blank, so "insert one blank cell here" can be entirely plausible and still be wrong, and a plausible wrong answer is the exact failure class this whole design exists to remove. The validator therefore refuses. It does not tidy.

What a refusal costs.

When a document fails validation on this pipeline the original page images are read again at high resolution, the whole document and never a page range, because tables span pages and because sending the whole document keeps the printed total row and column in the same request as the data they validate. The re-read records the page each row was read from, which is why page is observed rather than inferred. Choosing what does the re-reading is a platform capability: when complex image assets are detected, Talonic routes them to the vision model its own benchmarks rank best for that kind of page. That is a capability of the platform.

Saying so when we are not sure.

Every row carries a confidence between 0 and 100, scored on that row’s own evidence and tied by rule to what its note says, and an extraction_note in plain language. A review threshold means a reviewer reads the flagged rows rather than all of them. The researchers asked for an alert when the system was not sure, and these three fields are the answer to that request.

When the source is the thing that is wrong.

Three arithmetic contradictions were found during QA in the sources themselves: printed grade values totalling one number, the printed per-subject total column totalling another. They were recorded exactly as printed, with a confidence between 50 and 70 and a note stating the contradiction, and they were raised with the researchers as a question. They were not corrected. This is the paragraph the case exists for.

A pipeline that "fixes" them would be falsifying the legal record; one that ignores them hides information the researcher needs.

06 / next

Send us the twenty worst documents you have.

You already know which folder it is. Send twenty documents from it: scans with no text layer, tables that span pages, footnotes that live three sections away from the table they govern. We run them through Talonic and send back structured data within five business days, with a source page, a confidence score and a note on every value we were not sure about. Judge it row by row against the page it came from.

SIX CASE STUDIES: INDUSTRIAL ENERGY · GETEC / PHARMA WHOLESALE · PHOENIX PHARMAHANDEL / LOGISTICS · BRIDGEWAY / AUTOMOTIVE · MARUTI SUZUKI / RESEARCH · WZB / INDUSTRIAL HYDRAULICS · BOSCH REXROTH

THE LEDGER

Everything above this line, with its receipt.

The argument ends here. What follows is the working: every population, every date, every denominator, and the places our own account needs checking. Nothing below is needed to understand the case. All of it is needed to check it.

L1 / what we got wrong

The obvious schema was the wrong schema. We built it first, and it drifted between runs.

01

The wrong design, named.

The first schema asked for one fully resolved row per school form, per class and per subject, with hours distributed across grade ranges, footnotes interpreted into values, and amendments reconciled against the regulations they modify, all in a single structuring pass. It was built. It was tested through several iterations on samples from three states.

02

How we knew.

Asked to extract, calculate and reason about legal scope at the same time, the structuring engine loses track partway through: values drift between runs and reproducibility fails. The researchers’ line-level feedback showed the errors were unsystematic, which is the signature of an extraction step carrying too much logic at once. Reproducibility was the team’s core trust test, and it was the test that failed.

03

The instruction we got wrong before that.

An early iteration distributed a phase allocation across the years inside the phase, and in doing so invented a thirteenth school year that Thuringia does not have. The researchers’ correction became rule GR-01, and the rule is one word long: never distribute.

04

The silent loss.

A prose preamble that echoed a template string once made a chunk’s structured output unparseable, and 402 rows across three sections were dropped without an error. The parser now tries successive starts, retries a failed chunk once, and aborts rather than returning partial data. Those 402 are rows lost by a bug that was found and fixed. They are not rows missing from a delivery.

05

The redelivery.

The researchers reviewed two of the twenty delivered files in July 2026 and raised four defects. The August redelivery recovered a 2009 complete re-issue from one row to 157 rows, added 123 rows to a 2003 amendment’s section 20, and changed 35 individual hours values, each one spot-checked against the scan. Row totals across the eighteen unreviewed files went from 3,056 to 3,179.

06

What is still open.

Two questions were open with the researchers as of 10 August 2026: what to do about the arithmetic contradictions in the sources, and how the correction notice document type should be treated. Both are recorded as open questions with the researchers at that date.

What the second delivery recovered

Three paired horizontal bars comparing row counts before and after the August redelivery, plus a standalone count of hours values changed.3 JULY 202610 AUGUST 2026Rows across the eighteen files3,0563,179The 2009 complete re-issue1157The 2003 amendment, section 20012335individual hours values changed, each spot-checked against the scan
Between two deliveries of the same twenty-document batch, 3 July and 10 August 2026, after the researchers reviewed two of the twenty files and raised four defects. These are not independent totals and they do not sum.

L2 / the receipts

Every number here, with the thing it counts.

The figures below are engineering and verification measurements, all of them Talonic’s own.

The engagement in plain rows.

VALUEWHAT IT COUNTSWHEN
20 documentsone German state’s batch, Mecklenburg-Vorpommern, 1996 to 2021.delivered 3 July 2026, redelivered 10 August 2026
1,000+ documentstargeted for the bulk run across all sixteen states, which follows acceptance of the batch above.the corpus target
16 federal statesthe corpus target. The delivered batch is one state.the corpus target
3,179 rowsrows in the Timetable Data sheets of 18 files from the 20-document batch. Rows, not documents and not fields.at the 10 August 2026 redelivery
3,056 rowsthe same 18 files before the August fixes. The before half of one pair, never an independent quantity.at the 3 July 2026 delivery
7 of 24 tables flaggedon one document, the 1996 founding decree: 24 recovered tables, 7 flagged by deterministic cell-count and row-sum checks before any structuring, including both of the two tables the researchers had reported by hand. A flag rate, on one document.before structuring, on one document
1,192 estimated against 1,197 delivered, 0.4 percentthe coverage guard’s row estimate against rows actually produced, on the healthy batch. A divergence between two row counts.on the healthy batch
12 of 20 documentspure page images at 300 dpi with no text layer at all, roughly 2,432 by 3,496 pixels, in the delivered batchas they arrived
1 row to 157 rowsrecovered on one document, a 2009 complete re-issuebetween the 3 July and 10 August deliveries of the same batch
123 rowsadded to one section of a 2003 amendment in the same redeliverybetween the 3 July and 10 August deliveries of the same batch
35 hours values changedspot-checked against the scanbetween the 3 July and 10 August deliveries of the same batch
402 rowsrows once silently dropped by a parse bug that was found and fixed.a fixed bug
17 rules, version 6the extraction spec’s rule count at version six, co-authored with the researchersat the 10 August 2026 redelivery
6 schema rounds, 2 batch reviewsrounds of co-authoring the spec with the researchers, and two reviewed deliveriesacross the engagement
2 structuring passesone for research data and one for footnotes, run separately over the same document and joined afterwardsper document
~1,200 rows from one 19-page decreerows from a single founding decree. Row counts vary by an order of magnitude across this corpus.counted on one named document
~170 rows from one printed pagethe worst case seen, one section of the 1996 decree: nine special-needs school variants by three lines by up to ten gradescounted on one printed page
~1,118 rowsthe internal regression document’s healthy row count, re-runnable from cached layout recovery without repeating OCR. A test baseline, counted separately from the 1,197 delivered.the regression harness baseline
76 pages, about 4 relevantone 2009 gazette issue of 76 pages, of which about four carry the timetable regulation. The same issue carries two other regulations and about forty pages of certificate templates.counted on one named issue
3 arithmetic contradictionsfound in the sources themselves during QA. Recorded as printed, with confidence 50 to 70 and a note stating the contradiction, and raised with the researchers as a question. Never corrected.during QA
Confidence 0 to 100confidence is scored per row on that row’s own evidence, with a review threshold at 70. A garbled cell in a table that spans the grade gets a row with 0 hours and a confidence below 40, so an omission and a zero never look the same in the output.per row
10 to 30 minutes per dense documentstructuring run time on the densest documents, streamed so nothing is lost to timeoutsper document
4 defects raisedraised by the researchers after reviewing two of the twenty delivered files in July 2026, and fixed in the August redeliveryJuly 2026

Three of these are row counts on different populations and they are not interchangeable. 3,179 is rows across eighteen files. 3,056 is the same eighteen files before the August fixes. 1,197 is rows delivered on one document. Nothing on this page divides one by another.

L3 / withheld · WHAT STAYS WITH THE RESEARCHERS

The acceptance bar.

A numeric acceptance bar for this batch was agreed with the research team in January 2026. The bar and its measurement stay with the research team.

The research itself.

The study is the researchers’ to publish. This page describes documents and a method; what the dataset shows stays with the research team.

Quotations.

The correspondence in the record is an English rendering of exchanges conducted in German, so no quotation from it runs here.

Names.

No individual is named on either side, and the partner institute in the research collaboration stays unnamed.

The rest of the corpus.

More than a thousand documents across sixteen states are targeted for the bulk run that follows acceptance of this batch.

SUMMARY

What this case shows

FOUR CLAIMS

Four things this engagement demonstrates, each with the figure that backs it. Every one of them is checkable against the ledger above.

  1. 01

    Exactly as printed is the product

    The transcription preserves the document’s own wording. Nothing is normalised or interpreted between the page and the dataset.

    Transcribed as printed · no interpretation
  2. 02

    The page proves its own reading

    The document prints its own checksum, so whether it was read correctly is arithmetic rather than opinion.

    Checksum · arithmetic, not opinion
  3. 03

    Every row carries its provenance

    Source page, confidence score, and a note on anything ambiguous, on every row, so a researcher can cite the value and check it.

    Source page + confidence on every row
  4. 04

    Errors are printed, not absorbed

    What we got wrong has its own ledger section, protected from later cuts, alongside the receipts and what is withheld.

    L1 · what we got wrong
Document testresponse within 1 business day · your data back within 5

Response within 1 business day.