LeafLab turns package leaflets into validated electronic product information — sections, images and QRD codes included. The conversion is rule-based and deterministic: the same file always produces the same result, and every decision can be traced to a rule.
The approach
A package leaflet already carries its own structure: headings, numbered sections, tables that pair a step with its illustration. Most conversion tools discard that structure by flattening the document to plain text — and then ask a language model to reconstruct what was thrown away.
LeafLab reads the structure instead of guessing it. Measured against a corpus of 53 real leaflets, this removed the language model from the conversion path entirely, while producing more sections, preserving the original headings and running roughly sixty times faster.
The same input always yields the same output. No sampling, no temperature, no drift between runs or model versions.
Every decision maps to a named rule. When a section is placed where it is, the reason can be stated — not inferred from model weights.
Output is checked against the source document, not only against a schema. That difference matters more than it sounds.
How it works
Word files are parsed directly from their XML — paragraphs, styles, tables, embedded images and their relationship IDs. PDFs go through layout analysis; scanned documents through OCR.
Four signals in order of reliability: named style, full bold, font size above the document median, and section numbering. Necessary because leaflets are frequently formatted by hand rather than by template.
Each image is tied to its text position through the document's own relationship IDs, not through file order. Sliced illustrations are reassembled; duplicates from Word's fallback branch are discarded.
Sections are matched against the standard wording of the EMA QRD template — 253 patterns across 26 EU languages plus Turkish. Section numbers serve as a second signal, since Word often renders them automatically and they never appear in the text.
FHIR R4 output, checked in three stages: well-formedness, XSD conformance, and resolvability of every image reference.
Validation
Three defects found during development passed every conventional check. XSD validation, well-formedness and image counts were green in all three cases — the documents looked complete.
| Defect | Consequence | Found by |
|---|---|---|
| Images shifted by two positions | Step 3 of a reconstitution procedure showed the illustration for step 1 | source comparison |
| Vector drawing silently dropped | A complete preparation guide vanished; the document still appeared intact | source comparison |
| Section codes on the wrong lines | An ingredients line was labelled “What X is and what it is used for” | template check |
LeafLab ships a verification tool that compares the converted result back against the original Word file — position by position. It is the only check that catches a misplaced image, and it is the reason we know these three defects existed at all.
Every action recorded with actor, timestamp, IP and context — uploads, conversions, corrections, approvals.
Reviewer and approver sign separately, each re-authenticating at the moment of signing.
A preflight tool records which versions of every component are actually in use, and writes the result as a machine-readable record.
PHP and MariaDB, no external services required. Word processing needs no Python and no network access — documents never leave the server.
Coverage
Section headings are recognised through the standard wording of the EMA QRD template, extracted directly from the official version 10.4 files rather than reconstructed by hand.
Each of the 253 patterns was verified against all 156 headings across all languages: none matches a section it does not belong to. The document language does not have to be known in advance.
Get started
The most useful evaluation is a batch of your own leaflets. We run them through and report what the verification tool finds — including anything it flags.