Explainer
What Drives S1000D Conversion Cost

The Drivers
Page Count Is the Multiplier, Not the Driver
Identical page counts can differ in conversion cost by an order of magnitude, which is why nobody honest quotes per-page sight-unseen. What actually moves the number: source structure — validated SGML at the cheap end, word-processor drift in the middle, scanned PDF at the top. Exception density — the tables, math, and foldouts that resist automation. Illustrations — converted, redrawn, or ingested with hotspots preserved, a budget line of its own. The business-rules bar — every rule is a per-module acceptance check. Splitting judgment — how documents become modules. And the rework loop — who fixes rejects, on whose clock.
The Discipline
Evidence Beats Estimates

The method that keeps programs off the cost-growth treadmill is unglamorous: sample first. A representative slice — worst tables, oldest files, real graphics — through the real pipeline, before commitment. Out come the automation rate, the exception catalog, and honest per-module effort: an estimate that survives contact with the library, instead of a guess that renegotiates itself at month four. Your levers on the buyer side are just as concrete: scope with the convert-or-ingest split so the settled tail never enters the pipeline, settle the governance decisions once, and purge the duplicates and superseded variants that convert at full price for zero value.
We're publishing this because the silence around conversion cost serves vendors, not buyers. A client who understands the drivers asks better questions, scopes tighter programs, and — yes — negotiates harder. We'll take that trade: the programs that go well are the ones priced from evidence, and evidence is the part of our quote we're proudest of.
FAQ
Questions We Hear
Why won't anyone just publish a per-page price?
Because the honest answer is that identical page counts can differ in cost by an order of magnitude. A thousand pages of validated SGML with clean tables converts on rails; a thousand pages of scanned PDFs with hand-drawn schematics and undocumented variants is a re-authoring project wearing a conversion costume. A vendor quoting per-page sight-unseen is either padding heavily against the unknown or planning to renegotiate mid-project — you pay for the blindness either way. The fixable problem is the blindness, which is what sampling exists for.
What are the cost drivers, ranked?
In the order they usually dominate: (1) Source structure — validated SGML/XML at the cheap end, word-processor files in the middle, PDF and scans at the top. (2) Exception density — the tables, math, foldouts, and special structures that resist automation; their rate matters more than page count. (3) Illustrations — converted, redrawn, or ingested, with hotspot-linked graphics a category of their own. (4) The business-rules bar — how strict the target BREX is, since every rule is a per-module acceptance check. (5) Content splitting decisions — how monolithic documents become modules, which is judgment work. (6) The rework loop — who fixes rejects and how fast. Page count is the input to none of these; it's just the multiplier on all of them.
How does sample-first pricing work?
A representative slice — deliberately including the worst tables, oldest files, and hairiest graphics — runs through the actual conversion pipeline before anyone commits. Out come three numbers: the automation rate (what converts untouched), the exception catalog (what breaks, and whether consistently), and per-module effort for the remainder. Those three numbers turn a guess into an estimate that survives contact with the library. It costs a little to produce and routinely saves programs from six-figure surprises — which is why we quote from samples and decline to quote from file counts.
How do we keep conversion cost down on our side?
Three levers before any vendor arrives. Scope honestly with the convert-or-ingest split: converting the settled tail of the library buys almost nothing — convert the living families, ingest the rest. Settle governance early: the target issue, business rules, coding scheme, and re-author line, decided once, prevent the per-module relitigating that quietly doubles projects. And clean your source of known junk first: duplicate variants, superseded revisions, and orphaned files convert at full price and deliver zero value. The cheapest module to convert is the one you correctly chose not to.
Get In Touch
Price From Evidence
Send a representative slice of your library. We'll run the sample and hand you the three numbers your budget actually needs — before you commit to anything.