Explainer
PDFs Are Like Cockroaches

The Respect
Why They Survive
The PDF earned its immortality honestly. It renders identically on every machine, prints exactly what it shows, holds a signature, and satisfies the people whose approval publishing actually needs — lawyers, regulators, auditors, records managers. Any pitch that begins “kill your PDFs” is selling something and skipping the part where the general counsel says no.
So keep them. The problem was never that PDFs exist; it's the job they've been given. A PDF is a rendering — a faithful photograph of a document at a moment in time. Most organizations use that photograph as the master: the thing that gets stored, emailed, uploaded to the website, and asked, years later, to answer questions it was never built to answer.
The Problem
A Photograph Answers No Questions

A PDF knows where the ink goes and nothing else. It doesn't know that paragraph 4.2 is a definition, that the same clause appears in three other documents, or that a newer revision supersedes it. Every machine that touches it — search engines, screen readers, and now the language models your readers increasingly ask instead of reading — must guess at structure the document refuses to state. That guessing is where answers go wrong, and it's why content built for humans and machines has become the dividing line between libraries that get cited and libraries that get skipped.
The practical path has two lanes, run in parallel. New content moves to a structured source that renders PDF as one of its outputs — pages for the people who need pages, HTML and APIs for everything else. And the legacy library gets ingested into a delivery platform that serves old PDFs like modern documents: rendered page-by-page so thousand-page files open instantly, searchable, grouped into the collections they travel with, addressable by API. The cockroaches live on — they just stop running the building.
FAQ
Questions We Hear
Are PDFs obsolete?
No, and pretending otherwise is a sales pitch. PDFs print faithfully, sign cleanly, and read the same everywhere, which is why lawyers, regulators, and archives keep choosing them. What they are is a terminal format: excellent as a final rendering of content, poor as the way content lives and travels. Publish PDFs as an output; stop using them as the source.
Why are PDFs bad for AI and search?
A PDF is effectively a picture of a document: it knows where the ink goes, not what anything means. There is no reliable structure saying this is a clause, this is a definition, this table's third column is a date — and no context connecting it to the documents around it. Machines can extract text from it, but extraction is guesswork where structured content is knowledge, which is why systems grounded on PDF libraries hallucinate structure that structured sources provide for free.
Can existing PDF libraries be salvaged without redoing everything?
Yes — this is the redemption path. A delivery platform can ingest legacy PDFs and serve them like modern documents: rendered page-by-page so even enormous files open instantly, searchable, organized into collections with the documents they travel with, and addressable through the same APIs as structured content. The PDFs don't become structured, but they stop being second-class citizens while the structured future is built around them.
What should we publish instead of PDFs?
Both — from one source. Author structured content once, then render the PDF for the people who need pages and serve HTML and APIs for the people and machines that don't. The mistake isn't producing PDFs; it's producing only PDFs, from a source that is itself a document instead of data.
Get In Touch
A Library Full of Photographs?
We build the two lanes: structured sources that render your PDFs as outputs, and delivery platforms that make the legacy library behave like it was born modern.