How conversion quality is measured

See your PDF as a web page.

Upload a PDF and watch it come back as real HTML: headings that are headings, lists that are lists, tables with header cells, and a reading order a screen reader can follow. Compare it with the original side by side, then download the page as a single file.

Free limits: the first 10 pages, 2 documents per day per network address, 10 MB per PDF. OCR runs when text recovery is needed, and a source-grounded AI comparison checks content and structure when configured. Image descriptions remain a human review task. Uploads may be sent to those processing providers. Sign up for whole documents, drafted alt text, hosted pages, and the review workspace.

Drop your PDF here

First 10 pages, up to 10 MB. 2 documents per day, no account.

What this tool does, and what it does not

Automated first pass

  • Structure read from the PDF tag tree when it has one, or inferred from visual layout when it does not.
  • OCR for pages without usable text, including a bounded redo when a recognizable existing OCR layer is poor.
  • Source-grounded AI comparison of the page images and final block model, with constrained OCR character repair.
  • Headings, paragraphs, lists, tables with header cells, figures with captions, and links.
  • A self-contained HTML file with images inlined, so it opens anywhere with no network access.

Still needs a person

  • Image descriptions. Purpose is an authorial judgement, and a guessed description that reads as finished is worse than a stated gap.
  • Confirming inferred structure against the original when the source PDF carried no tags.
  • Complex tables, uncertain column order, and unreliable scans. When fidelity cannot be established, the tool keeps the attempted text inside a collapsed diagnostic and routes the source to expert remediation.

A hosted HTML page can be live in minutes because the conversion is automated. A tagged, remediated PDF is a different piece of work and follows later, after a specialist has been through it. Anyone promising both instantly is describing something other than what automation can actually do.

Questions

How many pages can I convert?
The first ten pages of any PDF, two documents per day per network address. Longer files are not rejected: they are converted up to page ten and labelled as a partial preview.
Is the HTML accessible?
When source fidelity can be established, the result is semantic HTML with real headings, lists, tables, and figure captions. It is not a conformance claim. OCR and AI source comparison reduce recognition and structure errors, but neither replaces manual comparison with the PDF. If the system cannot preserve the source reliably, the attempted extraction is shown only inside a collapsed, explicitly unverified diagnostic section; publication and finished-HTML download remain blocked.
Do you write alt text for my images?
Not in the free tool. Figures come back with empty alternative text and an open review item. A description invented by a machine and presented as finished is worse than an honest gap, so drafting alt text is part of the reviewed path, not this one.
What happens to my file?
The PDF is processed in memory and is not retained. It may be sent to the configured OCR and AI processing providers for text recovery and source comparison. The resulting HTML, or a safe failure report with a collapsed diagnostic attempt, is kept for 24 hours so the share link works, then deleted automatically. Shared previews are marked no-index for search engines.
Why does my PDF need expert remediation?
Some documents have features automation should not be trusted to reconstruct alone: complex or multi-level tables, uncertain multi-column reading order, unreliable scans, and specialist notation. The tool tells you when it sees one rather than producing a confident-looking result you cannot rely on.