DocAccessible
Research methodology

How we measure PDF-to-HTML conversion quality

The complete evidence method for 50 exact-revision PDF-to-HTML records: automated text, visual, and markup diagnostics plus separate project-owner human review.

Updated August 6, 2026. Reviewed by the DocAccessible team under our editorial policy.

Where the benchmark PDFs came from

We identified the 50 PDFs through public web research, including results indexed by Google, and acquired each frozen research copy only from an official publisher URL. None were supplied by customers, users, or partners, and none came from DocAccessible uploads.

The corpus intentionally spans ten document classes and includes recognizable public-sector publications. It is a purposive diversity sample, not a popularity ranking or a random sample. Public access does not grant redistribution rights: the source PDFs are not republished here, and each corpus record points to the publisher's official source.

Current pilot status

As of August 6, 2026, all 50 official source PDFs are download-verified, malware-scanned, hash-frozen, and source-scored. Every HTML working conversion also completed the same deterministic markup preflight and deterministic rendered visual comparison. The project owner then completed source-grounded review and a final route decision for every exact conversion revision. Automated diagnostics and human review are reported as separate states.

Official-source candidates
50
Download verified
50
Source-scored
50
Text threshold met
23
Below text threshold
27
Median source fidelity
89.16%
Markup preflights
50
Rendered pages compared
373
Mean visual coverage
72.33%
Human reviews complete
50

Download verification means the research copy passed a PDF signature check, ClamAV scan, and basic file inspection, and its hash is frozen in the corpus register. Candidate PDFs are not republished on this website or committed to the product repository; each working conversion links to the publisher's official PDF as the document of record.

Every candidate, official source link, verification state, routing hypothesis, HTML working version, and automated source-review result is listed in the 50-document benchmark corpus. Human review and automated diagnostics remain separate states.

How the automated source review works

Method version 1 uses Unicode-normalized text extracted from each official PDF and its HTML working conversion. The frozen score is 40% token multiset F1 and 60% ordered five-token shingle multiset F1. A score of 90% or higher is labelled “Threshold met”; a lower score is labelled “Below threshold.”

  • The mean source-fidelity score is 82.7% and the median is 89.16%.
  • Five image-only archival PDFs used disclosed 150 dpi English OCR; the other 45 used their embedded text layer.
  • Source and HTML image-object counts are recorded separately. They do not affect the text score and are not treated as image-to-image equivalence.

This score detects text loss, insertion, and local-order drift. It does not establish heading semantics, table relationships, meaningful alt text, form behavior, visual equivalence, or assistive-technology usability.

Download the complete 50-document source-review results for per-document scores, component metrics, frozen hashes, extraction methods, and visual-inventory diagnostics.

How rendered visual coverage works

Method version 1renders sampled source pages with Poppler at 110 dpi and renders each HTML fragment in a fixed Chromium viewport. Network requests are blocked; locale, timezone, viewport, device scale, and font stack are fixed; animations and transitions are disabled. This makes repeated HTML screenshots deterministic within the recorded renderer version.

Ordered five-token sequences align each sampled PDF page with the corresponding HTML scroll region. The renderer then rebases that region to the viewport origin before capture, including for extremely tall reflowed documents. The weighted visual-coverage score combines:

  • 35% source text-sequence region alignment.
  • 20% normalized horizontal block-position and width alignment.
  • 25% normalized grayscale raster and edge-structure similarity.
  • 10% coverage of sampled PDF pages that contain image objects by an aligned HTML image, figure, SVG, or canvas.
  • 10% representation of planned figures or images, semantic tables, and native form controls.

All pages are sampled when a PDF has at most 12 pages; otherwise 12 evenly spaced pages include the first and last. Null inapplicable components are excluded and the remaining weights are renormalized. The run compared 373 rendered page pairs across 50 documents. Source PDFs, temporary page renders, and HTML screenshots are not published; hashes and per-page evidence are retained in the benchmark data.

Visual attention is raised when sampled text-sequence alignment is below 70%, a sampled PDF page with image objects lacks an aligned HTML visual element, or a planned visual representation is absent. The 70% value is a triage trigger for inspection, not a visual pass or conformance gate.

Download the complete visual-coverage results for per-document component scores, attention reasons, sampled page numbers, and frozen source and conversion hashes.

What the current automated preflight establishes

Method version 1 parses the committed fragment and metadata sidecar for every candidate. It records artifact hashes and inventories text, headings, lists, links, figures, tables, form controls, page sections, language metadata, unsafe markup, and planned-feature representation. The method is deterministic and network-free.

  • 36 artifacts have no issue detected by the markup rules.
  • 14 artifacts have at least one high-confidence markup attention signal.
  • All 50 were routed to human review at preflight; markup “clear” was not treated as a pass.

Current attention signals include form documents without native HTML controls, planned visual representations without image or figure markup, table representations that need semantic attention, and one conversion without a fragment heading. These initial signals remain visible after human review so the record stays reproducible.

Download the completed 50-document human-review record for exact hashes, dimension results, final routes, reviewer role, and review date. A reusable blank copy remains available as the manual review template.

Why one average score would be misleading

A clean one-column memo and a scanned archival record do not present the same conversion problem. The pilot fixes five documents in each of ten primary classes so results can be reported by document type. A strong result on simple files cannot hide a dangerous routing error on forms, complex tables, or scans.

Document classCountText threshold metFinal routePrimary risk
Simple linear documents55 / 5HTML-firstText and headings
Long structured reports53 / 5Automation + reviewHierarchy and long-range order
Multi-column brochures55 / 5Automation + reviewColumn order and callouts
List- and link-dense guides52 / 5HTML-firstLists and links
Figure-, chart-, and map-heavy files51 / 5Automation + reviewFigures, captions, and purpose
Simple tables50 / 5Automation + reviewCells, headers, and footnotes
Controlled multilingual set54 / 5Automation + reviewScripts, diacritics, CJK, and RTL order
Complex tables51 / 5SpecialistSpans and header associations
Interactive forms52 / 5SpecialistFields, labels, and tab order
Scanned archival records50 / 5SpecialistOCR, noise, and handwriting

The route mix is 10 HTML-first, 25 Automation + review, and 15 Specialist records. The original sampling hypotheses were retained in the preflight evidence; the completed project-owner review records the same route mix as the final route decision.

What the completed human review covered

Every document was reviewed against the frozen source and committed conversion. Applicable dimensions receive their own result so one dimension is not averaged away by another.

  • Meaningful-text fidelity against the frozen source.
  • Reading order, including columns, notes, and repeated furniture.
  • Headings, lists, links, and semantic relationships.
  • Visual purpose and representation where applicable.
  • Table relationships and navigation where applicable.
  • Form controls, labels, instructions, and tab order where applicable.
  • Language, glyph, and direction behavior where applicable.
  • Keyboard and assistive-technology behavior of the HTML revision.
  • A final remediation route for every document.

False-safe routing is the primary safety measure

A false-safe occurs when a document that requires human review or a specialist is allowed through the HTML-first route. The pilot treats that as more serious than a conservative escalation because unsafe output must not become publishable through a browser-side override.

HTML-first

Responsive HTML is the primary reading experience and no specialist constraint is present.

Automation + review

Automation can prepare a draft, but a person must resolve structural or contextual uncertainty.

Specialist

Forms, scans, exact-layout records, complex tables, maps, or similar constraints require specialist handling.

How the human-review evidence was frozen

  1. Freeze each acquired source by SHA-256 digest and record its page count, language, tagged state, and scan state.
  2. Compare every meaningful source block with the committed HTML and record a result for each applicable review dimension.
  3. Review keyboard and assistive-technology behavior, then assign a final route without using the automated score as the decision.
  4. Apply the category-specific table, form, visual, multilingual, and OCR checks to the relevant document classes.
  5. Freeze reviewer role, date, dimension results, route, source digest, and conversion digest in the public review record.

Publication rules

The published source-fidelity number is intentionally narrow. Any broader accuracy or accessibility claim must identify the method, sample, date, conversion revision, metric denominator, document-class breakdown, and limitations. Product tuning requires a separate frozen holdout so the same documents are not used both to improve and to validate the system.

The completed project-owner review permits these exact HTML revisions to be published, but product confidence indicators remain advisory workflow signals, not measured conversion accuracy, accessibility scores, or probabilities of conformance. Independent audit or broader accuracy claims require additional evidence. For the practical output decision, read the guide to choosing HTML, a rebuilt PDF, or specialist remediation.

Keep reading