Document remediation

Tags added. Layout untouched.

For eligible, simple born-digital PDFs, the engine adds a reviewable structure tree only after every rendered page agrees with the canonical reading order. Complex, scanned, or OCR-derived files are refused; every candidate proves that not one drawing instruction changed.

No account needed: 5 documents per day, up to 20 pages each. In the full product the hosted accessible HTML page ships alongside it.

Before: untagged

A screen reader hears an unbroken text stream. No landmarks, no headings, no alternative for the image.

After: tagged in place

/Alt "Quarterly totals chart"
TH /Scope Column
page footer, decorative rule

Same pixels, real structure: headings to jump between, described figures, tables with header meaning, and page furniture kept out of the reading order.

How it works

Structure first, then proof.

Rebuilding a PDF to tag it throws away the document people recognize. Tagging it in place keeps the file you published and adds the layer assistive technology needs.

  1. 01

    Prove the source is eligible

    For a new tag tree, the engine inventories every page and rejects scans, sparse pages, OCR-derived text, existing tags, raster images, forms, tables, formulas, unsafe links, incomplete geometry, and blocking graph findings. Already-tagged PDFs enter a separate inspect, patch, or explicit regeneration workflow.

  2. 02

    Verify the rendered layout

    One visual pass reconstructs each rendered page; a separate pass compares the exact graph projection that would become tags. Text sequence, every number, semantic type, and reading order must agree. Provider failure or incomplete output abstains.

  3. 03

    Tag in place

    Marked-content boundaries are injected around the existing drawing instructions and wired into a new structure tree. Nothing is redrawn, rasterized, or reflowed. Repeated headers, footers, and decorative art become artifacts.

  4. 04

    Verify before shipping

    The candidate must keep every drawing operator identical, pass a structural audit, prove its text survives tag-tree re-extraction, and retain the independent veraPDF result without turning it into a conformance claim.

  5. 05

    Bind the exact candidate

    Source, canonical graph, and tagged output receive separate SHA-256 identities. The candidate remains release-blocked until a person completes the exact-version keyboard, zoom/reflow, screen-reader, and source-comparison tasks.

  6. Two deliverables from one pipeline

    The canonical evidence graph projects both hosted HTML and the tagged-PDF plan. AI supplies visual evidence; it never becomes a second tag or release authority.

Honest automation

It refuses before it guesses.

Typed refusals

Scans, OCR layers, encrypted files, images, tables, forms, formulas, unsafe links, and unreadable pages produce ordered reason codes before a tag is generated. Existing tag trees switch to the separate constrained editor instead of being overwritten.

No silent claims

Automated output never carries a PDF/UA conformance claim. That claim requires human review of reading order, tables, and the actual assistive-technology experience, and it stays with the expert remediation path.

Uncertainty is visible

Merged table headers, inferred heading levels, and content without a confident match are listed with page and position in the review queue. Nothing ambiguous is resolved by guesswork.

Questions

How automatic tagging behaves

Does automatic tagging change how my PDF looks?

No. The tagger only inserts marked-content boundaries and a logical structure tree around the drawing instructions that are already in the file. Before any result ships, the drawing operators of the tagged file are compared against the original, operator for operator. If anything about the visible page would differ, the candidate is discarded.

Does this make my PDF PDF/UA conformant automatically?

No, and automated output never writes a conformance claim. Even a machine-validated candidate still needs exact-version keyboard, zoom/reflow, source-comparison, and representative assistive-technology review. The output is review infrastructure, not certification.

How is the accuracy of the tag structure checked?

The engine first restricts automation to simple authored text. It reconstructs every rendered page, separately compares the proposed graph projection with that page, and requires exact text, numbers, structure, and reading-order evidence before tag generation. The candidate must then preserve every drawing operator, pass structural and round-trip checks, and record veraPDF results. Human approval is still required for the exact output hash.

What happens when a PDF is already tagged?

The workspace first opens the authored tree without replacing it. Native P and H1-H6 roles can be corrected in place, or the user can explicitly request a visual-AI regeneration. Regeneration renders every page, builds a new proposal, runs a separate visual review, compares source text and numbers, and writes only to a new copy when every strict gate passes. Low-confidence, disputed, table, figure, form, formula, or other complex regions return review or specialist guidance instead. Every copy remains blocked for human and assistive-technology review.

What about scanned documents and forms?

They are rejected before tag generation. The same is true for pre-existing OCR layers, raster images, tables, figures, formulas, rotated pages, and other complex semantics in this first strict service. Those files need reviewed OCR or specialist remediation; the engine does not lower the gate to increase coverage.