Free tool · No account · Nothing to install

Find every PDF on your website.

Most teams cannot answer the first question a document accessibility plan needs: how many PDFs are published, and where are they linked from. A crawler can answer it in about a minute, because the files are already public.

You get the full list, not a teaser: every file the crawler reached, the pages linking to it, and an automated check on the ten most-linked documents.

Public pages only. The scanner obeys robots.txt, identifies itself, and never signs in, guesses URLs, or stores your documents.

One live crawl per domain per day. If somebody scanned this site today, you get that result instead of the site being read twice.

Count, claim, switch

Only the last step needs anything installed on your site.

01This tool

Count

A crawler reads your public pages and lists every PDF they link to. No install, no account, no permission needed to read what is already published.

02Free plan, 1 domain

Claim

Verifying the domain turns a one-off count into a monitored inventory: more pages, repeat checks, source-change alerts, and conversion into private drafts.

03When you are ready

Switch

The observe-only script points an existing PDF link at the approved accessible version, and finds the files a crawler cannot see. It changes nothing until you publish.

What the report gives you

  • Every public PDF the crawler reached, with its URL
  • The link text a visitor actually sees
  • How many pages link to each file
  • Tags, title, language and page count on the sample
  • A shareable link that lasts seven days

What a crawler cannot see

  • Links added by JavaScript after the page loads
  • Anything behind a login or an intranet
  • Files nothing on the site links to
  • Pages the site's robots.txt puts off limits

The monitoring script covers those, and is also the only way to switch a live link to an accessible version.

What it will not claim

  • That your site is or is not accessible
  • That an unopened file has any particular problem
  • Conformance with WCAG, PDF/UA, or Title II

Counts describe the files this scan opened. Read what to do with them.

How the crawler behaves

We ask crawlers to respect our rules, so we respect yours.

The scanner identifies itself as DocAccessible-SiteScan with a link back to this page, obeys robots.txt including a rule aimed only at that name, paces its requests, follows links rather than guessing URLs, and never signs in or submits a form. It reads at most 50 pages of a site and opens at most 10 documents.

One live crawl per domain per day is shared by everyone who asks about that domain, so the site is read once no matter how many people run the tool. If you would rather we did not read your site at all, a robots.txt rule is enough, and we will also add your domain to a refusal list on request.

A checked PDF is downloaded, scanned for malware, audited out of process, and discarded. What is kept is the list of public URLs and the findings, for seven days.

Questions

What the scanner does, and what it refuses to do

How does it find the files without a plugin or a login?

It reads the website the way a search engine does: it starts at the homepage and the sitemap, follows links on that domain, and records every link or embed that points at a PDF. Nothing is installed, and nothing is guessed at. A URL nothing links to is not found, because looking for it would mean probing a site rather than reading it.

Do you respect robots.txt?

Yes, and more strictly than the monitored crawl does. If a site's robots.txt disallows the root for our crawler, the scan is refused before any page is requested. Individual disallowed paths are skipped and counted in the report. A site can block this tool alone with User-agent: DocAccessible-SiteScan.

How much traffic does a scan send to the site?

At most 50 pages plus up to 10 PDF files, paced with a delay between requests, and one live crawl per domain per day no matter how many people ask. A second request inside that window is answered from the first scan's results rather than by crawling the site again.

Why did it only check ten documents?

Opening a file costs the site bandwidth and costs us processing, so the free scan opens the ten most-linked documents, which are the ones most visitors actually meet. Every other file is listed with its URL and where it is linked from, and the report says plainly which ones were opened.

Does a clean automated result mean the document is accessible?

No. Automated checks find missing tags, missing titles, missing language, and structural failures. They cannot tell whether the reading order matches the page, whether a table's meaning survives, or whether a description is useful. Those need a person, and the report says so next to every count.

Can I scan a website I do not own?

The scanner only reads pages that are already public, so technically yes, and the report is about published URLs rather than anything private. If you run the site, connecting the domain is better: an owner-verified crawl reads more pages, monitors for changes, and can turn a file into a remediation draft.

What does the site script do that this does not?

Two things. It sees links a crawler cannot: pages built by JavaScript, pages behind a login, and files linked only from a search result. And it is the only way to switch a live PDF link to an approved accessible version, which no crawler can do from the outside.

What do you keep?

The list of public URLs, the automated findings, and a salted hash of the requesting address for abuse reports. The report expires after seven days. No PDF from the scanned site is stored: a checked file is downloaded, audited, and discarded.

How do I ask you not to scan my domain?

Add a robots.txt rule for DocAccessible-SiteScan, or contact us and we will add the domain to the scanner's refusal list, which takes effect immediately and applies to everyone.

Already know the file you need to fix? Convert one PDF into an accessible web page, free, without an account.

Open the PDF to HTML converter