Skip to content
SitesProof

The free scan

What the public scanner runs, the one-scan-per-URL-per-hour rule, how robots.txt is honoured, and what the shareable report link does and does not carry.

By The SitesProof team · Updated

The free scan is the whole product on one address, with nothing set up: paste a URL, pass a captcha, get a report on a link you can send to anyone. No account, no card, no call.

What it runs

All six checks run on the crawl, exactly as they do for a monitored site. Two of them have less to say than they would in an account, and it is worth knowing why:

  • Payment-page scripts has no approved-script list yet, so every script on a payment page is reported as unapproved. That is the inventory, not an accusation — in an account you approve the ones you expect, and from then on the check reports changes.
  • Known plugin vulnerabilities needs a slug and a version to match against an advisory. It runs, but a site that is not WordPress, or that does not expose versions, produces nothing from it.

Everything else — cookies before consent, accessibility, headers, TLS, legal pages — behaves the same as it does on a paid plan.

The rules we hold ourselves to

A public scanner points our browser at a stranger's website, so it is deliberately polite:

  • One report per address per hour. Ask again inside the hour and we say so and tell you when to come back, rather than crawling the site twice. The hour belongs to a report you actually got: an attempt that failed costs you a minute of quiet, not the hour, and while a scan of that address is still running we ask you to wait for it instead of starting a second one.
  • robots.txt first. We read the site's robots.txt before anything else and obey the rules that apply to our own user agent. If it asks crawlers not to fetch that address, we do not fetch it. If robots.txt cannot be read at all, we treat that as "do not crawl", not as permission.
  • We say who we are. Every request carries the SitesProofBot user agent with a link, so a site owner can identify us and rate-limit us.
  • Only public addresses. The scanner takes http and https URLs that are reachable on the public internet. Other schemes, and addresses that resolve into a private or internal network, are refused before any page is rendered.
  • A captcha that fails closed. If the captcha cannot be verified, the scan is refused rather than let through.

A finished scan lives at a /r/ link with a random twelve-character slug. That link is the only thing protecting the report, so treat it as the secret it is — and we keep those pages out of search engines entirely.

The public report carries the observations, their severity, the standard each one relates to, and suggested wording for whoever maintains the site. It deliberately carries no evidence detail: no page bodies, no element selectors, nothing that would make a link anyone can open into a copy of someone's website.

Open that link before the crawl has finished and the page waits with you: it says how long the crawl has been going and what it looks at, then shows the report by itself when it is ready. There is nothing to reload. If it is still not there after a few minutes, the page stops checking and says so rather than spinning forever.

If a scan cannot be finished, you get a page that says so. We would rather show you nothing than a report we cannot stand behind.

Turning it into monitoring

A scan is one moment. The point of an account is the second scan: weekly crawls, a diff against the previous one, and a client-ready PDF. That needs the domain verified — how to verify a site.

FAQ

Questions, answered

Do I need an account for the free scan?

No account and no card. There is a captcha, and one scan per address per hour.

Why was my scan refused?

Either a report for that address was made in the last hour, a scan of it is still running, its robots.txt asks crawlers not to fetch that address, robots.txt could not be read at all, or the address is not reachable on the public internet.

Who can see my scan report?

Anyone with the link, which is twelve random characters. The pages are kept out of search engines and carry no page bodies or element selectors.

Start with one page.

Paste a client’s checkout address and see what a crawl reports. No account, no card, and the report has a link you can send to anyone.

Scan a page free