Security and your data
SitesProof crawls websites on your instructions and keeps the result. That is other people's data in two senses — your clients' sites, and your relationship with those clients — so this page says what we hold, what we deliberately do not hold, and how the one genuinely dangerous part of the product is fenced in. The formal version, including hosting, access and backups, is the security page.
What a crawl actually stores
A crawl produces one artifact per scan, and the checks read that artifact instead of the live site. It holds hashes and metadata wherever a hash answers the question:
- Cookie values are never stored. We record that a cookie called
_gawas set before consent, on which domain, and nothing about what was in it. - Page bodies are never stored. Not the HTML, not the text of a legal page — only its length, which is how we tell a real privacy notice from a placeholder.
- Scripts are stored as sha-256 hashes, which is the whole point: a hash is enough to notice that an approved script's contents changed, and it is not a copy of anyone's code.
- A script's
nonceattribute is recorded as present or absent, never its value, which is a per-request secret.
Artifacts are gzipped JSON in private object storage with no public access, reachable only by our own worker. Our retention limit for an artifact is 90 days — long enough to re-run a check against an old crawl after we improve a rule, short enough that we are not sitting on a year of other people's websites.
Your clients are your business
Client names, client logos, site addresses and findings are customer data. We never use them as examples — not in marketing, not in a screenshot, not in a demo. The product has no shared directory of scanned sites, and every database query in the dashboard is scoped to your own organisation, with automated tests that check one account cannot read another's.
Reports are white-label because the relationship is yours. Your client never gets an account here and never sees our name on the deliverable.
The public scanner is the sharp edge
A URL box that anyone on the internet can point at any address is the most dangerous thing in this product: left naive, it is a tool for reaching things that are not meant to be reachable from outside. So every outbound request it makes goes through one guarded fetch path:
- Private, loopback, link-local and other non-public address ranges are refused, including when a hostname resolves to one.
- The resolved address is pinned for the request, and re-checked on every redirect, so a name that answers differently the second time cannot be used to slip past the first check.
- Redirects are capped, response size is capped, and every request has a timeout.
- Non-
http(s)schemes are refused outright. - A captcha protects the endpoint, and if the captcha cannot be verified the scan is refused rather than let through.
The crawler is also polite by design: robots.txt is read first and obeyed for our own user agent, an unreadable
robots.txt is treated as "do not crawl", requests identify themselves as SitesProofBot with a link, and the same
address can be scanned only once an hour. More on that in the free scan.
Shareable reports
A public report link is twelve random characters — sixty bits, because that link is the only thing protecting the report. Those pages are kept out of search engines, and they carry no evidence detail at all: no page bodies, no element selectors, nothing that would turn a link anyone can open into a copy of someone's website.
The application
- A strict Content Security Policy with a per-request nonce on both the website and the dashboard: a script runs
only if we put it there. Framing, plugins and
<base>rewriting are blocked outright. - HTTPS with HSTS everywhere, plus
nosniff, a referrer policy and a permissions policy on every response. - Sign-in is email and password, with a verified address, or Google. Passwords are stored only as salted, slow hashes; we never see them.
- Logs record facts, never secrets, request bodies or payment data, and error reports are scrubbed before they leave the server.
- We do not run ad pixels or cookie-based analytics on this site, so there is no consent banner to click. Analytics is cookieless and honours Do Not Track and Global Privacy Control.
What we do not have
One person builds and runs this. There is no security team, no SOC 2 report, no 24/7 on-call rotation, and no paid bug bounty. The security page is explicit about all of it, because a trust page that only lists strengths is not a trust page.
Reporting a vulnerability
Email security@sitesproof.com with details and steps to reproduce; the
machine-readable contact is at /.well-known/security.txt. Please give us reasonable time to fix an issue before
disclosing it, do not access or change other people's data, and do not run load or denial-of-service tests. We will
not take legal action over good-faith research that follows those rules, we aim to acknowledge a report within three
business days, and we will credit you if you want us to.