Crawler ethics
This page is linked from every request we send. If you found us in an access log, you are in the right place.
Contact: census@radixia.ai
This is not a security scanner
We never call an MCP tool, never send an Authorization header or any credential,
never attempt authentication, never fuzz or brute-force anything, and never probe a path that is not
on our published candidate list. Every request is a plain unauthenticated read of a document you
published on purpose, at a location a specification or a public proposal told you to publish it.
How politely
- One request per second, maximum, per domain.
- At most 64 domains in flight crawl-wide, which bounds how many different sites we visit at once; yours never sees more than the rate above.
- 5s connect, 10s total timeout. Exponential backoff on 429 and 5xx, then we give up.
- One redirect hop, and only if it stays on your domain.
robots.txtrespected for every path, including.well-known.Crawl-delayhonoured when stricter than our own limit.
How to opt out
Any one of these works, and we honour it within 24 hours for your domain and all subdomains:
- Email
census@radixia.ai. One line is enough. - Disallow us:
User-agent: MCPCensusthenDisallow: / - Open a pull request against data/optouts.txt.
If you are already in a published dataset when you opt out, we remove your rows from the live site and from later releases. We will not rewrite an already-citable frozen snapshot, because that would break the reproducibility the project rests on. We note the removal instead.
Corrections
If we got something wrong about your domain we would rather hear it than not. Being publicly wrong about a named domain is the failure mode we care most about avoiding.