Methodology
Versioned, published before any score, and argued with in public. Currently
0.5.0.
Conflict of interest, first
This census is run by Radixia, a commercial AI and cloud
consultancy, and radixia.ai is included in the measured population
rather than excluded from it. We sell services related to the thing being measured and have an
obvious interest in the subject appearing important. Prior work in this space tends to exclude the
authors' own properties; excluding ourselves precisely where we happen to score well would read as
less honest, not more. Discount our own row accordingly. Everything needed to check it is public.
Our own row has been wrong twice, and both are recorded in the repository. A catch-all made a
correct absence look like a broken document; our two server cards sat on paths that no current
specification names. Both are fixed, and fixing them is the only reason we now publish
/.well-known/ai-catalog.json, which the census had been measuring other people
against while not publishing it ourselves.
What is measured
For each domain in a frozen population: could an agent, starting from nothing but the domain name, discover and connect to an MCP server for that brand? It is a census, not a site audit, and emphatically not a security assessment.
Where the check identifiers come from
D1, F1, Q1 are ours. No specification assigns them and
nobody outside this project uses them. They exist because the identifiers ship as columns in the
published dataset and have to survive revisions of this document, so they are deliberately dull
and deliberately stable.
The authority, where there is any, belongs to the thing measured and never to the label. The table below gives it per check, and it ranges from a MUST in RFC 9728 to nothing at all. Citing "a D4 failure" as though it were a standard designation would be citing us.
- D, discovery.
D1toD6, in dependency order: find a document, find an endpoint, connect, list what is there. Each depends on the one before it, which is why a later check so often reports a skip. - Q, quality of the tool surface once an agent is connected. One check today.
- F, fallbacks and posture. The two things that help an agent that never speaks MCP at all: text an agent can read, and whether a crawler is welcome.
The original brief also specified an S1 for shadow servers. It is not
in this list because it turned out not to be a check: it measures the registry against a domain
rather than measuring the domain, so it lives in a separate pipeline. The gap in the sequence is
deliberate and recorded here so nobody looks for a missing column.
What a negative result means
A check reports pass, fail, skip or error. Fail is the one that can be misread, so from methodology 0.3.0 every failed candidate check also records why, from a closed vocabulary, worked out from responses we had already received.
| Outcome | Meaning | What you may conclude |
|---|---|---|
| absent_at_every_candidate | Every candidate answered 404 or 410. | The document is not at any path we publish. Still not proof of absence. |
| inconclusive_blocked | A candidate answered 401, 402, 403, 407, 429, 451 or 5xx, or failed at the transport. | Nothing about the domain. We were refused, or the server broke. |
| invalid_document | Something was served and did not parse. | A publisher meant to do this. A conforming client cannot read it. |
| mixed_negative | Negative, but not uniformly, and nothing was blocked. | Read the per-candidate evidence. |
Until 0.3.0 the code recorded every non-2xx as not_found. That was a measurement
error rather than a wording one: a 403 and a 404 support opposite conclusions, and collapsing them
allowed "we were refused" to be published as "they have nothing". Re-reading the frozen evidence
from the first full census under the new taxonomy moves 328 of 5,283 D1 failures out of absence,
145 of them in the Absent band, and finds 630 domains that served something which did not parse.
The published bands are not restated. Scoring reads the four statuses and never these labels, so no score moves; the reclassification is published beside the run it describes and the release stands as issued.
402 counts as a refusal because it is common here: hosting that has been suspended answers every path with it, and reading that as absence blames the brand for a billing dispute.
Scoring, in one sentence
A domain earns 70 of 100 points for being connectable at all, and the remaining 30 for the quality of what an agent finds once connected. Discovery is weighted far above everything else because that is the census question.
Discovery tiers are exclusive: a confirmed handshake is 70, a published discovery document 35, and an endpoint-shaped 405 only 20. A 405 is consistent with any POST-only endpoint, so on its own it is a hint rather than a finding.
When we refuse to score
A domain gets no score at all, and specifically not a zero, when we were not permitted or not able to look. Recording our own exclusion as a negative would publish a finding about our crawl dressed up as a finding about somebody's site. Unassessed domains are their own category and are excluded from every denominator.
Limitations that weaken our own findings
The full list lives with the methodology in the repository and is not a formality. The ones that
matter most: a negative is not proof of absence, because a brand may be reachable through a
marketplace or private agreement; we see one vantage point, so bot mitigation makes some rows
time-dependent; a D4 failure is inconclusive, since RFC 9728 binds only servers that
implement authorization and authorization is itself optional; and the population is
registry-derived, so it measures organisations that both run and registered a server.
The full methodology · what the specification actually says today · the decisions and why