← Radixia MCP Census

Methodology

Versioned, published before any score, and argued with in public. Currently 0.5.0.

Conflict of interest, first

This census is run by Radixia, a commercial AI and cloud consultancy, and radixia.ai is included in the measured population rather than excluded from it. We sell services related to the thing being measured and have an obvious interest in the subject appearing important. Prior work in this space tends to exclude the authors' own properties; excluding ourselves precisely where we happen to score well would read as less honest, not more. Discount our own row accordingly. Everything needed to check it is public.

Our own row has been wrong twice, and both are recorded in the repository. A catch-all made a correct absence look like a broken document; our two server cards sat on paths that no current specification names. Both are fixed, and fixing them is the only reason we now publish /.well-known/ai-catalog.json, which the census had been measuring other people against while not publishing it ourselves.

What is measured

For each domain in a frozen population: could an agent, starting from nothing but the domain name, discover and connect to an MCP server for that brand? It is a census, not a site audit, and emphatically not a security assessment.

Where the check identifiers come from

D1, F1, Q1 are ours. No specification assigns them and nobody outside this project uses them. They exist because the identifiers ship as columns in the published dataset and have to survive revisions of this document, so they are deliberately dull and deliberately stable.

The authority, where there is any, belongs to the thing measured and never to the label. The table below gives it per check, and it ranges from a MUST in RFC 9728 to nothing at all. Citing "a D4 failure" as though it were a standard designation would be citing us.

The original brief also specified an S1 for shadow servers. It is not in this list because it turned out not to be a check: it measures the registry against a domain rather than measuring the domain, so it lives in a separate pipeline. The gap in the sequence is deliberate and recorded here so nobody looks for a missing column.

What a negative result means

A check reports pass, fail, skip or error. Fail is the one that can be misread, so from methodology 0.3.0 every failed candidate check also records why, from a closed vocabulary, worked out from responses we had already received.

OutcomeMeaningWhat you may conclude
absent_at_every_candidateEvery candidate answered 404 or 410. The document is not at any path we publish. Still not proof of absence.
inconclusive_blockedA candidate answered 401, 402, 403, 407, 429, 451 or 5xx, or failed at the transport.Nothing about the domain. We were refused, or the server broke.
invalid_documentSomething was served and did not parse. A publisher meant to do this. A conforming client cannot read it.
mixed_negativeNegative, but not uniformly, and nothing was blocked.Read the per-candidate evidence.

Until 0.3.0 the code recorded every non-2xx as not_found. That was a measurement error rather than a wording one: a 403 and a 404 support opposite conclusions, and collapsing them allowed "we were refused" to be published as "they have nothing". Re-reading the frozen evidence from the first full census under the new taxonomy moves 328 of 5,283 D1 failures out of absence, 145 of them in the Absent band, and finds 630 domains that served something which did not parse.

The published bands are not restated. Scoring reads the four statuses and never these labels, so no score moves; the reclassification is published beside the run it describes and the release stands as issued.

402 counts as a refusal because it is common here: hosting that has been suspended answers every path with it, and reading that as absence blames the brand for a billing dispute.

Scoring, in one sentence

A domain earns 70 of 100 points for being connectable at all, and the remaining 30 for the quality of what an agent finds once connected. Discovery is weighted far above everything else because that is the census question.

Discovery tiers are exclusive: a confirmed handshake is 70, a published discovery document 35, and an endpoint-shaped 405 only 20. A 405 is consistent with any POST-only endpoint, so on its own it is a hint rather than a finding.

When we refuse to score

A domain gets no score at all, and specifically not a zero, when we were not permitted or not able to look. Recording our own exclusion as a negative would publish a finding about our crawl dressed up as a finding about somebody's site. Unassessed domains are their own category and are excluded from every denominator.

Limitations that weaken our own findings

The full list lives with the methodology in the repository and is not a formality. The ones that matter most: a negative is not proof of absence, because a brand may be reachable through a marketplace or private agreement; we see one vantage point, so bot mitigation makes some rows time-dependent; a D4 failure is inconclusive, since RFC 9728 binds only servers that implement authorization and authorization is itself optional; and the population is registry-derived, so it measures organisations that both run and registered a server.

The full methodology · what the specification actually says today · the decisions and why