Methodology

How the number is produced, and what it will not tell you

Every scorecard on this site carries a score. This page is what that score means, how to read it, and where it stops being reliable — because a number without stated limits is worse than no number.

What it measures

Out of the questions your buyers ask, how many name you

Not rankings. Not traffic. A set of buying-intent questions in a category is put to several AI engines, and each answer is read for which companies get named. The score expresses how often the domain in question is one of them.

In one line

The score is the share of category questions where an engine names the domain. Higher means your brand appears more often at the moment a buyer is deciding what to consider.

How to read it

Four bands

Treat the band as the finding and the exact number as noise around it. A 63 and a 68 are the same result.

BandWhat it means
0–20AI engines almost never name the domain. Competitors hold the category.
21–50Intermittent — usually branded or long-tail questions. Most category questions miss it.
51–80Strong presence on category questions. A frequent answer, not the default one.
81–100The default answer in its category across most of the engines audited.
Limits

What this number cannot do

Four things worth knowing before you act on one.

It is a sample

Engines do not answer the same question the same way twice. A score is a sample of something that moves, so two runs a week apart can differ without anything on your site having changed.

It is not causal

A score going up after you shipped a page is evidence, not proof. Engines retrain, competitors publish, and the answer set drifts on its own schedule.

It is category-relative

A 40 in a crowded category can be a stronger position than a 70 in a thin one. Comparing scores across different categories does not mean much.

Being named is not being chosen

It measures whether you enter the consideration set, not whether you win the deal. It is a leading indicator, and it should be read as one.

The second number

Whether an agent could actually buy from you

Being named is half the question. The other half is what happens when an AI agent, acting for a buyer, tries to complete the purchase — and that half is measured separately, on its own scale, for reasons worth stating plainly.

A level is a gate, not a total

Level 0 to 4: invisible, listed, discoverable, buyable, operated. Each level has requirements, and you are at the highest one whose requirements you meet — all of them. A shop scoring 71 points with no reachable return policy is Level 0, not “nearly Level 2”. The points exist to show distance to the next gate and buy nothing on their own, because a number that flatters you into launching a channel that cannot take an order is worse than no number.

What counts depends on what you sell

A bakery and an API vendor do not become buyable the same way. We classify what a site sells from its own published markup and grade it against the route buyers in that category actually use — a payment link for a service, a machine-payable endpoint for an API. Demanding the wrong one would mark a perfectly buyable site as broken. The classification and the signals behind it are shown in the report, and you can correct us.

The check never pays, ever

Where a site advertises a machine-payable endpoint, we send exactly one request with no credentials of any kind, and we never follow the payment link it offers back. Not a test payment, not a sandbox one, not once. An auditor that can be talked into completing a purchase is a liability, and the rule is enforced by tests written from the attacker’s side rather than by good intentions.

What we cannot see, we do not score

Whether your catalog feed actually refreshes every fifteen minutes, or your agentic checkout is configured, lives on servers we have no view of. Those are recorded as declared by you, with a date and a name, and are never computed into a pass from anything on the page. A claim about your infrastructure that we invented would be the fastest way to make the whole report untrustworthy.

Level 3 is your rails, not ChatGPT’s

Reaching Level 3 means an agent can complete a purchase on your own site, through your own payment provider. Buying inside ChatGPT is a separate programme, currently US-only, with its own approval — we do not claim it for anyone, and you should ask anyone who does to show you the paperwork.

“Level 1” means something different if you quote every job

A manufacturer who prices every project after a conversation is not failing to publish a price — that is how their market works. Grading them against a price tag would mark a perfectly buyable business as invisible, on the first page of something they paid for. So where a site publishes no prices but offers a real way to ask, three requirements substitute for the price ones: a starting figure a buyer can plan against, a request path an agent can actually take on its buyer’s behalf, and the terms that govern it. The substitutes carry the same points, so nothing is added and nothing is free — and a quote-based site with no price signal and no way to ask still fails at Level 0. The report always names which set of rules produced your number, because a number you cannot check is not worth having.

Some catalogues have a ceiling, and we say so

Cannabis, prescription products, weapons, most regulated financial products: the platforms that complete a purchase inside an assistant exclude these categories, so a merchant in one of them cannot reach Level 3 however good their site is. That decision was never theirs. We state the real ceiling, say whose limit it is, and point the plan at being found, understood and recommended — which is most of the value and all of the work anyone can actually do. Selling a plan aimed at a gate that is closed is a refund with extra steps.

When what you tell us and what we see disagree

We read your site, you answer a few optional questions, and some things are computed from your own results. Where two of those disagree — you say one platform, the page shows another — we do not quietly pick a winner. Both are kept, the report shows which one it used and why, and the disagreement itself is written up as a finding. It is often the most useful line in the audit: it usually means something on your site is not what you think it is. Every finding also carries what it is made of, and those are not the same thing — “you published a link”, “the link answered when we called it” and “money actually moved” are three different facts, and a single score that flattened them would be hiding the only one that settles an argument.

Our benchmarks never describe a shop we did not work for

We publish what typical readiness looks like by platform, by category, by market. Every figure behind those is a real merchant’s site, measured without them asking, so a group is only published once it contains at least ten different shops — and that is ten shops, not ten measurements of one. Below that, an average says too much about the individual businesses inside it, none of which hired us. Any shop that asks to be left out is removed before the average is computed, not hidden afterwards.

Orders are counted, never modelled

Where we report agent-driven revenue, an order counts only where the payment record itself carries the evidence. Sales we cannot attribute stay in an untagged column at full size — they are never quietly redistributed into the agent column to make the channel look better than it is.

What is not published

The check set stays private, on purpose

Two reasons, and only one of them is mine.

The first is commercial: the engine is the part that is not for sale. The second is that a published check list is a list of things to game. A site could raise its score without becoming any easier for an agent to actually use, which would make the number worthless to you and to me.

What you do get, on every report: each failing check named in plain language, why it matters, and what to change. That is the part that is useful. The part that stays private is only the bookkeeping.

Check my work

Every finding names the page it applies to and every engine result is dated, so you can open the page and ask the engine the same question yourself. If a result does not reproduce, tell me. That is a bug worth more to me than a compliment.

Related

See it on a real site

The scorecards are reports on real domains, produced this way.