How the register works
Methodology · scoring · the monthly cycle
What this measures
When people decide where to spend money, they increasingly ask AI assistants — ChatGPT, Claude, Perplexity, Google AI Mode — instead of scrolling search results. The register tests whether an entity (a hotel, a brand, a product line) actually appears in those AI answers for the questions real travelers ask. An entity that never appears is, commercially speaking, a ghost: real, open for business, and invisible at the moment of decision.
The scoring vocabulary
Every prompt is tested on each platform in a fresh conversation and stamped with exactly one of three states. SIGHTED — the entity appears in the response: named, listed among options, or recommended. VANISHED — the entity does not appear anywhere in the response. UNTESTED — not yet run this cycle. Visibility is the share of tested platform-prompt pairs stamped SIGHTED. No partial credit, no judgment calls beyond "does it appear."
Commercial intent weighting
Each prompt carries a funnel stage — Dreaming, Planning, Booking — weighted 1×, 2×, and 3× respectively in the intent-weighted score. A miss on a booking-intent prompt (a traveler ready to pay) costs three times a miss on a dreaming-intent one. The register's most important number is not average visibility — it is the count of booking-intent prompts where the entity appears on no platform at all.
Testing protocol
Prompts are unbranded — they never contain the entity's name, because the test is discovery, not name recognition. Each is written exactly as a real traveler would type it and pasted verbatim, with no rewording, into a fresh conversation on each platform (chat history biases AI answers toward previously discussed brands). Testing runs on a monthly cycle; after content ships against a logged gap, the same prompt is re-run the following cycle to confirm whether visibility moved. Some platforms can answer for themselves, through their own APIs with web search enabled. Those results are marked AUTO until a person confirms them — a machine grade is a claim awaiting review, not a stamp. In every case an API answer is a close proxy for what the consumer app returns, not an identical one, and the caveat differs per engine: Claude — the Anthropic API with server-run web search. ChatGPT — the OpenAI Responses API with its built-in web search tool. The consumer app layers its own retrieval and personalisation on top, so treat differences as directional. Platforms without a runner stay manual, and a platform whose key is not configured on the server shows its control disabled rather than hidden. One grader, always the same one. Whichever engine produces an answer, the pass that decides whether the entity appeared is the same Anthropic call with the same instructions. Scores are only comparable across engines if the thing doing the comparing does not change; an engine grading its own answer would be marking its own homework. Cited sources are read from each response's own citation metadata rather than inferred from its prose. Auto-tests run before 27 August 2026 under-reported citations — the earlier version asked the grading model to spot domains in text that never contained any, so it almost always found none. Re-running a prompt repopulates them.
Fan-out — one prompt is never one prompt
An answer engine rarely retrieves for the question it was asked. It expands that question into several narrower sub-queries, retrieves for each, and writes one answer from everything it found. A traveler asking about planning sees a single paragraph; underneath it may sit half a dozen searches about amenities, budget, neighbourhood, and season. That is where visibility is quietly lost. An entity can be perfectly findable on the prompt itself and still miss the answer, because it was absent from the sub-queries that actually fed it — and testing the seed alone would never show it. So a prompt in the register can carry a family: the seed, plus the facets an engine would plausibly derive from it, spread across amenity, audience, price, occasion, and geography. Facets are stamped exactly like any other prompt, on the same platforms as their seed. A family counts as covered only when every tested member was sighted somewhere — one invisible facet is enough to leave the family incomplete, which is the point. Families are one level deep; facets of facets describe nothing a reader can act on.
The monthly cycle checks itself off against a real ledger once you're signed in.
Open the register
What the register does not claim
AI answers are non-deterministic — the same prompt can vary between runs, so a single stamp is a sample, not a verdict; trends across cycles matter more than any one test. Visibility measures presence in AI answers, not bookings or revenue. Month-over-month movement is directional evidence, not proof of causation. And where testing uses AI platform APIs rather than the consumer apps, results are a close proxy, not a verbatim match — consumer apps should be spot-checked manually on a slower cadence. Engines' internal fan-out is not observable: a family is a modeled approximation of the sub-queries an engine would run, not a record of the ones it did.
Haunting grounds — choosing who you appear for
An AI answer is a shortlist of three to five names, which makes the generic ask unwinnable and the niche ask winnable. So before a property is tested against everyone, the register reads the property's own website, names the one sentence a local competitor could not copy, and ranks three candidate audiences by whether the business case is any good — not by how much traffic each would bring. What the scrape does. One live read of the site, returning six to eight differentiator signals rated STRONG (well evidenced), WEAK (mentioned but thin) or ABSENT (an obvious segment gap the site never mentions). Absent signals are the ones that become work items, because they are the only kind somebody can act on. Match and rivals are AI estimates, not measurements. Nothing counted the rival properties in a market, and nothing measured how much of an audience's language a site can back. Both numbers are a model's judgement, and they are directional — useful for choosing between three options, not for reporting as fact. The recommendation itself is derived from them by a stated rule (best evidence, then smallest denominator), so it is reproducible even though its inputs are estimates. “Who wins now” and the guest quotes are AI-searched. They are what a model found or inferred from public sources at the time of asking. Verify them before quoting either in published content — a review attributed to a platform that never carried it is a worse problem than having no quote. One haunt per cycle. Confirming an audience makes it that cycle's haunt and returns any previously confirmed one to parked — parked, not deleted: its cached detail and armed targets survive, so a decision can be revisited. Armed targets are inserted into the following cycle's prompt set, tagged with their origin, and their priority sets the order the battleground runs them in. A ghost that appears everywhere gets believed nowhere.
Previously tracked
Microsoft Copilot was tested as a sixth platform until August 2026 and is no longer part of the register. Its historical stamps are kept — they are real observations about what Copilot answered at the time, and deleting them would rewrite the record — but it is not tested going forward and does not count toward any current score. A platform leaving the grid stops being counted rather than counting as a miss.
METHODOLOGY v1.0 · AUGUST 2026 · CHANGES TO SCORING OR PROTOCOL ARE VERSIONED ON THIS PAGE
How it works · The Ghost Register