AI SEARCH · MEASUREMENT

AI Search Ranking for Hospitals — Why There Is No Rank

2026-09-14 · Boily

There is no ranking to occupy. Generative engines answer in sentences, the answer changes between runs, each engine reads different sources, and there are no ad slots. What replaces rank is an appearance rate: the share of questions in a frozen set where the clinic is named, measured separately per engine and compared across rounds. Boily measures four engines every two weeks for Korean hospitals and reports the split rather than a single number. We do not guarantee visibility or position.

Why ranking does not apply

  • · There is no ordered list. The engine writes a sentence naming one to five places, and the order inside that sentence is not a published ranking.
  • · The same question returns different clinics on different runs. Two runs an hour apart can disagree, so any single observation is a sample, not a position.
  • · Each engine draws on different sources — a search index, map and business data, directories, or its own index — so there is no shared leaderboard to be ranked on.
  • · There are no ad slots. Nothing can be bought into the answer, which also means no vendor can place a clinic at a given position.
  • · Answers are personalised and contextual in ways the clinic cannot see, so a screenshot from one phone does not represent what other patients get.

What to measure instead — five numbers

Appearance rate

The share of questions in a fixed set where the clinic is named at all, per engine, per round. This is the primary number because it is the only one that survives the randomness of generative answers. A single query returns a sample of one; a 40-question set run every two weeks returns a rate you can compare.

Engine split

The same clinic can sit at 40% on one engine and 0% on another in the same round, because each engine reads different sources. A combined figure hides this. Reporting one number per engine is the only honest form.

Position within the answer

When a clinic is named, it can be first, third, or a passing mention in a closing sentence. Weighting by position gives a finer signal than a binary appear/not-appear, but it is noisier — position moves between runs even when appearance does not.

Share of voice

Across all answers in a round, how often each clinic in the area is named. This tells you whether your rate is low because the whole category is thin or because competitors occupy the slots.

Cited sources

Which documents the engine attached to the answer. This is the diagnostic layer: it shows whether the engine read the clinic's own pages, a map entry, a directory, or a competitor's site.

Tracking ChatGPT hospital recommendations — five rules

Fix the question set
Write the questions patients actually type, freeze them, and never edit them mid-series. Changing a question breaks comparability with every prior round.
Run all engines the same way
Same questions, same round, same settings. If web search is on for one engine and off for another, the numbers are not comparable.
Keep the raw answers
Store the full response text and cited sources, not just a yes/no. Diagnosis happens in the text — which competitor took the slot, and what was cited about them.
Run on a fixed interval
Every two weeks is enough to see movement without drowning in run-to-run noise. Weekly mostly measures randomness.
Judge on trend, not on one round
Three to six rounds before concluding anything. A single round moving up or down is usually variance.

A one-off manual check is a different exercise — useful for a first look, not for tracking. That method is in how to check if your clinic appears in ChatGPT. What to look for in a tool is in AI search visibility tools for clinics.

What the numbers look like in practice

Boily runs its own service against itself with a frozen 100-question set. In the seventh round the four engines split as follows on the same questions in the same round. The spread, not the average, is the point.

EngineAppearance rate, round 7
Perplexity42.7%
Gemini3.1%
ChatGPT1.0%
Claude0.0%

Same subject, same questions, same round. A combined figure would have read as a single modest number and hidden the fact that one engine carries nearly all of it. Small sample; no causal claim.

Frequently asked questions

Q. Is there an AI search ranking for hospitals?
A. No. Generative AI answers are sentences, not ordered lists, and there is no published position to occupy. The same question returns different clinics on different runs, each engine reads different sources, and there are no ad slots to buy. What replaces rank is an appearance rate: the share of questions in a fixed set where the clinic is named, measured per engine and compared across rounds. Any vendor quoting a ranking position for AI answers is describing something that does not exist.
Q. How do I track ChatGPT hospital recommendations?
A. Fix a set of questions patients actually ask, run them on a fixed interval, and record for each question whether the clinic was named, where in the answer, and which sources were cited. Keep the raw answer text so you can see which clinic took the slot when yours did not. One manual check tells you almost nothing because answers vary run to run. Note that API results and what a patient sees in the app can differ, so treat the API figure as an index that moves consistently rather than an exact reproduction of a patient's screen.
Q. Can a vendor guarantee a position in AI answers?
A. No, and the mechanism explains why. The vendor cannot control the model, cannot buy placement, and cannot stop the answer from varying between runs. What can be verified is whether the same fixed question set, measured before and after work on the clinic's information, moves. That is a before-and-after comparison, not a guarantee.
Q. Why do the four engines disagree so much?
A. They read different things. One leans on a web search index and on map and business data for local questions. One is grounded in another search index and reads business profiles as documents. One favours directories and pages that state a clinic's specific capability. One uses its own index and cites many sources per answer. In our own measurement of Boily itself, the four engines ranged from roughly 43% to 0% in the same round on the same 100 questions.
Q. How long before changes show up?
A. Observed range is a few days on the fastest engine to six to twelve weeks on the ones that depend on a search index being re-crawled and re-ranked. Because of that spread, work done this month is often invisible in this month's round and appears two or three rounds later. This is the main reason to keep the question set frozen — otherwise there is nothing to compare the later rounds against.

Measured every two weeks, reported per engine

Boily is a Korea-focused AI search visibility service for hospitals and clinics. We measure ChatGPT, Claude, Gemini and Perplexity against a frozen question set and, on request, do the remediation and re-measure with the same set. We do not guarantee visibility or position.

Written by Boily. Observations come from our own biweekly measurement of Korean clinics; samples are small and no causal claim is made. Engine behaviour changes without notice. Measurements are taken via API and can differ from what an individual user sees in an app. Boily is a vendor in this market and does not guarantee visibility or position.