Research · Benchmark

Hospital AI search visibility benchmark — what a first measurement usually looks like, and how much it moves

Published 2026-09-24 · data as of 2026-09-23 · 23 hospital sets · 74 rounds · Boily, Korea

Across 23 hospital sets, the median first-measurement visibility was 16.5%, and 6 started below 5%. Lower starting points moved more: sets starting below 5% rose by a median of +18.8 points, while those starting at 25–45% moved by +9.6. The first month was usually quiet, and 29% of all round-to-round transitions went down. Question sets differ per hospital, so this is a reference distribution, not a ranking. Hospital names are withheld and no visibility is guaranteed.

What these numbers are

  • · Sample: 23 hospital sets (specialty × area question sets of contracted hospitals; 15 with two or more rounds), 74 rounds, June 2026 to 2026-09-23.
  • · Visibility: share of a hospital's fixed question set on which it is named in the answer, across ChatGPT, Claude, Gemini and Perplexity. Questions containing the hospital's own name are excluded from the denominator.
  • · Not a ranking: each hospital has its own question set, so absolute values are not comparable between hospitals. Only the distribution and the change per starting bucket are shown.
  • · No work details: the per-hospital records note only whether remediation was in progress. What was done, where and in what order is not published.
  • · Two discontinuities: an engine enabled web grounding on 2026-08-07 and moved every hospital at once; one measurement model was changed on 2026-08-19, so values before and after are not directly comparable.

Where hospitals start

Median 16.5%, mean 20.7%. The spread is wide.

First measurementHospital sets
5% and below6
5~25%8
25~45%5
45% and above4

A first measurement near 0% is not unusual — newly opened clinics and sites whose facts live only inside images land here.

How much it moved, by starting point — first to latest round

Starting bucketSetsMedian changeMeanRange
5% and below5+18.8 pts+21.2 pts+10.6 to +31
5~25%4+11 pts+17.9 pts+6.5 to +43
25~45%3+9.6 pts+13.7 pts+3.9 to +27.6

Low starting points have the most to fill and move the most; the higher the start, the smaller the change. Sets that started at 45% or higher have too few tracked rounds to publish a delta yet and will be added as rounds accumulate. None of these figures is a promise of improvement.

The first month versus later

Round 1 → 2 (about one month)

median +3.9 pts

only 9 of 15 rose.

Round 2 → latest (12 sets with 3+ rounds)

median +9.8 pts

the lag between indexing and answer refresh shows here.

One in three rounds goes down

Of 51 round-to-round transitions, 15 (29%) were lower than the previous round. 10 of 15 tracked sets are at their own peak in the latest round, and the median gap between a set's peak and its latest value is 0 points. No hospital rose in a straight line. A single drop is not evidence; three rounds in the same direction is.

By specialty and region type — small samples are reference only

GroupSetsFirst medianLatest medianUp / tracked
Dental1724.1%29.4%9 / 11
General hospitalref.20.6%18.2%2 / 2
Dermatologyref.27.4%34.2%2 / 2
Ophthalmologyref.228.4%28.4%0 / 0
Seoul metro area1524.1%37.9%
Other regions84.7%18.2%

Non-dental specialties and non-metro regions have fewer than five sets each; read them as direction only. Lower starting points outside the Seoul metro area go with lower competitive density and larger moves.

What was done in this period — not published

Most of these hospitals had on-site remediation running alongside measurement; some were measured before remediation started. The per-hospital records keep only that distinction. What was done, where and in what order is not published, and no causal link between the work and the numbers is asserted. The purpose of this page is to show how numbers measured on a fixed question set moved over time — not the method.

FAQ

Q. What AI search visibility is 'normal' for a hospital?
A. There is no normal, but there is a distribution. Across 23 hospital sets Boily measured, the median first-measurement visibility was 16.5%; 6 started below 5% and 4 started at 45% or higher. A first measurement near 0% is common and is not a problem in itself. Because every hospital has its own question set, this is a reference distribution, not a ranking.
Q. How much does visibility change after remediation?
A. It depends on the starting point. Hospitals that started below 5% moved by a median of +18.8 percentage points from first to latest round; those starting at 25–45% moved by +9.6. Lower starts move more. For hospitals starting at 45% or higher the tracked rounds are still too few to publish a delta. Of 15 sets tracked for two or more rounds, 13 are above their first measurement and 2 are below. No causal claim is made.
Q. When does it start moving?
A. Usually not in the first month. The median change from round 1 to round 2 was +3.9 points and only 9 of 15 rose; from round 2 to the latest round the median was +9.8 points (12 sets with three or more rounds). Indexing and answer refresh lag, so we judge on rounds 3–6, not 1–2.
Q. Is it normal for a round to drop?
A. Yes. Of 51 round-to-round transitions, 15 (29%) were lower than the previous round — roughly one in three. AI answers vary run to run, and external changes (an engine turning web grounding on in early August 2026) moved every hospital at once. A single drop is not a signal; three rounds in the same direction is. 10 of 15 tracked sets are currently at their own peak.
Q. Can I use these numbers in my clinic's marketing?
A. We advise against it. Korean medical advertising law restricts comparisons with other providers and unsupported superlatives, and because question sets differ per hospital this distribution is not a valid basis for comparison. Use it internally to place your own first measurement and decide what to do first. This is not legal advice.

Where does your hospital sit in this distribution?

The first measurement answers that. Boily measures Korean hospitals biweekly on a fixed, hospital-specific question set across four AI engines and keeps the baseline. No visibility or ranking is guaranteed.

Written by Boily. Figures are aggregated automatically from Boily's per-hospital measurement records; hospital names are withheld and regions are grouped by type. Question sets differ per hospital, so cross-hospital comparison of absolute values is meaningless, and the sample is small, so no causation is claimed. Measurements are API-based and may differ from what an individual user sees in an app. Boily is itself a vendor in this market and guarantees no visibility or ranking.