Cover of The Med Spa AI-Readable Website Study, showing a face split between human and machine beside AI citation and schema readouts
med-spas
Generative Engine Optimization
Original Research

We Crawled 117 Med Spa Websites. The Ones AI Never Cites Are Perfectly Crawlable.

Original research: we crawled 3,942 pages across 117 California med spa websites and scored each on 300+ machine-readability signals against citations on seven AI surfaces. Never-cited clinics score 94.1 on crawl accessibility. Cited clinics score 94.6. The real gap is provider transparency, content depth and schema.

August 2, 2026
12 min read
Chris Panteli

Total Authority crawled 3,942 pages across 117 California med spa websites and found that the clinics AI assistants never cite are, technically speaking, in perfect health. Their median crawl accessibility score is 94.1 out of 100. The clinics that do get cited score 94.6. That difference is statistically nothing.

The gap is somewhere else entirely. Never-cited clinics score 52.5 on provider transparency against 80 for cited clinics — the widest single gap in the study. Machines can reach their pages perfectly well. They just find nothing worth quoting when they get there.

This is the second study in a series on the same 117 clinics. The first measured what happens off your website: earned media placements and AI citations correlate at ρ = 0.58. This one measures the site itself.

117med spa websites
3,942pages crawled & parsed
7AI surfaces tracked
36features tested
0survive authority controls

What the study found

Five findings, each written so you can quote it without the rest of the page:

  • In Total Authority's 2026 study of 117 California med spa websites, clinics in the bottom quartile of machine readability recorded a median of 0.5 AI citations, and only 50% were cited anywhere at all.
  • Never-cited med spas and cited med spas have statistically identical crawl accessibility scores — 94.1 against 94.6 out of 100 — so crawlability does not separate them.
  • Provider transparency shows the widest gap between cited and never-cited clinics: a median component score of 80 against 52.5.
  • Sitemap URL count (ρ = 0.60) and median words per page (ρ = 0.46) are the strongest correlates of AI citations among 36 website features tested.
  • After controlling for referring domains, Domain Rating, organic traffic and earned media, no single website feature remains statistically significant — readability travels with authority.

Only 4% of med spa sites block any AI crawler. The industry is wide open to these systems. It is simply not saying much to them.

The never-cited clinics are perfectly crawlable ghosts

Twenty-four of the 103 fully-crawled clinics — 23% — have zero recorded citations across all seven AI surfaces. Comparing their component scores against the clinics cited at least once shows precisely where the failure sits, and where it doesn't.

Where the never-cited clinics actually fail

Median 0–100 component scores, 24 never-cited clinics against the 79 cited at least once. Mann–Whitney, BH-corrected.

Provider transparency+28
never cited52.5
cited80
Structured data+14
never cited54.2
cited68.5
Crawl accessibility±0
never cited94.1
cited94.6

Crawl accessibility is statistically identical between the two groups. The never-cited are not unreachable — they are unreadable.

Read the third bar again. Crawl accessibility is the metric most technical SEO advice optimizes for, and it is identical between the winners and the invisible. The median clinic in this cohort scores 92 out of 100 on crawl accessibility and 99 on renderability. Almost everyone passes. It has stopped being a differentiator.

What separates them is whether there is anything substantive on the page: named clinical staff with credentials, real treatment descriptions, complete organization schema, content that changed recently. Being fetchable is table stakes; being understandable is the contest.

How big is the readability gap?

Total Authority scored every site on a ten-component AI Readability Score, then split the 103 complete-crawl clinics into quartiles. The bottom quartile is close to invisible: half those clinics have never been cited on any surface tracked.

The readability gap

103 complete-crawl clinics split into quartiles of the AI Readability Score. Note that medians are not monotonic — Q3 outranks Q4.

Median AI citations

Across all seven AI surfaces.

0612180.5Q1n=269.5Q2n=2616Q3n=256.5Q4n=26 AI Readability quartile

Cited at least once

On any of the seven surfaces.

025507510050%Q1n=2677%Q2n=2692%Q3n=2588%Q4n=26 AI Readability quartile

Score ranges — Q1 48.1–74.2 · Q2 74.3–79.5 · Q3 79.6–83.7 · Q4 84.6–90.9.

The honest complication is in the left chart: the relationship is not monotonic at the top. Q3's median of 16 citations beats Q4's 6.5. Several of the most-cited clinics sit in the middle quartiles, carried by brand authority rather than site structure.

That shape is the actual finding. Readability behaves like an admission ticket, not a ranking algorithm — it strongly predicts whether you get cited at all, and much more weakly how much.

Which website signals actually correlate with citations?

Thirty-six features were tested simultaneously. Fifteen survived Benjamini–Hochberg correction for multiple testing. The strongest are all measures of a site being bigger, clearer or fresher to a machine.

What actually correlates with AI citations

Spearman ρ against log-citations, n=103. Every value below stayed significant after Benjamini–Hochberg correction across 36 simultaneous tests. Only the figures the study states numerically are charted.

Crawl accessibility and renderability are greyed because they barely differentiate — almost every site passes them. Being crawlable is table stakes; being understandable is the contest.

Note what fails to register. Crawl accessibility manages ρ = 0.16 and renderability actually trends slightly negative at ρ = −0.10. Neither differentiates, because almost every site already passes.

One caution on the top line: a ρ of 0.60 for sitemap URL count means clinics with more indexed URLs rank consistently higher on citations. It does not mean publishing more URLs produces citations. Hold that thought — the next section does something uncomfortable to it.

Does readable actually cause cited?

On this evidence, no — and the study is built to say so. The hard question is whether readable sites get cited because they are readable, or because successful businesses tend to have both better websites and more authority.

Every correlation was recomputed holding four authority measures constant: referring domains, Ahrefs Domain Rating, organic traffic, and the earned-media placements measured in the companion study. No single feature stays significant at adjusted p < 0.05.

What survives when you control for authority

Partial Spearman ρ after holding referring domains, Domain Rating, organic traffic and earned media constant. These are the strongest residual signals — and none reaches significance.

Sitemap URL count is shown for contrast: the loudest raw correlate at ρ=0.60 collapses to 0.17 once authority is held constant. Big sitemaps mostly mark big businesses. All residuals land at adjusted p≈0.10 — suggestive, not proof.

Sitemap URL count is the clearest casualty. The loudest raw correlate in the entire study, at ρ = 0.60, collapses to 0.17 once authority is held constant. Big sitemaps mostly mark big, established businesses.

The variables that hold their shape best are content variables — how much you publish, how deep it goes, and how recently it changed. Freshness retains ρ = 0.27, treatment content depth ρ = 0.25, days-since-update ρ = −0.26. All land at adjusted p ≈ 0.10: suggestive, not proof.

There is also a genuine surprise in that chart. Mobile LCP correlates positively with citations after controls (ρ = 0.28) — slower-loading, content-heavy pages are cited more, not less. "Make your site faster to get recommended by AI" is not supported by this data.

Does the pattern hold on every AI platform?

Yes, on all seven independently. The readability–citation association replicates across ChatGPT, Perplexity, Gemini, Copilot, Grok, Google AI Overviews and Google AI Mode, with every adjusted p below 0.05.

Every surface rewards readability — Google's most of all

The association replicates independently on all seven AI surfaces. All adjusted p < 0.05, n=103.

PlatformTop correlated featureρadj-pClinics cited
GrokAI Readability Score0.380.00366%
Google AI OverviewsAI Readability Score0.330.01167%
Google AI ModeAI Readability Score0.320.01163%
PerplexityAI Readability Score0.310.01152%
GeminiAI Readability Score0.290.01640%
ChatGPTProvider transparency0.290.01641%
CopilotAI Readability Score0.250.04130%

ChatGPT is the only surface where provider transparency outranks the composite score — consistent with assistants preferring named, credentialed sources for medical questions.

Google's surfaces cite the widest share of the cohort at 63–67%; Perplexity reaches 52% and Copilot the fewest at 30%. The strategically interesting row is ChatGPT — the only surface where provider transparency outranks the composite score, which fits assistants preferring named, credentialed sources for anything medical.

Where the industry actually stands

Adoption of machine-readability fundamentals across all 109 successfully crawled sites. The low numbers are the competitive whitespace.

The industry scoreboard: where 109 med spa sites stand

Adoption of machine-readability fundamentals across every successfully crawled site, August 2026. The low numbers are the competitive whitespace.

94%publish an XML sitemap
84%have valid JSON-LD
53%name their medical director
36%publish an llms.txt file
34%declare LocalBusiness schema
29%use any medical schema type
4%block any AI crawler
486median words per page

That median splits sharply by outcome: cited clinics run 543 median words per page against 282 for the never-cited.

The basics are solved: 94% publish a sitemap and 84% have valid JSON-LD. The differentiators are not. Only 34% declare LocalBusiness schema, 29% use any medical-specific schema type, and barely half name their medical director anywhere on the site — despite provider transparency being the widest cited-versus-invisible gap in the whole study.

The llms.txt figure is worth a raised eyebrow: 36% of med spa sites publish one, which is far higher adoption than the format's uncertain standing warrants.

The exceptions that prove the rule

Six clinics score above 85 out of 100 on readability and hold at most one citation between them.

The exceptions that prove the rule

Six clinics score above 85/100 and hold at most one citation. One scores 90.9 on a crawl-hostile site and is cited 23 times. One sits mid-table on readability and is cited 1,104 times.

DomainPatternReadabilityAI citations
reddyaesthetics.comStrong site, almost invisible89.21
savageserenityspa.comStrong site, zero citations88.20
thegspasf.comStrong site, almost invisible87.41
loungeofbeautymedicalspa.comStrong site, almost invisible86.71
noveluwellness.comStrong site, zero citations86.30
synerchimedspa.comStrong site, zero citations85.50
greyaestheticsoc.comEntity-clear but crawl-poor — cited anyway90.923
facebeautyscience.comAuthority juggernaut, mid readability75.91,104

Those six share a profile: excellent structure, thin external authority. Machines can read them perfectly, and nothing on the wider web tells an AI they matter. That is the same lesson the earned-media study reached from the opposite direction.

The reverse case is equally instructive. FACE Beauty Science holds 1,104 citations — the most in the cohort — on a middling readability score of 75.9. At the extreme, brand demand overrules site structure entirely.

What to do with this if you run a clinic

The order matters, because the data implies a sequence rather than a checklist.

  1. Stop optimizing crawlability. You have almost certainly already won it — the median site scores 92 out of 100. Time spent there is time not spent on the gaps.
  2. Name your providers. Provider transparency is the single widest gap between cited and never-cited clinics, and only 53% of sites name a medical director. Named, credentialed humans with dedicated pages.
  3. Write real treatment pages. Cited clinics run 543 median words per page against 282 for the never-cited. Thin pages give a machine nothing to quote.
  4. Complete the organization entity. Structured data is the second-widest gap (68.5 against 54.2, adjusted p = 0.009), and entity clarity is what lets a machine resolve who you are.
  5. Then build authority. It dominates everything above. Readability makes authority legible to machines; it does not substitute for it.

If you want the technical half of that as a working list, our technical GEO audit checklist covers the same ground procedurally.

Methodology

Total Authority collected site measurements in August 2026. The full method:

  1. Crawling. An async crawler took up to 40 pages per domain at a 1.5-second delay with exponential backoff, obeying robots.txt per RFC 9309 with matched-rule evidence stored, identifying itself honestly with a contact address on every request. Logins, CAPTCHAs and bot protection were never bypassed.
  2. Raw versus rendered. Every homepage and one treatment page were rendered in headless Chromium, then word, link, heading, schema and essential-fact coverage were compared against raw HTML to quantify JavaScript dependency.
  3. Deterministic extraction. Structured data via extruct (JSON-LD, Microdata, RDFa); entities, providers, treatments across a 17-category taxonomy, freshness and link-graph metrics parsed from source with evidence snippets. No LLM guessing of measurable values.
  4. Third-party ground truth. AI citations per platform from Ahrefs, validated at ρ = 0.99 against an independent export; authority controls from Ahrefs; lab performance from Google PageSpeed Insights.
  5. Outcome-blind scoring. Ten 0–100 components with published formulas and equal weights, fixed before any correlation was computed. A PCA variant derived from site features only, plus twelve alternative weightings, all land between ρ = 0.31 and 0.40 — no cherry-picked weighting drives the result.
  6. Statistics. Spearman correlations with 2,000-resample bootstrap confidence intervals and 5,000-permutation p-values, Benjamini–Hochberg correction across 36 tests, partial correlations, negative-binomial and logistic regressions with robust standard errors, leave-one-out sensitivity and random-forest importance. Seed 20260731.

Of the 117 clinics, 109 were crawled successfully — eight block automated access outright — and 103 had a complete enough crawl to enter the analysis set. Failed measurements carry an explicit status and are excluded from the affected statistic rather than silently converted to zero, so the analysis n varies between 98 and 103 by feature.

Limitations

  • Cross-sectional and observational. No intervention was tested, so these are associations only.
  • Authority controls reduce but cannot eliminate confounding. Brand demand, reviews and offline reputation are unmeasured.
  • Site measurements were taken in August 2026 and postdate part of the citation-accumulation window.
  • A minority of sites throttle or block crawlers; their technical metrics are recorded as missing, not failing.
  • n = 117 caps statistical power, especially per-platform and in the outlier cells.
  • This measured public website disclosure, never clinical quality. Nothing here says anything about the care a clinic provides.

Frequently asked questions

Does a more machine-readable website get you cited by AI?

In Total Authority's 2026 study of 117 California med spas, clinics in the top readability quartile were cited at least once 88% of the time against 50% for the bottom quartile. The design is correlational, and no single website feature remained statistically significant once authority was controlled for, so readability is best understood as an admission ticket rather than a cause.

Do AI assistants care whether my site is crawlable?

Crawlability no longer differentiates. Never-cited med spas scored a median 94.1 out of 100 on crawl accessibility against 94.6 for cited clinics — statistically identical. The median site in the cohort scores 92 on crawl accessibility and 99 on renderability, so nearly everyone already passes.

What website feature best separates cited clinics from invisible ones?

Provider transparency, by a distance. Cited clinics record a median component score of 80 against 52.5 for the never-cited — a 28-point gap covering named medical directors, credentialed staff and dedicated provider pages. Structured data is second at 68.5 against 54.2.

Does site speed affect AI citations?

Not in the direction usually claimed. Mobile LCP correlates positively with citations after authority controls (ρ = 0.28), meaning slower-loading, content-heavy pages were cited more rather than less. This study does not support "make your site faster to get recommended by AI".

Should med spas publish an llms.txt file?

36% of the med spa sites measured already publish one, but llms.txt was not among the features that separated cited from never-cited clinics in this study. Provider transparency, content depth and structured data mattered considerably more.

About the Author

Chris Panteli is the founder of Total Authority and Linkifi, host of the Market Movers Pod, and an AI visibility researcher. His work focuses on repeatable methods for understanding brand discovery, citation and recommendation in AI answers.

See what machines actually read on your site

This study scored 117 clinics on ten components. The same measurement runs on one. Grade a priority page free, or take the LLM Visibility Audit for a full readability and citation baseline across all seven AI surfaces.