Your Site Looks Fine to You. Most AI Crawlers Get the View-Source Version.

The page you see is not the page they get

You open your own website. The phone number is there, top right. The hours are in the footer. The FAQ has twelve questions with real answers under each one. Everything looks complete, so you assume everything is complete.

Your browser is a generous reader. When it loads a page, it downloads the HTML the server sends, then runs any JavaScript that came along, and that JavaScript is allowed to rewrite the page. A booking widget drops in your phone number. A reviews plugin builds the star rating. An accordion script fetches the FAQ answers from somewhere else and pastes them in. What you see on screen is the finished result of all of that.

A lot of AI crawlers are not that generous. They ask for the page, take the HTML that comes back, and stop. The script never runs. If a fact only exists after the script runs, then as far as that crawler is concerned, the fact is not on your site.

This is the sort of problem that never announces itself. Nothing is broken. Your site is not blocked, not slow, not penalized. It just has a few holes in it that only show up from the outside, and the holes tend to be exactly the facts a customer would ask an assistant about: who you are, where you are, when you're open, what you charge.

What the server logs show

The best-known data here is a late-2024 study from Vercel, working with a research firm called MERJ. They looked at a month of traffic across Vercel's network and counted what the crawlers actually did. Googlebot made about 4.5 billion fetches. GPTBot made 569 million, Claude's crawler 370 million, AppleBot 314 million, and PerplexityBot 24.4 million.

The finding that matters for you is a flat sentence from the study: none of the major AI crawlers currently render JavaScript. That covered OpenAI's crawlers (GPTBot, OAI-SearchBot, and ChatGPT-User), Anthropic's ClaudeBot, and the ones from Meta, ByteDance, and Perplexity.

There was an interesting wrinkle. These crawlers do request JavaScript files. GPTBot fetched them in 11.5% of its requests, and Claude's crawler in 23.84%. So they download the script. They just don't run it. It's the difference between picking up a recipe card and cooking the meal.

Two exceptions came out of the same study. Gemini leans on Googlebot's infrastructure, which does full JavaScript rendering, and AppleBot renders pages through a browser-based crawler. Everybody else in the list read raw HTML and moved on. Vercel's recommendation was to prioritize server-side rendering for critical content, which is a developer's way of saying "make the server send the real page, not a skeleton the browser fills in."

This is a single vendor's sample, and it's nearly two years old. So I went looking for something newer.

A test from this June

A published test from June 2026 went at the question from the other side. Instead of reading crawler logs, the author asked twelve AI assistants to read a page and report back what it said.

The setup was clever. Each assistant got its own unique secret URL. The raw HTML of the page held a decoy value. An external script then swapped the decoy for the real value. So if an assistant reported the decoy, it had read only the raw HTML. If it reported the real value, it had run the script. Server logs and a canary check confirmed the fetches actually happened.

Seven assistants reported the decoy: ChatGPT, Claude, Gemini, Perplexity, Meta AI, Copilot, and Grok. Five reported the real value, meaning they rendered the JavaScript: DeepSeek, ERNIE, Qwen, Kimi, and Mistral.

If your customers are asking ChatGPT, Claude, Perplexity, or Copilot, the first group is the one that matters, and every one of them read the unrendered page.

Gemini did not render JavaScript in the live chat fetch, even though Google owns some of the best rendering infrastructure on the planet. That looks like it contradicts the Vercel finding, but it doesn't. Those are two different jobs. Vercel's point was about the crawler that builds an index over time. The June test measured an assistant fetching one page, live, in the middle of answering a question. The same company can run both, and they can behave differently.

Grok is messy. One of its proxy nodes did run the JavaScript, and the answer still quoted the decoy. Whatever is happening between the fetch and the answer, "the fetcher can render" and "the answer includes the rendered content" are not the same statement.

The test was one pass, one prompt per assistant, one moment in time, one production domain. Perplexity even said it couldn't access the page while the logs showed it had retrieved it. Treat it as a snapshot, not a law. But it's a snapshot from this year, and it points the same direction as the logs from two years ago.

Where the Google exception comes in

Google's own documentation says Googlebot renders JavaScript, with some caveats about when and how, and the JavaScript SEO basics page on Google Search Central lays it out. Bingbot renders too. Bing moved to an evergreen, Chromium-based crawler back in 2019.

That's why your site may look perfectly fine in classic Google results while being thin everywhere else. The search engines that have been around for twenty years built rendering into their crawlers because they had to. The newer assistants mostly didn't.

It's commonly described that Google's AI Overviews and AI Mode benefit from the same rendering, since they sit on Google's index. I found no primary source that says so.

Do not stretch the Bing and Google point too far. Copilot did not render JavaScript in the June test, and Bing's index feeds several AI products. A crawler that renders and builds an index is one thing. The assistant that fetches your page live while answering a question is another. "Bing renders" does not mean "Copilot sees your JavaScript content." Results also move around. A separate test published in December 2025 found Gemini rendering JavaScript on a live fetch. By the June 2026 test it didn't. Neither study is wrong. Different tests, different months, different setups. The behavior changes over time and between tests, so do not build your plan on any single assistant's habits.

Schema won't rescue the missing text

The tempting fix is to say, fine, I'll put the phone number and hours in JSON-LD structured data, and the machines can read that instead. It's a reasonable instinct. It doesn't hold up well.

That December 2025 test, from searchVIU, built a fictional product page with eight price variants, each shown in a different format. One of those variants existed only in JSON-LD. On direct fetch, no system extracted the price that lived only in the schema. Gemini picked up the right price in 4 of 8 variants and ChatGPT in 3 of 8, and the schema-only one was not among them. Perplexity and Google AI Mode scored lower still, 12.5% and 25%, though those were measured after indexing rather than on a direct fetch, so they aren't a like-for-like comparison. Hidden microdata and RDFa were recognized by none.

It doesn't say schema is worthless. Structured data may matter during indexing and training, and in Google's AI Overviews, and that wasn't measured. It does say that on a live fetch, a fact that exists only in a schema block is a fact an assistant may never see.

This connects to something I wrote about FAQ schema recently. The advice that holds up is the same here. Put the fact in the visible text, in the HTML the server sends, and put it in the schema as well. Schema is the second copy, not the first.

Which sites are actually at risk

Fewer than the headline suggests.

A plain site where the server sends finished HTML is fine. Most hand-built brochure sites, most WordPress sites running a normal theme, and pages from at least some of the big site builders fall in that bucket. Wix, for instance, says in its help center that it uses server-side rendering, and it states that most AI crawlers don't run JavaScript. That's a platform that knows the problem and designed around it.

The risk sits in a few specific places.

The first is the whole-site shell: a site built as a single-page app, where the server sends a nearly empty page and a JavaScript framework builds everything in the browser. Those are the sites Vercel's recommendation was written for.

The second is the one I see as more common for a small business, and much quieter: a normal, server-rendered page with a few important pieces injected by a widget. The phone number from a call-tracking script. The hours from a Google Business Profile embed. The FAQ answers loaded by an accordion plugin that fetches them after the page loads. The reviews from a third-party badge. The rest of the page is fine. What's missing is exactly what a customer would ask about.

I have no verified data on how Squarespace or GoDaddy handle this. Test your own site.

The two-minute check

You need a browser and the address of one page, preferably your homepage and a service page.

Do not use Inspect or Inspect Element. That shows you the page after the JavaScript has run, which is the generous version your browser already showed you. You want the other one. Put view-source: in front of the address, so view-source:https://yoursite.com, or right-click the page and choose View Page Source. A wall of raw code opens. That is, roughly, what a non-rendering crawler gets.

Now use your browser's find command, Ctrl+F or Cmd+F, and search that source for three things.

Your phone number, typed exactly as it appears on the visible page. Then the name of one service you sell, like "water heater" or "teeth whitening." Then the first few words of one FAQ answer, copied from the visible page.

If all three turn up, you're in good shape for this problem. The facts are in the HTML the server sends.

If the phone number is missing but the service name is there, you have a partial-injection problem. A script is adding the number. Ask whoever built the site to put it in the page itself, or add it as plain text somewhere server-rendered, like the footer.

If the FAQ answers are missing but the questions are there, same story, and it's the one worth fixing first, because those answers are exactly the text an assistant wants to quote. Some FAQ plugins load the answers only when a visitor clicks. Whether the answer text sits in the page already, just hidden, or gets fetched on the click, is something the source will tell you.

If none of the three turn up, and the source is mostly a few lines and a lot of script tags, you're looking at a whole-site shell. That's the larger conversation, and a call to your developer.

If you're comfortable at a command line, the same test works with curl. Something like curl -sL https://yoursite.com | grep -i "your phone number" fetches the raw page and searches it. No result means the number isn't in the HTML. It's the same check, just without the browser in the way.

Changing your browser's user agent is not a substitute. Your browser will still run the JavaScript, so you'll still see the finished page. View source is the test.

What all of this does not prove

None of this proves that fixing it will get you cited. Raw-HTML visibility is a precondition. If an assistant can't see your hours, it can't quote your hours. If it can see them, that gives it a reason it could, not an obligation it will.

The evidence is also uneven. The schema test used a fictional product page. None of them used a plumber's website. They agree with each other on the direction, which is why I trust the direction, and none of them gives you a number to expect from your own site.

And five of the twelve assistants did render the script. If your customers use those, you may be fine.

How the audit sees this

The AEO Checker audit fetches your pages with a plain HTTP request, no JavaScript execution, and works from the raw HTML that comes back. So it sees roughly what a non-rendering AI crawler sees. It does not impersonate any particular bot; it sends its own user agent.

There is a narrow JavaScript check built in. On the homepage, if the page looks like an empty shell (very little text, at most one heading, and either a nearly empty React, Vue, Angular, Next, or Gatsby container or a mostly-script page), the paid report adds a "Heads up" banner and marks the affected Structured Data and Content Structure checks as "Not Determinable" with partial credit, instead of failing them. The free Grade Check result page does not show that banner.

That check will not catch the case this whole post is about. It does not compare raw HTML to rendered HTML. It does not look at inner pages. A server-rendered page where only the phone number or the FAQ answers are injected by a widget will not trip it. If your whole site is a JavaScript shell, the report will flag it. For partial injection, the view-source check above is the tool, and you are the one holding it.

Do the check first

Run the three searches. It takes less time than reading this post did. If everything turns up, you've crossed one thing off the list for free.

Then, if you want a number to go with it, the free Grade Check at aeocheck.net runs your homepage through the same scoring engine as the paid report and hands you back a score out of 100 and a letter grade. It can take up to a minute and costs nothing. Use it as the baseline, and use view-source to find out what the score can't tell you.