Orphan Pages, 'Click Here' Links, and Click Depth: What's Actually Sourced About AI Crawlers and Your Internal Links
"AI crawlers ignore pages four clicks deep"
Search for how AI crawlers treat internal links and you'll find confident answers. One vendor post says pages more than four clicks deep are "rarely processed" and that orphan pages are "effectively invisible" to Perplexity. Another says AI crawlers have tighter budgets than Googlebot and don't go as deep. Neither one shows crawl logs or a study. The second cites "audit observation," which is a polite way of saying somebody looked at some sites and formed an opinion.
I went looking for primary data on how deep AI crawlers follow links. I didn't find any. What I can do is separate what's documented from what's repeated, and hand you a check that costs nothing and is worth doing anyway.
The three things an audit looks at
Say you run a plumbing company with a 15-page site. A few years ago somebody built you a "Water Heater Repair" page. It's a good page. It's in your sitemap. Nothing on your site links to it. That's an orphan page: it exists, but there's no road leading there.
Now say your service pages are reachable, but the links to them all say "Read more." A person scanning the page can guess what's behind each one. A crawler reading the link text can't, because the text carries no information.
And say your "Emergency Drain Service" page is reachable only by going homepage, then Services, then Residential, then Drain Work, then the page itself. That's click depth: how many clicks it takes to get from the homepage to the page. Four clicks for something you'd like a stressed customer to find fast.
Orphans, vague link text, and click depth. That's the whole internal-link check.
What the claims say and what backs them
OpenAI's crawler documentation is the closest thing to a primary source on how one AI company's bots behave. It describes OAI-SearchBot as the crawler for surfacing sites in ChatGPT search, and says sites that opt out "will not be shown in ChatGPT search answers, though can still appear as navigational links." GPTBot is for content that may be used in training. ChatGPT-User is user-initiated, so robots.txt rules may not apply to it.
Read that again looking for link following, sitemaps, or crawl depth. It's not there. The documentation says what each bot is for and how to opt out. It says nothing about how far any of them wander once they arrive.
So the claim that AI crawlers give up at depth four isn't contradicted by anything I found. It's just not supported by anything either. Repeated is not the same as sourced.
What Google actually documents
Google has written this down, for its own crawler. Three pieces are worth knowing.
On crawlable links, Google's documentation says "Generally, Google can only crawl your link if it's an <a> HTML element...with an href attribute." Links built from JavaScript events or framework attributes such as routerLink may not be parsed reliably. If your navigation is a clickable something that isn't a real link, you may have pages that look connected to a visitor and aren't connected to a crawler.
On anchor text, the same page advises against vague link text like "click here," "read more," and "website." Link text should set the expectation for the destination. Google's example is "list of cheese types." It also warns that stuffing keywords into link text violates its spam policies, so "Water Heater Repair" is good and "best cheap water heater repair plumber near me" is not.
On orphans, the same page says: "Every page you care about should have a link from at least one other page on your site."
Click depth has an older source. John Mueller said it "is more a matter of how many links you have to click through to actually get to that content rather than what the URL structure itself looks like." That was said in a Google Webmaster Central hangout and reported by Search Engine Journal in June 2018. It's guidance about Google Search, not about AI crawlers, and it dates from 2018. It does tell you one useful thing: depth is counted in clicks, not in slashes in the URL.
Why it's still worth fixing
None of that proves an AI benefit. The fixes are the same ones Google documents for its own crawler, and they also make the site easier for a human to use. A customer with a leaking water heater who can't find your repair page has the same problem as a crawler that can't.
Link every page you care about. Use link text that says where the link goes. Don't bury your money pages. If AI crawlers turn out to care about all that, you're covered. If they don't, you still fixed things Google asks for and your visitors will notice.
It's also a small fix. In our audit, internal linking health is worth 2 of the 100 points. That's not a lever that moves your grade from a C to an A. It's an afternoon of tidying.
The ten-minute check
Open a text file and list every page that earns you money. Service pages, the contact page, the location pages if you have them.
For each one, find a link to it from somewhere else on your site. If you can't, you've found an orphan. Add a link from a page that makes sense, which is usually the page a customer would be reading right before they wanted this one.
Then confirm the link is real. Right-click the page, choose View Page Source, and look for an <a> tag with an href pointing at the page. If what's there is a button or a script doing the navigating, that's the kind of link Google says it may not parse reliably.
Next, count clicks from the homepage to each page, taking the shortest path. Anything past three is worth a second look. Three isn't a magic number backed by an AI study. It's the threshold our audit uses, and a reasonable rule of thumb.
Last, scan for "click here," "read more," "learn more," and bare URLs used as link text. Rewrite them so the text names the destination.
Where our audit scores this
Internal linking is one check inside Technical Health, which is 12 points of the 100. The check passes when fewer than 10% of crawled pages are orphans, at least 80% of internal links use descriptive text, and no crawled page is more than three clicks from the homepage. Meet all three and you get 2 points. Miss one and you get 1. Miss two or more and you get 0.
On anchor text, our list of generic link text is short: "click here," "here," "read more," "learn more," "more," "this page," "link," bare URLs, and links with no text. Google's other examples, like "website," aren't on it. Every internal link on every crawled page counts, including the ones repeated in your navigation and footer.
Our audit samples your link structure. It fetches your homepage, the pages the homepage links to, and the URLs in your sitemap, up to the limit for the tier you bought. An orphan, to us, means no link from any page we crawled. A page linked only from a page we didn't crawl can be flagged, and a page that appears in neither your sitemap nor your homepage links won't be seen at all. So it won't find every orphan on your site, and your ten-minute list is a better map than we can draw.
The free Grade Check crawls one page, so it can't evaluate links between pages. It passes this check automatically with a note that it's only assessed on multi-page scans. The paid tiers crawl up to 20, 50, or 100 pages.
Start with the list
The free Grade Check at aeocheck.net gives you a 0 to 100 score and a letter grade for your homepage in about a minute, and it costs nothing. Run it for a baseline, then do the ten-minute link check by hand. You may find a page nobody links to.