AI crawlers, 404s and redirects: how broken URLs cost citations
Vercel measured ChatGPT's crawler spending 34.82% of fetches on 404s and 14.36% on redirects, against 8.22% and 1.49% for Googlebot. How to cut the waste.
LLaunchScaler·Published ·7 min read
AI crawlers waste a large share of their requests on dead links and redirects: Vercel measured ChatGPT's crawler spending 34.82% of its fetches on 404 pages and Claude's 34.16%, against 8.22% for Googlebot, and ChatGPT spent an extra 14.36% of fetches following redirects, against 1.49% for Googlebot. Every wasted fetch is one that did not reach a page you want quoted.
You cannot change how these crawlers pick URLs, but you can stop feeding them bad ones. Link internally to final URLs that return 200, keep your sitemap free of redirecting and dead URLs, and return honest status codes. Below are the numbers and the fixes, checked in September 2026.
How often do AI crawlers hit 404s and redirects?
Much more often than Googlebot. Vercel and MERJ measured crawler traffic across Vercel's network and on nextjs.org and published the results in December 2024. Roughly a third of ChatGPT's and Claude's fetches ended in a 404, about four times Googlebot's rate, and ChatGPT spent almost ten times Googlebot's share on redirects.
Crawler
Fetches in one month on Vercel's network
Share of fetches on 404 pages
Share of fetches following redirects
ChatGPT (GPTBot)
569 million
34.82%
14.36%
Claude
370 million
34.16%
Not reported
Googlebot
Questions, answered
What people ask about this
01
Do AI crawlers waste requests on 404 pages?
Yes, far more than Googlebot. Vercel measured ChatGPT's crawler spending 34.82% of its fetches on 404 pages and Claude's 34.16%, against 8.22% for Googlebot, across its network in one month.
Vercel's own reading: "These high rates of 404s and redirects contrast sharply with Googlebot," which "suggests Google has spent more time optimizing its crawler to target real resources." The data is from late 2024 and crawler behaviour changes, so treat the figures as a measured pattern rather than today's exact rates.
Why does link hygiene matter more for AI crawlers than for Googlebot?
Because AI crawlers make fewer requests, pick URLs less carefully and read only the raw HTML. Googlebot made about 4.5 billion fetches in Vercel's month against 569 million for GPTBot, so every wasted request is a bigger share of what an AI crawler spends on your site. And what it fetches, it reads raw.
Three differences stack up:
Selection. Vercel found AI crawlers "show less predictable patterns in their URL selection" and that high 404 rates suggest they "may need to improve their URL selection and validation processes."
Rendering. Vercel found that none of the major AI crawlers it measured render JavaScript, apart from Google's Gemini and Apple's crawler. A non-rendering crawler judges a page by its status code and raw HTML alone.
Documentation. Google publishes exactly how it treats each status code, including that its crawlers "follow up to 10 redirect hops." OpenAI, Anthropic and Perplexity document their user agents and robots.txt handling but not their redirect or error handling, so you cannot count on the same tolerance.
Do 404s and redirects actually cost you citations?
Indirectly, through crawl waste. No study has tied a 404 rate to a citation count, and no AI operator publishes how many pages it will fetch from your site. But a page a crawler never reaches cannot be quoted, and Bing, whose crawl guidance addresses AI grounding directly, spells out the cost of waste.
Bing's webmaster guidelines say it "allocates crawl capacity based on site health, efficiency, signal quality, and crawl value," and that excessive crawl waste can "Delay indexing of important content" and "Reduce grounding visibility for priority URLs." Grounding is the step where Copilot picks the pages its answer cites. Its guidelines also ask sites to remove deleted or redirected URLs from sitemaps promptly and to update changed URLs through IndexNow.
Vercel's recommendation for site owners who want to be crawled makes the same point: "Efficient URL management matters more than ever," because "The high 404 rates from AI crawlers highlight the importance of maintaining proper redirects, keeping sitemaps up to date, and using consistent URL patterns across your site." The fixes below are cheap, and they help Googlebot and Bingbot at the same time.
Where do AI crawlers' wasted fetches come from?
From URLs that no longer resolve to a real page: old asset files, pages you moved or deleted, internal links to retired URLs, and sitemap entries nobody updated. Vercel found that, excluding robots.txt, AI crawlers "frequently attempt to fetch outdated assets from the /static/ folder," which means part of the waste comes from URLs they learned long ago.
The part you control is every URL you still publish:
Internal links in navigation, footers and old blog posts that point to moved pages.
Sitemap entries for URLs that now redirect, return 404 or carry noindex.
Canonical tags that point at a URL which itself redirects.
Links from your own profiles and listings on other sites to an old domain or path.
Variants of the same URL: http against https, www against bare domain, trailing slash against none, each adding a redirect hop.
How do you fix 404 waste?
Find every internal link and sitemap entry that returns a 4xx, then fix the link, redirect the old URL, or remove it. Return a real 404 or 410 for content that is gone for good. Bing's guidelines state the same rule for its crawler and Copilot: use 301 redirects for permanent moves, and "Return a 404 status code" when content is permanently removed.
Crawl your site with any link checker and export every internal link that returns 4xx.
For each broken link, change it to the page's current URL. Fixing the link is better than adding a redirect.
For pages that moved, add one 301 from the old URL to the new one, so links elsewhere on the web still resolve.
For pages that are gone with no replacement, return 404 or 410 and remove every internal link to them.
Remove dead URLs from your sitemap.
Re-run the crawl after each deploy that changes URLs.
Link directly to the final URL, and collapse any chain to a single 301. Every internal link, canonical tag and sitemap entry should point at a URL that returns 200 on the first request. Redirects are for outside links and old bookmarks you cannot edit, not for your own navigation.
Each HTTP/ line is one response. One HTTP/... 200 line means no redirect; 301 followed by 200 is one hop; anything longer is a chain to collapse. Common chains come from stacked rules: http to https, then bare domain to www, then adding a trailing slash. Combine them so the first response already points at the final form. Loops and longer chains are covered in redirect chains and loops.
Why are soft 404s worse for non-rendering crawlers?
Because a soft 404 answers 200, and the status code is the main signal a non-rendering crawler has. Google detects when a 200 page "suggests an error for Google Search, an empty page or an error message" and reports it as a soft 404. No AI crawler operator documents a similar check, so an error page served with 200 can be taken as a real page.
Single-page apps are the usual source. The server answers every path with the same HTML shell and a 200, and JavaScript decides later that the route does not exist. A crawler that does not run JavaScript never sees that decision; it sees a 200 and an almost empty page.
Fix it at the server:
Return a 404 status from the server for unknown routes, not only a "not found" component in the browser.
Return 410 for content you deleted on purpose.
Make sure your framework's not-found handler sets the status code, then confirm with curl -sI https://example.com/this-page-does-not-exist.
How do you check your sitemap for dead and redirecting URLs?
Request every URL in it and list the ones that do not return 200. Google says it uses a sitemap's <lastmod> only when it is "consistently and verifiably" accurate, and Bing asks for sitemaps that "List only canonical URLs" and "Remove deleted or redirected URLs promptly." A sitemap full of redirects tells every crawler to waste fetches.
This loop prints each sitemap URL whose first response is not 200:
Remove or replace every line it prints. For a sitemap index, run it on each child sitemap. Then keep it clean by generating the sitemap from the same source as your routes, so a deleted page leaves the sitemap in the same deploy.
Check your links, redirects and status codes in one scan
Run the free scan on LaunchScaler to see what crawlers hit on your site. It takes a URL, needs no account and runs 156 checks across 6 of its 7 categories at no cost. It flags broken internal links and dead sitemap URLs that waste AI crawl budget, internal links that depend on redirects, redirect chains and loops, soft 404s served with a 200, and sitemaps that list redirecting, 404 or noindexed URLs.
They do, but at a cost. In the same Vercel data ChatGPT spent an additional 14.36% of its fetches following redirects, against 1.49% for Googlebot. Every hop is a fetch that could have gone to a real page.
03
How do I reduce 404s from AI crawlers?
Point internal links at final URLs that return 200, redirect moved pages with a single 301, return a real 404 or 410 for pages that are gone, and keep your sitemap free of redirecting and dead URLs.
04
What is a soft 404 and why does it matter for AI crawlers?
A soft 404 is a 'not found' page served with a 200 status code. Google detects many of them and reports them in Search Console, but AI crawler operators do not document doing the same, so a crawler that trusts the 200 can treat an error page as real content.
05
Does Googlebot handle redirects better than AI crawlers?
Google documents that its crawlers follow up to 10 redirect hops and use the final target's content. AI crawler operators do not publish equivalent limits, which is one reason to keep every redirect to a single hop.
When ChatGPT names competitors and not you, list who wins each prompt, open the sources it cited, and get onto those pages. A step-by-step gap analysis.
Bing's AI Performance report counts how often Copilot and Bing's AI summaries cite your pages and shows the grounding queries behind them. How to use it.