Google's crawl budget guide is for sites with a million pages, or 10,000 changing daily. What a small site should fix instead, and how to read Crawl Stats.
LLaunchScaler·Published ·7 min read
Crawl budget is the set of URLs Google can and wants to crawl on your site, and a small site almost never needs to manage it. Google's own crawl budget guide says it is written for sites with about a million unique pages that change weekly, or 10,000 or more that change daily, and that if your pages are crawled the day they are published, you don't need it.
What a small site should care about instead is narrower: not wasting the crawls it gets on junk URLs, keeping the server fast and error-free, and giving Google reasons to crawl the pages that matter.
What is crawl budget?
Google defines a site's crawl budget as "the set of URLs that Google can and wants to crawl." It is the product of two things: the crawl capacity limit, which is how much crawling your server can handle without slowing down, and crawl demand, which is how much Google wants to crawl your URLs. Budget is set per hostname, so www.example.com and docs.example.com each have their own.
Both halves move:
Part
What sets it
What raises it
What lowers it
Crawl capacity limit
How much Googlebot can crawl without overloading your server
Consistent, fast responses
Slower responses, 5xx errors, 429 rate limiting
Crawl demand
How much Google wants your URLs
Popular URLs, content that changes, a site move
Questions, answered
What people ask about this
01
What is crawl budget?
Google defines crawl budget as the set of URLs on a site that Google can and wants to crawl. It combines the crawl capacity limit, how much crawling your server can take, with crawl demand, how much Google wants to crawl your URLs.
Many duplicate or unimportant URLs Google already knows about
Google says every site starts with the same conservative default capacity limit, which it adjusts up over time if there is demand and the site stays healthy. It also notes that even if capacity isn't reached, low demand means Google crawls less, which describes a new site that few other pages link to.
When does crawl budget matter, according to Google?
Google's guide lists three kinds of site it is written for: large sites with about 1 million or more unique pages whose content changes moderately often (once a week), medium or larger sites with 10,000 or more unique pages that change very rapidly (daily), and sites where a large share of URLs sit in "Discovered - currently not indexed". It calls these rough estimates, not exact thresholds.
The guide opens with a test that rules out many sites: "If your site doesn't have a large number of pages that change rapidly, or if your pages seem to be crawled the same day that they are published, you don't need to read this guide." For those sites, it says keeping the sitemap up to date and checking the Page indexing report regularly "is adequate."
Search Console's own help draws a similar line. It says the Crawl Stats report is aimed at advanced users and that "if you have a site with fewer than a thousand pages, you should not need to use this report or worry about this level of crawling detail."
You can run Google's test yourself. Publish a page, link it from your homepage, and inspect it in URL Inspection a day or two later. If the Last crawl date is filled in, crawling is keeping up with your publishing, and crawl budget is not your problem.
What does "Discovered - currently not indexed" mean on a small site?
It means Google knows the URL but hasn't crawled it yet. Google's help says that typically Google wanted to crawl it "but this was expected to overload the site," so it rescheduled. On a small site that rarely means the site is too big; it points at a slow or erroring server, or at URLs Google has little reason to prioritise.
Work through it in this order:
Check the server. Open Settings > Crawl stats and look at host status and average response time. Slow responses and server errors lower the crawl capacity limit, which is the "overload" Google's message refers to.
Check the demand. Are the discovered pages linked from pages Google already crawls, or only listed in the sitemap? Google's link guidance says every page you care about should have a link from at least one other page on your site.
Check the inventory. If the discovered URLs are parameter variants, tag archives or near duplicates, Google is right to deprioritise them. Fix the source of those URLs rather than the crawl.
A small site can still generate thousands of URLs that are worth nothing: filter and sort parameters, calendars and other infinite spaces, redirect chains, soft 404s and duplicate pages. Google's troubleshooting guide lists these as the URLs that "can negatively affect a site's crawling and indexing." Cleaning them up costs nothing and helps even when budget isn't tight.
Crawl waste
How it arises on a small site
Fix
Parameter traps
?sort=, ?filter=, ?view= or session IDs on list pages, each combination a new URL
Disallow the parameters in robots.txt, and link to one unfiltered listing.
Infinite spaces
Links that generate endless new URLs, such as a calendar's next-month link or internal search results
Block them in robots.txt; Google's guide says to block infinite spaces from crawling.
Redirect chains
http to https to www to trailing slash, one hop each
One 301 from any variant to the final URL. Google's guide says long chains hurt crawling.
Soft 404s
Empty or "not found" pages that return 200
Return a real 404 or 410. Google says soft 404 pages keep being crawled.
Duplicate content
The same page on several URLs
Consolidate with redirects and canonicals.
Server errors
5xx responses or timeouts
Fix the cause; Google lowers the crawl rate in proportion to the URLs returning errors.
Two of Google's recommendations surprise people. First, don't use noindex to save crawling: the guide says Google "will still request, but then drop the page," wasting the crawl. Use robots.txt for URLs you never want fetched. Second, don't add and remove robots.txt rules to shuffle budget between sections; Google says it won't shift the freed crawls elsewhere unless you're already at the capacity limit.
Open Settings, then Crawl stats, in a Domain property or a root-level URL-prefix property. The report charts total crawl requests, total download size and average response time, shows a host status, and breaks requests down by response code, file type, purpose and Googlebot type. For a small site, host status and response codes are the parts worth reading.
What each part tells you:
Total crawl requests counts every request Google made, including page resources on your site and each hop of a redirect chain separately.
Average response time is the mean for every resource fetched. A rising line is worth a look, because slower responses lower the capacity limit.
Host status flags significant availability problems in the last 90 days in three categories: robots.txt fetching, DNS resolution and server connectivity. The report counts a category as having an issue when, for example, DNS resolution fails for more than 5% of requests on a day.
Crawl responses groups requests by status. Google's help says most responses should be 200 unless you are moving or reorganising the site.
The robots.txt line matters more than it looks. The report explains that if robots.txt returns a 429 or 5xx, Google stops crawling the site for the first 12 hours, then uses the last good copy for up to 30 days. A 404 for robots.txt is fine: it means there is no file and everything may be crawled.
What should a small site do instead of optimising crawl budget?
Make every URL worth crawling and every important page easy to reach. For a site with a few hundred pages, that means a clean sitemap, internal links to every page that matters, a fast server that returns the right status codes, and no parameter or redirect clutter. Google's guide lists the same practices for large sites; a small site just needs fewer of them.
Keep the sitemap current, listing only canonical URLs that return 200, with <lastmod> on pages that changed.
Link every page you want indexed from at least one page Google already crawls.
Return 404 or 410 for removed pages, and 301 for moved ones, in one hop.
Block parameter and infinite URL spaces in robots.txt, and never CSS or JavaScript your pages need.
Keep response times steady and support 304 Not Modified, which Google says lets it reuse its cached copy.
Check Crawl stats host status after hosting changes, and the Page indexing report monthly.
Check the crawl waste on your site
Most crawl waste on a small site comes from a handful of settings. Run the free scan on LaunchScaler with your address and no account; it runs 156 checks across 6 of its 7 categories at no cost, including redirect chains that take more than one hop, soft 404 pages that return 200, broken internal links and dead sitemap URLs, a sitemap that lists redirecting or non-canonical URLs, and robots.txt rules that block pages you want ranked.
Rarely. Google's crawl budget guide says it is for large sites (about 1 million or more unique pages changing weekly) or medium sites (10,000 or more pages changing daily), and that if your pages are crawled the day they are published you don't need it.
03
How do I check my crawl budget?
Open Settings > Crawl stats in Search Console. It shows total crawl requests, download size, average response time and host status over time. Google's help says a site with fewer than a thousand pages should not need this level of detail.
04
What wastes crawl budget?
URLs that don't deserve crawling: faceted and sorted parameter URLs, infinite spaces such as calendars, long redirect chains, soft 404 pages, and duplicate content. Slow responses and server errors also lower how much Google crawls.
05
Does noindex save crawl budget?
No. Google's crawl budget guide says noindex pages are still requested and then dropped, which wastes crawling time. To stop crawling of URLs you never want fetched, use robots.txt.
Crawled - currently not indexed means Google fetched the page and chose not to index it. Triage by URL pattern, then merge, noindex or improve each page.
Discovered - currently not indexed means Google knows the URL but postponed the crawl. Check server capacity in Crawl Stats, then raise demand with links.