Crawl budget is the number of pages a search engine or AI bot is willing to crawl on your site in a given period — and if you waste it on junk URLs, your important pages get crawled less often and updated more slowly in the index. For small sites this rarely matters. But once you pass tens of thousands of URLs, or you generate endless combinations through filters and parameters, crawl budget becomes a real constraint on how fresh and complete your presence is in search and AI answers.
What crawl budget is and when it matters
Every crawler operates under two limits. The first is crawl rate — how fast it can request pages without straining your server. The second is crawl demand — how much it actually wants to crawl you, based on your site's popularity and how often your content changes. Together these set a practical ceiling on pages fetched per visit.
If your site is a few hundred pages, bots will comfortably cover everything and budget is a non-issue. It starts to bite when you have a large catalog, a big publisher archive, or a site that programmatically generates URLs. The danger is not that bots stop crawling; it is that they spend their limited attention on low-value pages and revisit your money pages less often. To understand the mechanics behind this, our guide on how search engines crawl and index your website lays out the full crawl-to-index pipeline.
Faceted navigation: the biggest budget drain
Faceted navigation — the filters on category pages for color, size, price, brand, and so on — is the classic crawl-budget killer. Each combination can generate a unique URL, and those multiply combinatorially. A category with a handful of filters can spawn thousands of near-identical, thin URLs that add nothing to the index but soak up crawl budget and dilute your signals.
The goal is to let crawlers reach the valuable pages and steer them away from the infinite combinations. Common tactics:
- Block worthless parameter URLs in robots.txt so bots do not crawl filter permutations you never want indexed.
- Use canonical tags to point filtered variants back to the clean category page where appropriate.
- Apply noindex to thin combinations you must keep accessible to users but do not want in the index.
- Avoid linking to low-value parameter URLs in your internal navigation so bots do not discover them in the first place.
Getting your robots.txt rules right is central here, and small mistakes can either fail to block the junk or accidentally block real pages — our complete robots.txt guide covers the pattern matching you need to target parameters precisely.
Point bots at what matters
Saving crawl budget is only half the job; the other half is directing what remains toward your best content. A clean, current XML sitemap is your most direct signal of which URLs deserve attention and when they last changed. Keeping it free of redirects, dead URLs, and non-canonical pages helps bots trust and act on it — our guide to XML sitemaps explains what to include and what to leave out. Accurate lastmod dates in particular help crawlers prioritize pages that have genuinely changed.
Beyond the sitemap, a few habits protect your budget: fix long redirect chains that make bots do extra work, return proper status codes so crawlers do not keep hammering dead URLs, and keep your important pages within a few clicks of the homepage so they are easy to reach. A shallow, well-linked structure spends crawl budget efficiently by design.
How to tell if you have a problem
Server logs are the ground truth. They show exactly which URLs bots request, how often, and how much of their activity lands on junk versus valuable pages. If you see crawlers spending the bulk of their requests on parameter URLs, faceted combinations, or pages you do not care about, that is budget you are losing. Search Console's crawl stats report gives a higher-level view of the same picture. The fix is rarely to beg for more crawling — it is to stop wasting the crawling you already get.
Curious how much of your crawl budget is being spent on the right pages? Run the CheckMy.site scanner to check your crawlability, spot the traps that waste bot attention, and make sure your important pages are the ones getting seen.