CRAWL-01
robots.txt asks for a delay the audit could not honor in full
What the check measures
Like SCOPE-02, this finding does not describe a fault on your site. It describes what we had to do during the audit, and why.
Your robots.txt carries a Crawl-delay directive asking for a longer gap between requests than one audit can absorb. We therefore capped the value at 5 seconds. The report states both numbers: what the file asked for and what we actually used.
Those 5 seconds are not a politeness guess but a calculated limit (config/). A paid audit crawls up to 200 pages and has 1,800 seconds for the whole run. At a 5-second delay the crawl alone takes roughly 1,000 s, leaving ~800 s for analysis and file generation, more than a normal audit takes end to end. At 10 s the crawl alone would want 2,000 s and the audit would hit the wall-clock limit for certain, even if everything else ran at full speed. Five seconds is the highest value that still fits the budget.
The slower of the two always wins: our own pace (two requests per second per domain, i.e. 0.5 s) and your value. Someone else's file can only slow us down, never speed us up; Crawl-delay: 0.1 does not move us. The group is picked the same way as for Disallow: the most specific match on our token beats User-agent: *.
What the check does not do:
- It does not appear when we honored the delay in full. A value of 5 seconds or less slows us down and produces no finding. We respected it completely.
- It does not judge whether your value is sensible. A server that genuinely cannot take more than one request every ten seconds is right to say
Crawl-delay: 10. The finding only says we did not honor it. - It neither ignores a broken value silently nor crashes on it.
Crawl-delay: abc, a negative number or “10 seconds” is skipped and we use our own pace.robots.txtis untrusted third-party input. - It does not shorten the audit. The delay affected pace, not page count. How many pages were crawled and why is
SCOPE-02. - It does not read other directives. Disallow rules are still enforced by the crawler independently. That is
ACC-01.
The finding is site-level and informational. It does not affect the score at all (config/, the info weight is 0). Your score should not drop because of how you configured your own server.
How strong the evidence is
We recommend it because it does no harm or has some other benefit, but we promise nothing about whether it makes language models cite you. Nobody has demonstrated that yet.
It has no bearing on visibility in AI answers in either direction and this finding recommends nothing. The “effect not documented” class is here because we have no other class for a finding that is not a recommendation.
Something else is documented, though, and it matters if you use Crawl-delay as server protection: as protection it is unreliable, because most crawlers do not read it.
| Crawler | Honors Crawl-delay? |
|---|---|
| Googlebot | No. Google's own documentation lists the supported fields and adds: “other fields such as crawl-delay aren't supported”. |
| Bingbot | Yes, per Bing's own documentation. |
| Yandex | No, since 22 February 2018. It did before; the rate is now set in their webmaster tools (their own documentation). |
| Amazonbot | No. Amazon's own documentation says so explicitly: “They do not support the crawl-delay directive” (verified 2026-09-14). |
| Other AI crawlers | We do not know. We went through the primary documentation of the crawlers we track (config/ai-bots.json). Apart from Amazonbot, none of them mentions Crawl-delay. That does not mean they ignore it; it means none of them promises otherwise. |
One more thing worth saying plainly: Crawl-delay is not part of the standard. RFC 9309, published in 2022, standardised User-agent, Allow, Disallow and Sitemap, nothing else. The directive is an extension originally introduced by Yahoo/MSN, and every crawler interprets it its own way: some as seconds between requests, some as a weighting.
The practical conclusion: against a crawler that is hammering your server, Crawl-delay is a request, not a rule. Anyone who does not read it is not constrained by it. What works is a server-side rate limit or a WAF rule. Those apply to everyone regardless of what they read.
How to fix it
First decide which of two situations you are in. They are entirely different cases and easily confused:
| Situation | What to do |
|---|---|
| The value matches your server's real capacity | Leave it alone. It only slowed the audit down, it did not cut it short. You got the full scope, it just took longer. Treat the finding as confirmation that we noticed your request. |
| The value is a typo or a leftover from an old configuration | Lower it. Values like Crawl-delay: 86400 (one request per day) do occur in the wild and are usually accidents, typically copied from someone else's template. |
How to tell which: look at your server logs for response times under concurrent requests. A site that handles visitor traffic will usually handle two requests per second from a single crawler.
If you really do need to slow crawlers down, use something that also binds those who never read Crawl-delay:
- A server-side rate limit or a WAF rule: the only option that works regardless of a crawler's goodwill.
- Google Search Console: Googlebot's crawl rate is managed there, not in
robots.txt(Google has a page on it, linked below). - Bing Webmaster Tools: Bingbot does honor
Crawl-delay, but you can also set the rate directly in their tools. - Respond
429 Too Many Requests: the standard way to tell a crawler to back off. Unlike403and401, crawlers act on it; Google says so explicitly in its crawl-rate documentation.
And what not to do: do not raise the value to keep our audit away. A higher number will not slow it beyond those five seconds, and all you gain is slower crawling from the crawlers that do care about your content.
What the report says about it
Finding description
Your robots.txt asks for a [declaredSeconds]s delay between requests (the Crawl-delay directive). We slowed the crawl down because of it, but only to one request every [appliedSeconds]s. A longer delay does not fit the time budget of a single audit, and the audit would have failed before it got through the site. So we crawled more slowly than we normally do (two requests per second per domain), just not as slowly as the file asks.
Recommendation
If the value is a typo or a leftover from an old configuration, lower it. Bingbot honors Crawl-delay by its own documentation, so a value this high means crawling a larger site takes it days or weeks. You cannot rely on it elsewhere: Yandex has not taken it into account since 22 February 2018 by its own documentation, and the AI crawlers we track either do not mention it in theirs at all or explicitly deny supporting it. Amazonbot, by its own documentation, does not support the Crawl-delay directive. Googlebot ignores the directive entirely by its own documentation (it sets its own crawl rate, adjustable in Search Console), so as protection against server load it is unreliable; a server-side rate limit or a WAF rule does that job better. Crawl-delay is also not part of the robots.txt standard (RFC 9309); it is an extension that every crawler interprets its own way. If the value does match your server's real capacity, leave it alone: it only slowed the audit down, it did not cut it short.
Sources
- Google Search Central: Robots.txt specification (supported fields) (accessed 2026-09-13)
- Google Search Central: Reduce the Googlebot crawl rate (accessed 2026-09-13)
- Yandex Webmaster: Crawl-delay (support discontinued 22 Feb 2018) (accessed 2026-09-13)
- Amazonbot: Amazon's own documentation (does not support Crawl-delay) (accessed 2026-09-14)
- RFC 9309: Robots Exclusion Protocol (accessed 2026-09-13)
Text verified 2026-09-13