Docket / Fix it / Rate limits

A rate limit is not a block

An audit is traffic. It asks for page after page, faster than a person would, from one address — which is the exact shape of the thing rate limiting exists to slow down. So a crawl can trip a limit, and then every measurement it takes afterwards is a measurement of the limit it tripped. The report that comes out says your server is turning crawlers away. Nobody is turning anyone away.

The audit that caused the problem it found

Docket crawls a batch of pages, then asks for the homepage once more as each AI crawler in turn to see which of them the edge lets through. On one retailer that sequence was enough: the crawl consumed the allowance, every crawler probe after it came back HTTP 429, and Docket reported that the server had refused the audit and that the site's bot protection distrusted the network it ran from.

A browser from the same machine loaded the page a couple of minutes later, twice, without complaint. A re-run of the audit passed clean. Nothing about the site had changed in between, because nothing about the site was ever the problem — the tool had caused the condition it was reporting, and then described it as a property of somebody else's server.

This is the shape to keep: a measurement that a re-run can change is a measurement about the run, not about the site. Before you act on any crawler-access finding, run it again, slower. If it goes away, it was about you.

A throttle names a wait; a refusal names a door

HTTP 429 means too many requests, and HTTP 503 means not right now. Both are temporary and both are usually correct behaviour by a server that is protecting itself. A 403 is a different sentence: it says this requester is not allowed, and it will still say that tomorrow at a slower pace.

The tell that settles it is the Retry-After header. A server that sends one is naming a time to come back — it is scheduling the crawler, not refusing it. Well-behaved crawlers honour it and return. A rate limit with a Retry-After is not a visibility problem at all.

Docket got this wrong for a while and it is worth saying how, because the error is easy to repeat. On a forum, every AI crawler received 429 with a short Retry-After while a browser from the same address got 200. Docket called that refusal, rated it critical, and sent the owner to their CDN to add an allow rule for a WAF block that did not exist. The observation was right and the diagnosis was wrong, which costs the reader an afternoon and their provider a support ticket.

The one-click test that settles a 403

A genuine 403 has two possible causes and they need opposite fixes: the rule is on who is asking, or the rule is on where they are asking from. Open the same URL in your own browser, on your own connection:

Docket had this backwards, and the reason is worth more than the fix. Its crawler sends a browser-shaped user-agent with an honest suffix naming the tool and linking to a page about it — which is precisely the substring a WAF rule matches on. But the check's own text called that "a plain browser request", so when it was refused it concluded the block could not be about identity and must therefore be about the network, and it told the reader to try a different connection. A clean browser string from the same machine got 200 and a full page. The network had never been involved.

An instrument that misdescribes its own request will misattribute the response. When a tool reports a refusal, the first question is what it actually sent.

A true observation with a false consequence

The third variant is the hardest to catch, because the fact is correct. Docket flags AI crawler tokens in robots.txt that no longer belong to any crawler — a retired name sitting in a file that has not been revisited. On a manufacturer's site it found several, all genuinely dead, and said those crawlers were therefore reading pages the owner meant to withhold from them.

They were not. A crawler that matches no group of its own falls back to User-agent: *, and on that file the wildcard group was the stricter of the two — it carried every rule the AI group carried, and more. The dead tokens cost nothing but a crawl-delay. The tokens were dead; the consequence was invented; and the severity of the finding rested entirely on the consequence.

So when a finding tells you what something costs, check whether the tool measured the cost or inferred it. Inferring it means modelling the fallback behaviour of whatever reads the file, and most tools do not.

When none of this matters

Said plainly, because the checks above are rated where they are for a reason and not every one of them is worth your afternoon:

The fixes that make it worse

How to make the finding disappear without changing anything

Worth knowing, because it tells you what the finding is worth. Crawl one page at a time with a pause between requests and the rate limit never fires, so the crawler probes all return 200 and the report is clean. Nothing about your site is different. The same is true in reverse: crawl harder and a clean site produces refusals.

Which means the absence of this finding proves less than its presence, and its presence proves less than a second, slower run that reproduces it. Any tool that reports edge access without telling you how hard it crawled is handing you a number it cannot support.

Where this sits in an audit

The two registered checks are ai.edge_access, which tests server access for AI crawlers, and ai.dead_crawler_directive, which reads retired tokens in robots.txt. Both live in Docket's AI search visibility area.

Their findings carry different identifiers from the checks that emit them — ai.throttled_at_the_edge and ai.edge_untestable both come out of ai.edge_access — which is worth knowing when you search for one and find nothing. A finding identifier is not a check identifier.

The untestable one is the useful habit here: when a run cannot measure something, the honest output is a notice saying so, not a pass. A gate that could not measure is not a pass →

Common questions

Is HTTP 429 a block?

No. It means too many requests, and it is temporary. A server that sends it with a Retry-After header is naming a time to come back rather than turning the crawler away. HTTP 403 is the refusal.

Why did my audit report crawler blocks that a re-run did not?

Because the first run probably caused them. A crawl is traffic from one address, which is what rate limiting slows down, so the probes that follow the crawl can all be measuring the limit the crawl tripped.

My site returns 403 to an audit tool. Is that a problem?

Only if you did not intend it. Open the same URL in your browser: if it loads, the rule is on the user-agent and lives in your bot protection; if your browser is refused too, the rule is on the network.

Should I remove retired AI crawler names from robots.txt?

It is tidiness rather than visibility, unless the wildcard group those crawlers fall back to is more permissive than the group the retired name sat in. Compare the two groups before editing.

How do I know a crawler-access finding is real?

Run the audit again, more slowly, and see whether it survives. A measurement that a re-run can change is a measurement about the run rather than about the site.