Seven sites kept the crawler that cites them and blocked the one that trains, and none did the reverse
Seven sites let in OpenAI's search crawler, the one that can name you in an AI answer, while turning away its training crawler, the one that feeds future models. Not one site did the opposite. On 17 September 2026 I took a random sample of 400 sites from the Crane Index™ and asked each the same thing twice from the same address at the same moment: once as an ordinary browser, then as four AI agents. Any difference between the two answers is the visitor's name and nothing else.
Seven is a small number and I will not dress it up as a trend. Its force is that across 400 sites nothing contradicted it. Where a business appears to have taken a position, it is the same position every time: keep the citation, decline the training.
One site in nine served a person and refused a machine
Of the 320 sites that served an ordinary browser, 36 refused at least one AI agent, and 24 of those shut out all four. That is 11.25 per cent of the sites I could read, and on a sample this size the true figure is very likely between 8.2 and 15.2 per cent. Same door, same instant: a person gets the page, a machine gets a refusal.
Only one site of the 400 faked a missing page rather than refusing outright, so deception is rare. The rest simply said no, cleanly.
One caveat, stated plainly. This is behaviour at the edge, the layer that answers before your website does, and not published policy. I did not read anyone's robots.txt, the file where a site states which crawlers it welcomes, and I probed homepages only. So I can tell you how often a door was shut, not why, and I will not name a company as an offender on evidence this narrow.
Most of those refusals show no sign of anyone deciding
Twenty-nine of the 36 turned an agent away with no pattern behind it: no split between the crawler that cites and the crawler that trains, nothing that reads as a position. I cannot see intent from the edge, so I will not call them accidental. What the numbers support is narrower and more useful: only a fraction of refusals look like anyone weighed them.
That fits how the setting usually arrives. New domains on Cloudflare can block AI crawlers by default, which is my own observation from configuring domains rather than a figure I can source to a Cloudflare announcement. A default is still a decision. It is just somebody else's.
The distinction those seven sites drew is the one worth understanding, because the two crawlers pay you differently. One fetches your page so it can be quoted, with your name on it, in an answer a buyer is reading now. The other collects text to train future models and returns nothing you can measure this quarter. They can be allowed and refused separately. A business that blocks without looking usually loses the first in order to stop the second.
The biggest companies are the hardest for a machine to read
Budget does not buy machine-readability, and if anything it buys the opposite. In the full census on 1 September 2026 I approached 3,336 sites and could read 2,291, so 31.3 per cent did not serve the scanner at all. The gap was widest in the FTSE 100 at 49.4 per cent and narrowest in professional services at 15.7 per cent.
One honest qualification. That census runs from cloud servers, and the same 400 domains refused about 20 per cent of the time from an ordinary home-style connection, so part of the gap is where the scanner stands rather than a change in the sites. It is still the right vantage point, because the AI crawlers that matter, GPTBot and ClaudeBot among them, are cloud crawlers too. They stand roughly where my scanner stands and see the harder version of your site. And the figure is flat: it moved a tenth of a point in a month, so this is a standing condition, not a trend to panic over.
Take the decision back in about a minute
Your edge has already answered the AI question, and you can find out what it said in about a minute. Ask whoever runs your hosting or your technical setup one question:
When an AI agent such as GPTBot, ClaudeBot or OpenAI's search crawler requests our homepage, does it get the page or a refusal, and did we choose that?
If the answer is a refusal nobody chose, you are in the eleven per cent, and the fix is a setting rather than a project. If it is a refusal you did choose, take the two crawlers separately: the one that cites you is the one that can send you a buyer. Either way, a commercial decision this size belongs with you, on the record, not with a default you have never seen.