Somebody Already Decided Whether AI Can Quote You. It Wasn’t You.
Every marketing team I talk to has spent the last year worrying about whether AI will cite them. AEO, GEO, whatever the acronym is this quarter. They are rewriting headlines, adding FAQ blocks, arguing about schema. Almost none of them have checked whether the AI systems they are courting are allowed to read the site at all.
This is not a content problem. It is a permissions problem, and the permissions live in files and settings nobody on the marketing team has ever opened.
The switch you think you’re flipping
On July 1, 2025, Cloudflare announced that every new domain signing up would be asked, right there in onboarding, whether to allow AI crawlers. They called it Content Independence Day. Read the scope carefully, because a lot of coverage got it wrong. It applied to new domains at signup, not retroactively to every existing customer. Which means somebody at your company once answered a question about your AI visibility strategy. Possibly IT. Possibly a contractor. Possibly whoever was clicking through a setup flow at 4pm on a Thursday. They were not thinking about AI visibility. They were thinking about getting DNS to resolve.
Go check what they picked. Not because Cloudflare did anything wrong. Because a decision got made about your brand’s presence in AI answers and marketing was not in the room.
It isn’t one switch. It’s three per vendor.
Here’s where the whole “should we block the AI bots” conversation falls apart. There is no single AI bot. Every major vendor now runs at least three, and they do completely different jobs.
OpenAI’s own documentation spells it out. GPTBot crawls content that may be used to train foundation models. OAI-SearchBot builds the index behind ChatGPT search. ChatGPT-User fetches a page live when a person asks about it. OpenAI states plainly that each setting is independent of the others, and that a site can allow OAI-SearchBot in order to appear in search results while disallowing GPTBot. Anthropic runs the same split with ClaudeBot, Claude-SearchBot, and Claude-User. Perplexity documents PerplexityBot and Perplexity-User, and is explicit that PerplexityBot is not used to crawl content for foundation models.
So the real question was never “do we let the AI companies have our stuff.” It’s “do we want to be in the answer, in the training set, or both,” and those are separate lines in a text file. If somebody blanket-blocked everything with an AI-shaped name because a headline scared them, they didn’t protect your content. They took you out of the answers and left the harder question untouched.
And two of them don’t care what your robots.txt says
Now the part that should annoy you. The user-triggered fetchers, the ones that grab your page because an actual human just asked about your company, largely ignore robots.txt on purpose. OpenAI’s docs say ChatGPT-User is not used for crawling the web in an automatic fashion, and that because these actions are initiated by a user, robots.txt rules may not apply. Perplexity says the same about Perplexity-User: since a user requested the fetch, this fetcher generally ignores robots.txt rules.
So your robots.txt is a polite request, honored by the bots you might have wanted to block and skipped by the ones showing up because a buyer is mid-evaluation.
Unless the blocking is happening at the network layer instead. If your WAF or your bot-fight setting is the thing swinging the bat, those requests fail for real, user-initiated or not. That’s the scenario worth twenty minutes of your week: you’re not blocked in the file, you’re blocked at the edge, and nobody finds out, because a failed bot fetch doesn’t generate a ticket. It generates nothing. Silence looks identical to “we just aren’t getting cited yet.”
Google’s version is worse, and it’s worse on purpose
Google publishes a token called Google-Extended, and plenty of people assume it’s the AI opt-out. Google’s own crawler documentation states that Google-Extended does not impact a site’s inclusion in Google Search, nor is it used as a ranking signal. What it governs is training future Gemini models and grounding in Gemini apps. Separately, Google’s guidance on AI features says AI is built into Search and integral to how Search functions, which is why robots.txt directives for Googlebot are the control.
Translated: there is no AI Overviews opt-out. You can block Googlebot and leave Search entirely. You can use nosnippet or max-snippet to starve the AI summary, which also starves the normal search snippet you depend on. Apple built the same structure with Applebot-Extended. The training opt-out is free and generous. The answer-surface opt-out costs you the thing you actually need. That’s a design choice, not a conspiracy, and it means your lever is a lot smaller than the internet promised.
About that traffic you were expecting
Cloudflare publishes a crawl-to-refer ratio on Radar: roughly, how many pages an AI platform pulls from sites versus how many visitors it sends back. The ratios are lopsided in the direction you’d guess, often thousands of crawls per referred visit on the training-heavy platforms.
I’m not going to quote you a number, and you should be suspicious of anyone who does. Cloudflare’s own figures have moved by orders of magnitude between reports a few months apart, and different cuts of the same data produce wildly different ratios for the same platform in the same period. The specific number in the LinkedIn post you saw is a snapshot with the date filed off. The useful part isn’t the ratio anyway. It’s that citation and traffic have come apart, so optimizing for AI visibility while expecting a click-shaped payoff is a category error.
And no, the new file won’t save you
The llms.txt proposal keeps getting sold as the fix: a tidy file at your root telling AI systems what matters. Ahrefs looked at server logs across roughly 137,000 domains and reported in mid-2026 that the overwhelming majority of published llms.txt files were never requested at all in the month they measured, and that AI bots don’t go looking for the file on sites that lack one. No major AI company has publicly committed to reading other sites’ llms.txt in production. Several of them publish one for their own developer docs, which is a different thing entirely and keeps getting passed around as if it were proof. Adding it costs ten minutes and hurts nothing. Treating it as strategy is how you end up with a checked box instead of a result.
Who owns this, honestly
Here’s the actual problem, and it isn’t technical. Whether AI systems can read your site touches robots.txt, your CDN’s bot rules, your WAF, your meta tags, and a checkbox somebody clicked during onboarding. Marketing owns the outcome. IT owns four of the five controls. Nobody has a view of the whole thing, so it gets set once by accident and revisited never.
So go look this week. Open your robots.txt in a browser and read every line. Ask whoever runs your CDN what the bot rules do, by name, per user agent, not “we block the AI ones.” Ask them to show you whether requests from the search-side agents come back 200 or 403. That’s an afternoon of work, and it beats another quarter spent rewriting headlines for a machine that can’t get through the door.
The shape of this is familiar, which is why it bugs us: the setting that decides your numbers is almost never in the tool you’re staring at. That’s the problem THE DASHBOARD was built for, your marketing, sales, and web stack in one view with an assistant you can just ask questions. It won’t fix your robots.txt. But it does mean the rest of your numbers sit somewhere you can actually see them change.
Prefer to listen? This post is an episode of THE DASHBOARD Confessional.
See your entire stack in one place.
$1,800 a month, flat. No AI tokens, no seats, no bullshit. Onboarding in days, not quarters.
Get THE DASHBOARD →