Aurora Owl / Blog / AI crawlers

Should you block AI crawlers? Which bots to let in, and why

Block the AI training crawlers if you want to, and leave every search crawler open. Training crawlers such as GPTBot and ClaudeBot collect pages to build future models; search crawlers such as OAI-SearchBot, Claude-SearchBot, PerplexityBot, Googlebot and Bingbot decide whether an assistant can find and recommend you today.

Google and Microsoft make that split harder than it sounds, and one Cloudflare setting can block the wrong bots for you.

Ao, the Aurora Owl mascot, sits on an ink plum branch beside a hand-drawn signpost. One board points on and reads: Allow: search. The other points back and reads: Disallow: training.

What’s the difference between a training crawler and a search crawler?

Each large AI company sends more than one bot, and each has its own name in robots.txt, the file that tells crawlers where they may go. They do three jobs.

  • Training crawlers collect pages that may be used to train future AI models, as GPTBot and ClaudeBot do. Google and Apple train on what their search crawlers collect, and give you a separate name to opt out with: Google-Extended and Applebot-Extended. Blocking any of these changes nothing about what an assistant can find for a customer today, with one exception, Google-Extended, below.
  • Search crawlers build the index an assistant looks things up in when it answers. Block one and you can drop out of that assistant’s answers. OpenAI says sites that opt out of OAI-SearchBot “will not be shown in ChatGPT search answers” (OpenAI, Overview of OpenAI crawlers).
  • User fetchers open a page at the moment a person asks the assistant about it: ChatGPT-User, Claude-User and Perplexity-User. Because a person asked, OpenAI says robots.txt rules “may not apply” to its fetcher, and Perplexity says its fetcher “generally ignores” them (OpenAI; Perplexity, Perplexity crawlers).

Our advice: a business that wants to be recommended should never block a search crawler or a user fetcher. Blocking training is a fair choice, and with OpenAI, Anthropic and Apple it costs you nothing in their answers, because each lets you refuse training and stay in search (Cloudflare, Have it both ways, 15 Sep 2026). Cloudflare also says fewer than 1% of the sites on its network block search bots, while 17% block training in some way (Cloudflare, Have it both ways, 15 Sep 2026).

Letting the right bots in is step one of our guide to why AI assistants don’t recommend your business.

What does blocking each bot do to being recommended?

  • OpenAI. Blocking GPTBot tells OpenAI not to train on your content, and ChatGPT search carries on, because OpenAI treats each bot separately. Blocking OAI-SearchBot takes you out of ChatGPT’s search answers (OpenAI, Overview of OpenAI crawlers).
  • Anthropic. Blocking ClaudeBot keeps your future pages out of its training data. Anthropic says blocking Claude-SearchBot or Claude-User may reduce your site’s visibility in Claude’s search answers (Anthropic, Does Anthropic crawl data from the web?, 7 Apr 2026).
  • Perplexity. Leave PerplexityBot open. Perplexity says it “is not used to crawl content for AI foundation models”, so blocking it only costs you Perplexity’s results (Perplexity, Perplexity crawlers).
  • Google. AI Overviews and AI Mode link to pages from Google’s ordinary search index, which Googlebot builds (Google Search Central, AI features and your website). Google-Extended is a name you can block in robots.txt to keep your pages out of Gemini training, and Google says it doesn’t affect inclusion or ranking in Search. The catch: the same name also covers grounding in the Gemini app, where Google feeds Search results to the model as it writes an answer (Google for Developers, Google’s common crawlers, 14 Jul 2026). Block it, and the Gemini app can’t use your pages that way. To leave AI Overviews and AI Mode while staying in Search, use the Search generative AI setting in Search Console, which Google had rolled out to every site by 31 August 2026 (Search Console Help, Search generative AI control). We’d leave it on include, the default.
  • Microsoft. Bingbot feeds Bing and its chat answers, and Bing has no robots.txt name for training alone. Its NOARCHIVE tag stops training and also drops the page from those answers. Its NOCACHE tag keeps you in them with only your title, link and snippet, and limits training to the same (Bing Webmaster Blog, 22 Sep 2023). Cloudflare says Microsoft is building a robots.txt option, targeted for early 2027 (Cloudflare, 15 Sep 2026).
  • Apple. Applebot powers search in Spotlight, Siri and Safari, and what it collects may also train Apple’s models. Blocking Applebot-Extended opts you out of that training, and Apple says those pages “can still be included in search results” (Apple Support, About Applebot, 4 Sep 2026).

Is Cloudflare blocking AI crawlers on your site?

It might be, and its settings changed meaning on 15 September 2026. Cloudflare began asking every new domain whether to allow AI crawlers on 1 July 2025 (Cloudflare press release, 1 Jul 2025). In July 2026 it split its controls into Search, Training and Agent (Cloudflare, Your site, your rules, 1 Jul 2026). Since 15 September, choosing Block for Training also stops Googlebot, Bingbot and Applebot, and Block on pages with ads stops them on any page that shows ads; Cloudflare says either setting “impacts search as well as training”. The option that stops training and keeps search is called Disallow AI Training (Cloudflare, Have it both ways, 15 Sep 2026).

Sites that had the old Block AI bots switch on moved to Disallow AI Training, with Agent traffic, meaning user fetchers and browser agents, blocked only on pages that show ads. New sites that show ads are offered that same setup. So most sites stayed findable; the risk is anyone who picks Block. Here’s how to check yours.

  1. In the Cloudflare dashboard, open your domain’s Security settings and find Search, Training and Agent. Search and Agent should say Allow. Training should say Allow or Disallow AI Training, never either Block option.

  2. Open AI Crawl Control and its Crawlers tab. It lists each AI crawler with its allowed and unsuccessful requests (Cloudflare docs, Manage AI crawlers, 28 Jul 2026). A search crawler with mostly unsuccessful requests is being turned away. Any firewall rule or error can cause those too, Cloudflare notes, so check your rules.

  3. Open yourdomain.com/robots.txt in a browser. Cloudflare can add its own lines to the file, so read what crawlers actually receive (Cloudflare, Have it both ways, 15 Sep 2026).

What should robots.txt say if you want to be found but not trained on?

With OpenAI, Anthropic and Apple you can have both. With Google you choose between Gemini training and the Gemini app, and with Microsoft robots.txt can’t say it yet. This is the file we’d start from:

# Search crawlers and user fetchers: welcome
User-agent: *
Allow: /

# Crawlers that only collect for training: no
User-agent: GPTBot
User-agent: ClaudeBot
User-agent: Applebot-Extended
Disallow: /

# Gemini training: remove the # from the
# next two lines only if that matters more
# to you than the Gemini app's answers
# User-agent: Google-Extended
# Disallow: /

Sitemap: https://www.example.com/sitemap.xml

It works because a crawler obeys only the group that names it most closely. GPTBot follows its own Disallow, while OAI-SearchBot, which isn’t named, falls back to the open group at the top. Keep any rules you already have in the top group. Several User-agent lines can share one rule (Google for Developers, How Google interprets the robots.txt specification, 31 Aug 2026).

Two limits. The file is a request: the robots.txt standard says its rules “are not a form of access authorization” (IETF, RFC 9309, Sep 2022). And changes take a day or so. OpenAI says about 24 hours, Perplexity up to 24 hours, and Google caches the file for up to 24 hours (OpenAI; Perplexity; Google).

For Bing, we’d change nothing until its robots.txt option arrives. If training matters more to you, add <meta name="bingbot" content="nocache"> to your pages, which is the form Bing gives, and accept shorter mentions in its answers (Bing Webmaster Blog).

If you’d like us to check your robots.txt and Cloudflare settings and fix what’s blocking the wrong bots, our services page explains how we work, or you can message us on WhatsApp.

Sources

Read on 5 October 2026. Cloudflare sells the bot controls it describes, so read its figures as one company’s data.

  1. OpenAI, Overview of OpenAI crawlers (undated)
  2. Anthropic, Does Anthropic crawl data from the web, and how can site owners block the crawler? (updated 7 Apr 2026)
  3. Perplexity, Perplexity crawlers (undated)
  4. Google for Developers, Google’s common crawlers (updated 14 Jul 2026)
  5. Google Search Central, AI features and your website (updated 10 Dec 2025)
  6. Search Console Help, Search generative AI control (rolled out 31 Aug 2026)
  7. Google for Developers, How Google interprets the robots.txt specification (updated 31 Aug 2026)
  8. Fabrice Canel, Bing Webmaster Blog, Announcing new options for webmasters to control usage of their content in Bing Chat (22 Sep 2023)
  9. Apple Support, About Applebot (4 Sep 2026)
  10. Cloudflare, Cloudflare Just Changed How AI Crawlers Scrape the Internet-at-Large (press release, 1 Jul 2025)
  11. Jin-Hee Lee and Bryan Becker, Cloudflare Blog, Your site, your rules: new AI traffic options for all customers (1 Jul 2026)
  12. Bryan Becker, Cloudflare Blog, Have it both ways: stay discoverable in search while disallowing AI training (15 Sep 2026)
  13. Cloudflare docs, Manage AI crawlers (updated 28 Jul 2026)
  14. Koster, Illyes, Zeller and Sassman, RFC 9309: Robots Exclusion Protocol (IETF, Sep 2022)

How do you start?

A short brief is enough. Tell us who you are and what you want, and the quote comes to your inbox. WhatsApp is the quickest way to reach us.

Based in Dubai, working with clients in the UAE and India.