
Updated September 22, 2026
You don't want AI companies using your website however they want. Fair enough.
So you see a setting that says Block AI or Block Training, and turning it on sounds like an easy decision.
Except that one setting can affect a lot more than AI training.
Cloudflare rolled out new AI crawler controls on September 15 that give website owners more control over how search engines, AI systems, and automated agents access their content. The update also exposes a problem most business owners have probably never had a reason to think about:
The bot crawling your website for Google Search may also be doing something else.
Googlebot, Bingbot, and Applebot are examples of what Cloudflare calls mixed-use crawlers. They can serve more than one purpose. That means a hard block intended to stop AI-related activity can also stop the crawler you need to remain visible in search.
Cloudflare now offers a way to separate those decisions more carefully. But if you use Cloudflare, this is a good time to check what your website is actually allowing and blocking.
Because "block AI" is no longer just an AI decision.
It's a search visibility decision too.
What Changed With Cloudflare's AI Crawler Controls?
Cloudflare has been expanding the controls website owners have over automated crawlers.
Its AI Crawl Control tools let website owners see which AI services are accessing their content and create policies for individual crawlers. Cloudflare also distinguishes between different reasons automated systems may want your content, including search, AI agents, and model training.
The distinction matters.
A contractor may be perfectly comfortable letting Google crawl an HVAC service page so someone searching for "AC repair near me" can find it.
That same contractor may feel very differently about allowing a company to collect thousands of pages for model training.
Until recently, separating those uses could get messy when the same crawler performed both jobs.
Cloudflare's September 15 update addressed part of that problem by introducing a new option called Disallow AI Training. It allows a website to express that its content should not be used for AI training while keeping mixed-use crawlers available for search.
But there is an important distinction between disallowing training and blocking the crawler.
Choose the wrong one, and you can still affect search.
Search, AI Agents, and AI Training Are Not the Same Thing
Here's the simplest way to think about the different types of crawler activity.
Search: A crawler accesses and indexes your content so it can appear in search experiences.
AI input or agent activity: An AI system accesses information in real time to help answer a person's question or complete a task.
AI training: Content may be collected to train or fine-tune an AI model.
Cloudflare's Content Signals system reflects similar distinctions with directives for search, ai-input, and ai-train.
For a home service company, those differences aren't academic.
Imagine you own a plumbing company in Dallas.
You probably want your water heater pages available when someone searches Google for a plumber. You may also want your company information available when someone asks an AI platform for plumbing companies serving their neighborhood.
Whether you want your website content used to train future AI models is a separate decision.
The problem comes when all three activities get treated as though they're the same thing.
They aren't.
Can Blocking AI Crawlers Block Googlebot?
Yes, depending on which Cloudflare setting you use.
Cloudflare says that its Block and Block on pages with ads settings now apply to mixed-use crawlers, including Googlebot, Bingbot, and Applebot. If you use those controls against a mixed-use crawler, you can affect its search access too.
Cloudflare introduced Disallow AI Training as the alternative for businesses that want to remain discoverable in search without allowing their content to be used for training.
That's the distinction worth checking.
Disallow AI Training is not the same as Block.
One expresses and synchronizes a preference against AI training while preserving search access for accountable mixed-use crawlers. The other can stop the crawler itself.
And if you stop Googlebot from reaching your website, you've created a much bigger marketing problem.
Why Crawler Access Matters to Your SEO
Google cannot reliably understand changes to pages it cannot crawl.
If you've added a new HVAC service, expanded into three new cities, changed your financing information, or published a detailed answer to a common homeowner question, search engines need access to discover and process that information.
Blocking crawling and preventing indexing are not even the same technical action.
Google specifically notes that if a page is blocked from crawling, Googlebot cannot access the page to see a noindex directive. In some circumstances, the URL could still appear in search based on information Google finds elsewhere.
Bing similarly warns that preventing its crawler from accessing a page can prevent that page from being indexed.
This is why crawler configuration shouldn't be treated as a random security toggle.
It sits upstream from your visibility.
And This Is Bigger Than Traditional SEO
There's another reason contractors should pay attention.
Your website isn't only being read by humans and traditional search engines anymore.
AI assistants, answer engines, agents, search engines, and training systems are all requesting web content for different reasons.
That changes the question from:
"Should we block AI bots?"
to:
"Which systems should have access to which content, and for what purpose?"
That's a much better question.
A home service company trying to grow visibility may want its public service information widely discoverable. You want systems to understand that you provide heat pump installation in Phoenix, emergency plumbing in Atlanta, or electrical panel upgrades in Denver.
You may want Google to index that information.
You may want an AI answer engine to reference it when a homeowner asks for a local contractor.
You may want an AI agent acting on behalf of a homeowner to access business information.
You might still decide that you don't want the same content used for model training.
Those are four separate considerations.
Your crawler strategy should reflect that.
What Should Home Service Companies Check Right Now?
You don't need to become an expert on bot infrastructure. But someone responsible for your website should know the answers to a few questions.
1. Is Your Website Using Cloudflare?
Start there.
If Cloudflare sits between your website and visitors, its security and crawler rules can affect whether automated systems ever reach your website.
2. What Are Your Current AI Crawler Settings?
If you use Cloudflare, review AI Crawl Control and your bot settings.
Don't just look for whether something says "AI."
Check what is happening with Search, Agent, and Training access and whether you've chosen Block or Disallow AI Training.
Cloudflare says existing customer preferences were migrated into its new controls, so don't assume your configuration is correct simply because nobody on your team changed anything last week.
3. Can Googlebot and Bingbot Still Reach Important Pages?
Verify rather than assume.
Check your Google Search Console data for crawling and indexing problems. Review Bing Webmaster Tools. If you have access to server or CDN logs, look at what response codes important crawlers are receiving.
A sudden pattern of blocked responses deserves investigation.
4. What Does Your Robots.txt File Say?
Cloudflare's controls aren't the only place crawler access can be managed.
Your robots.txt file may contain instructions about which bots can access different parts of your site. Cloudflare also now supports Bot Preference Sync, which can align robots.txt information with the AI bot preferences configured in Cloudflare.
That makes it worth reviewing the effective file rather than assuming the version someone remembers configuring months ago is still exactly what crawlers are seeing.
5. Are Other Security Rules Blocking Bots?
Cloudflare's Web Application Firewall and other bot-management rules can also affect crawler traffic.
Cloudflare notes that AI Crawl Control uses WAF rules to enforce crawler blocks, and other WAF configurations can interact with crawler management.
So if Google Search Console says Googlebot can't reach pages but your crawler settings look fine, don't stop there.
The block may be happening somewhere else.
Don't Make Crawler Decisions in Isolation
This is where the conversation should move beyond a September Cloudflare update.
For years, crawler management was mostly an SEO technicality.
Make sure Google can crawl the site. Block obvious junk. Maintain robots.txt. Move on.
That world is changing.
Now a crawler might be indexing information for search, retrieving information for an AI answer, acting on behalf of a user, or collecting information for model training.
And sometimes one crawler serves more than one of those purposes.
Cloudflare's September update is one example of the web infrastructure adapting to that reality.
Businesses need to adapt too.
"Block AI" Is Now a Marketing Decision
The instinct to protect your website content makes sense.
But don't let a broad technical rule undermine the visibility you're paying to build.
For home service businesses, your website needs to do more than exist. Search engines need to discover it. AI systems need enough reliable information to understand your company. Potential customers need to find accurate information about what you do, where you work, and why they should consider you.
That requires a deliberate crawler-access strategy.
So before someone clicks Block, ask what you're actually blocking.
Because the question isn't whether your website should be open or closed to "AI."
The question is who needs access, what they're using that access for, and what happens to your visibility when you take it away.
Frequently Asked Questions
Can blocking AI crawlers hurt SEO?
Yes. If the rule also blocks a crawler used for traditional search, it can interfere with crawling and eventually affect search visibility. Cloudflare specifically warns that its Block controls can apply to mixed-use crawlers including Googlebot, Bingbot, and Applebot.
Does blocking AI training automatically block Google?
Not necessarily. Cloudflare's Disallow AI Training setting is designed to let websites refuse AI training while keeping accountable mixed-use crawlers available for search. A hard Block, however, can block the crawler itself and therefore affect search access.
What is a mixed-use crawler?
A mixed-use crawler is a crawler that serves more than one purpose. Cloudflare identifies Googlebot, Bingbot, and Applebot as examples because their operators use them for search as well as AI-related purposes.
What is the difference between AI search and AI training?
AI search or AI input uses web content to help retrieve or generate answers for users. AI training uses content to train or fine-tune models. Cloudflare's Content Signals distinguish among search, AI input, and AI training uses.
Should contractors block AI crawlers?
There isn't one setting that makes sense for every business or every crawler. Contractors should determine whether a crawler supports search discovery, AI answers, agent activity, model training, or multiple purposes before blocking it. For businesses that depend on organic and AI-driven discovery, crawler access should be reviewed as part of the overall search visibility strategy.



