Should I block AI crawlers from my website?
Short answer
For most local service businesses, no. Blocking AI crawlers stops assistants from reading your own pages, and being read is how you get named. It doesn't remove you from AI answers, which can still draw on directories and reviews. The choice is made in robots.txt or at the firewall, and is worth checking for accidental blocks.
For most local service businesses, no: blocking AI crawlers keeps your own website out of what assistants can read, and being read is how you get named. A home-service company publishes its services and service area, with a phone number, because it wants them found. There are sound reasons to block, but they mostly belong to businesses whose content is the thing they sell.
What blocking does and doesn’t do
A block stops a crawler from fetching your pages. Your business stays in AI answers, because an assistant can still describe you from directories, review sites and articles written by other people. So for a contractor, the practical effect of a block is that your own account of your business is the one source left out of the answer.
The trade-off
| Choice | What you gain | What it costs |
|---|---|---|
| Allow AI crawlers | Assistants can read your services, cities and hours from you | Your text may be used in training, or summarized without a visit to your site |
| Block only crawlers that collect training data | Your content stays out of future training, where the crawler honors the request | Little today, though a model may hold less about you over time |
| Block all AI crawlers | The most control over copying | Your pages can’t be read or cited when a customer asks about you |
The middle row is possible because several AI companies publish the names of their crawlers and say what each is for, and some separate the one that gathers training data from the one that fetches a page when a person asks a question. Those names and purposes change. Check each company’s current documentation before acting on a list from a blog post.
Google is a separate case. Its AI Overviews are part of Google Search and rely on Google’s normal crawling, so blocking Google’s main crawler would remove you from ordinary search results as well.
Where the choice is made
One place is the robots.txt file, a text file at the top level of your site that tells named crawlers what they may fetch. It’s a request. Well-behaved crawlers follow it, and nothing forces the others to.
The other is the firewall or content delivery network that sits in front of your site, which can refuse a request outright. That’s enforcement, and it’s also where accidents happen.
Check that you aren’t blocking by accident
For a small business, an accidental block is more likely than a deliberate one. A security plugin or a hosting setting can challenge or refuse automated visitors, and so can a bot-protection feature. Any of them may have been switched on by a previous developer or by a default you never chose. Nobody notices, because the site looks normal in a browser.
Two checks take a few minutes. Type your web address followed by /robots.txt and read what it says. Then ask whoever manages your hosting which bots are being challenged or blocked. I look at both in an AI visibility audit, alongside the wider indexing and crawling question of what search engines can reach.
When blocking is reasonable
A publisher or a photographer has something to lose if the work is absorbed and reproduced, and so does anyone selling a paid course. Some owners also object to their writing being used for training as a matter of principle. That’s a legitimate position, and blocking the training crawlers is the way to act on it. I’d only ask that you make the decision knowing what it costs in visibility. I can’t advise on the legal side of how published content may be used. Rules in that area are unsettled, so check the current position before relying on it.