What does a robots.txt file do?

Short answer

A robots.txt file is a plain text file at the root of your website that tells search engine crawlers which parts of the site they may fetch. It controls crawling, not indexing: a blocked page can still appear in Google as a bare address. It's also not security, because obeying it's voluntary.

A robots.txt file tells search engine crawlers which parts of your website they may fetch and which parts to stay out of. A crawler is the program a search engine sends to read web pages. The file is plain text and sits at the top level of the site, at an address like yourcompany.com/robots.txt. Anyone can read it by typing that address into a browser, which is the quickest way to see what yours says.

What is inside the file

A handful of short lines. Each group begins by naming a crawler and then lists what that crawler should leave alone.

  • User-agent names the crawler a group of rules applies to. An asterisk means every crawler.
  • Disallow gives a folder or address that shouldn’t be fetched.
  • Allow makes an exception inside a disallowed folder.
  • Sitemap gives the address of the site’s XML sitemap.

The line to know is a Disallow followed by a single forward slash. Under a User-agent of asterisk, it tells every crawler to stay out of the entire site. Developers use it on a test copy of a site so the unfinished version stays out of search. If that file is carried over to the live site at launch, the business slowly disappears from Google while everything looks normal to visitors.

It controls crawling, not indexing

Blocking a page in robots.txt stops Google from reading it. It doesn’t stop Google from listing it. If other pages link to a blocked address, Google can still show that address in results as a bare link with no description, because it knows the page exists and was told not to look inside.

To keep a page out of results, the right tool is a noindex instruction, a line of code on the page itself that tells search engines not to store it. The catch is that Google has to be able to crawl the page to see that instruction. Blocking a page in robots.txt and marking it noindex at the same time defeats the second with the first.

What you want The right tool
Keep crawlers out of admin screens and internal search results robots.txt
Keep a page out of search results A noindex instruction on the page, with the page left crawlable
Keep a page private A password
Retire an old page Remove it and redirect its address to the closest current page

Obeying it’s voluntary

Google and the other major search engines respect the file. A scraper or a malicious bot ignores it. Because the file is public, listing a private folder in it also tells anyone curious where that folder is. Nothing sensitive should rely on it.

The same file is where a site states whether the crawlers run by AI companies may fetch its pages, since each of them announces itself under its own name. Whether a local business should block AI crawlers is a separate decision.

What a home-service site should block

Very little. The admin area and internal search results are reasonable. Beyond that, the file on a contractor’s site should be short. Don’t block the style and script files the pages are built from, since Google needs them to see the page the way a visitor does.

You may read advice about using robots.txt to save crawl budget, the amount of a site Google is willing to fetch in a given period. That’s a concern for very large sites. A site of a hundred pages has nothing to save.

How I check it

I read the file, then confirm with Google. Search Console reports which robots.txt file Google found and whether it could read it, and the URL Inspection tool says whether a particular page is blocked. I check it again after every redesign or change of host, since that is when a test-site file is most likely to be left in place. It’s one of the first checks under indexing and crawling, and a routine step in site migrations.

Check what your robots.txt file is blocking.

Send me your web address, and I'll read the file and tell you whether it's keeping Google out of anything it should see.

Prefer to talk? Call (619) 675-7678, 9am to 6pm.

Step 1 of 2: your business

Your business

Where to reach you

No long contracts. You talk to Simon.

Call Get My SEO Audit