robots.txt and sitemap.xml: what to put and what to avoid

They are two text files that almost every website has and almost nobody checks, and both are easy to get wrong. The most misunderstood is robots.txt: it does not hide pages from Google, and blocking a page with it can achieve exactly the opposite of what you want. What each file does, the real rules from Google's documentation and the mistakes that keep recurring.

Reviewed on 3 October 2026 · 5 min read

What each one does

robots.txt tells crawlers which parts of your site they may request and which they may not. sitemap.xml does the opposite: it shows them which pages you want them to discover. One restricts, the other invites. Neither is mandatory, but a small site without them relies on internal links being enough to reach everything, and a badly written robots.txt can leave the whole site out of crawling.

robots.txt: what to know

Where it goes and what it must look like

The file is called exactly robots.txt, sits at the root of the site (https://yourdomain.com/robots.txt, never in a folder) and there is only one per site. Its rules apply only to the protocol, host and port where it is served: those in https://example.com do not apply to https://m.example.com or to http://example.com. It must be plain text in UTF-8; if you write it in a word processor, curly quotes or other characters can sneak in and invalidate the rules. Google enforces a 500 KiB limit: anything beyond it is ignored.

The minimum syntax

User-agent: *
Disallow: /admin/
Disallow: /cart/
Allow: /admin/help.html

Sitemap: https://yourdomain.com/sitemap.xml

Each group starts with User-agent (who it applies to; the asterisk means all) and continues with Disallow rules (do not crawl) and Allow rules (an exception inside a blocked folder). Comments start with #. Paths are case-sensitive: /Admin/ and /admin/ are different. Google supports two wildcards: * (any sequence of characters) and $ (end of the address), so Disallow: /*.pdf$ blocks every address ending in .pdf.

When two rules clash, the most specific one wins, the one with the longest path; and if they still conflict, Google applies the least restrictive. That is why, in the example, /admin/help.html can be crawled even though /admin/ is blocked.

Google only recognises four fields: user-agent, allow, disallow and sitemap. Others, such as crawl-delay, it ignores.

The most repeated mistake: using it to hide pages

Google puts it bluntly in its documentation: robots.txt is mainly for managing crawl traffic, not for keeping a page out of Google. If other pages link to yours, Google can index the address without visiting it, and it will appear in results with no description. To keep a page out, you need noindex (in a meta tag or an HTTP header) or password protection.

And here is the trap: if you block the page in robots.txt, the crawler never gets to see the noindex, and the page can keep appearing. The two cancel each other out. To get a page out of results, leave it crawlable and mark it noindex.

Besides, robots.txt is not a lock: serious crawlers respect it, but nothing forces the rest to. If there is private data, the answer is a password, not a line of text.

Other common failures

  • Leaving a forgotten Disallow: /. It is the rule that blocks the whole site and is often copied over from the test version. Check your robots.txt after every major release.
  • Blocking resources the page needs. If you stop Google loading the CSS or JavaScript without which the page cannot be understood, it will analyse it worse. Google recommends not blocking those files.
  • Writing Sitemap: with a relative address. It must be the full address, with protocol and domain, exactly as it is visited (with or without www).

AI crawlers

Besides search engines, there are now crawlers that collect content to train AI models, and robots.txt is the route for deciding what to do about them. Two examples documented by the companies themselves: OpenAI uses GPTBot for training content and OAI-SearchBot to appear in ChatGPT's search features, and they are independent: you can block one and allow the other. Google offers Google-Extended, a control token for deciding whether the content it crawls may be used to train its Gemini models; according to Google, it does not affect inclusion or ranking in Google Search.

User-agent: GPTBot
Disallow: /

User-agent: Google-Extended
Disallow: /

As with all robots.txt, it is a request that respectful crawlers follow. Think before blocking: leaving out a search crawler also means not appearing in those results.

sitemap.xml: what to know

A sitemap is a list of addresses in XML. The minimal version looks like this:

<?xml version="1.0" encoding="UTF-8"?>
<urlset xmlns="http://www.sitemaps.org/schemas/sitemap/0.9">
  <url>
    <loc>https://yourdomain.com/page</loc>
    <lastmod>2026-10-03</lastmod>
  </url>
</urlset>
  • Limits: a sitemap holds up to 50,000 addresses or 50 MB uncompressed. If you go over, it is split into several and linked from an index file.
  • Encoding and location: UTF-8, and preferably at the root: a sitemap only affects the addresses in its folder and below, unless you submit it through Search Console.
  • Full addresses: with protocol and domain, as they are visited. Google will crawl them exactly as you write them.
  • Only what you want to appear: canonical addresses that return 200 and carry no noindex. Including redirects or error pages confuses more than it helps (if you doubt what code a page returns, see the HTTP status codes guide).
  • lastmod only if it is true. Google uses it when it is consistent and verifiable, and it should reflect a significant change to the page: the content, structured data or links, not the footer's copyright year.
  • priority and changefreq do nothing for Google: it ignores them. Do not waste time tuning them.

To tell search engines you have two routes, which can be combined: the Sitemap: line in robots.txt and submission through Search Console, which also gives you a report on how many addresses were read.

To generate them without slips

The robots.txt and sitemap.xml generator builds both files from a form, with templates for the most common cases and rules for AI crawlers in one click. Check the result carefully, especially the Disallow lines, and upload the files to the root of your site. Then open yourdomain.com/robots.txt in a browser to confirm it shows. And if you are looking after the page's SEO, also see how its title and description will look with the length checker and its preview when shared with Open Graph.

Sources and further reading

Figures checked on 3 October 2026.

Do it now, free, in your browser. Your files are not uploaded.

Create a robots.txt and a sitemap.xml for your site and download them.

Frequently asked questions

Does robots.txt hide a page from Google?
No. Google says robots.txt is for managing crawl traffic, not for keeping pages out of results: if other sites link to the page, the address can be indexed without a description. To hide it, use noindex or a password.
Why shouldn't I block a noindex page in robots.txt?
Because if the page is blocked, the crawler never sees the noindex tag and the page can still appear in results.
Where should the robots.txt file be?
At the root of the site (yourdomain.com/robots.txt), with that exact name, in UTF-8 and only one per host, protocol and port.
How many addresses can a sitemap hold?
Up to 50,000 addresses or 50 MB uncompressed. If there are more, they are split into several sitemaps linked from an index.
Does Google use the priority and changefreq tags?
No, Google ignores them. It does use lastmod, but only when the dates are consistent and verifiable.
How do I block AI training crawlers?
With rules in robots.txt for each one, for example User-agent: GPTBot and User-agent: Google-Extended, both with Disallow: /. Google-Extended does not affect Google Search, according to its documentation.