SEO

Robots.txt and XML sitemaps: how to get your pages indexed

How robots.txt and XML sitemaps work, which mistakes stop indexing and how to ask Google to index a page through Search Console.

Two small files control a lot of how Google finds your website. Robots.txt tells search engine robots which parts of the site they may visit. The XML sitemap lists the pages you want them to find. If either of them is wrong, important pages can be missing from Google entirely, without visitors noticing anything.

This guide covers both files, the most common mistakes and how to get a page indexed.

How indexing works

For a page to be shown on Google, it has to pass through three steps.

  1. Discovery: Google finds the address, usually via a link from another page or via your sitemap.
  2. Crawling: Googlebot fetches the page and reads the content. This is where robots.txt controls what the robot may fetch.
  3. Indexing: Google assesses the page and adds it to its index. Only then can it appear in the search results.

Google does not index everything it crawls. Pages with thin content, duplicates and pages with a noindex instruction are left out. So indexing a page means Google has chosen to keep it, and that is a decision you can influence but not make.

What is robots.txt?

Robots.txt is a plain text file in the root of your domain, for example yourdomain.com/robots.txt. It has to be exactly there for search engines to read it. Each hostname has its own file, so a subdomain such as shop.yourdomain.com needs its own robots.txt.

The file is public. Anyone can open it in a browser, so it should never be used to hide sensitive content. It is an instruction to well-behaved robots, not password protection.

The syntax

The file consists of groups of rules. Each group starts with which robot the rules apply to, followed by what it may and may not visit.

  • User-agent states which robot the group applies to. An asterisk means all robots.
  • Disallow states a path the robot may not visit.
  • Allow makes an exception to a Disallow rule.
  • Sitemap gives the full address of your sitemap. The line can go anywhere in the file.

A typical example for a business site might look like this:

  • User-agent: *
  • Disallow: /admin/
  • Disallow: /cart/
  • Allow: /admin/public-page/
  • Sitemap: https://www.yourdomain.com/sitemap.xml

Rules match the start of the path, and case matters. When several rules match the same address, Google follows the most specific one, meaning the one with the longest path. Google also supports the asterisk as a wildcard and the dollar sign to mark the end of an address.

An empty Disallow line means everything is allowed. If the file is missing entirely, Google also takes that to mean the whole site may be crawled.

Common mistakes

Disallow: / left over from development. While a new site is being built, the whole site is often blocked with a single line, Disallow: / under User-agent: *. If the line is left in at launch, Google stops crawling the site. Check robots.txt the same day a new site goes live.

Blocked CSS and JavaScript files. Google renders pages roughly the way a browser does. If robots.txt blocks folders with stylesheets and scripts, Google cannot see the page the way visitors see it, and its assessment of the page can be wrong. Keep resource files accessible.

Rules that are too broad. Disallow: /service also blocks /services/ and /service-van/, because the rule matches the start of the path. End with a slash when you mean a folder.

Unsupported rules. Google ignores crawl-delay and does not read noindex in robots.txt. Lines like that do nothing for Google.

Robots.txt does not remove pages from Google

If you block a page in robots.txt, Google stops reading it, but the address can still remain in the index, or get into it, if other pages link to it. It then appears in the search results without a description.

If you want a page to disappear from Google, use noindex instead, either as a meta tag in the page's head or as the X-Robots-Tag HTTP header. Google has to be able to crawl the page to see the instruction, so the page must not also be blocked in robots.txt. If you need to hide an address quickly, Search Console has a Removals tool, but it only hides it temporarily.

Want to be more visible on Google? Read about SEO →

What is an XML sitemap?

An XML sitemap is a file that lists the addresses on your site in a format search engines can read. It usually sits at yourdomain.com/sitemap.xml. The sitemap helps Google discover pages, especially new pages and pages with few internal links.

It does not replace good internal linking. A page that only exists in the sitemap but is not linked from any other page on the site will struggle to be prioritised.

What belongs in the sitemap

Only include pages you want to appear on Google. Every address in the file should meet three requirements.

  • Canonical: the address should be the version you want indexed, not a variant with parameters or a duplicate.
  • Indexable: the page must not have noindex and must not be blocked in robots.txt.
  • Returns 200: no redirects, no 404 pages and no server errors.

So remove thank-you pages, internal search results, filter pages and old addresses that redirect. Use full addresses with the same protocol and the same www variant that the site uses.

lastmod

Each address can have a lastmod value stating when the page was last changed. Google uses the value if it proves accurate over time. So only set it when the content has actually changed, and not to today's date on every page at every build. Google ignores the changefreq and priority fields.

Sitemap index for larger sites

A single sitemap may contain at most 50,000 addresses and be at most 50 MB uncompressed. If you have more pages, split them across several files and gather them in a sitemap index, a file that lists your sitemaps. This can be practical on smaller sites too, for example with one file for service pages and one for blog posts. You can then see in Search Console how indexing is going for each part.

Most publishing systems create the sitemap automatically. On sites built in Next.js, both the sitemap and robots.txt are often generated directly in the code, so new pages are included automatically when they are published.

Submit the sitemap in Search Console

Point to the sitemap in two places.

  1. Add a Sitemap line to robots.txt, so all search engines find it.
  2. Submit it in Google Search Console under the Sitemaps report.

Search Console then shows whether the file could be read, when it was last fetched and how many addresses were discovered. You do not need to resubmit it every time it changes. Google fetches it again at regular intervals.

In the page indexing report you can filter by your sitemap. You then see how many of the addresses you yourself have listed are actually indexed, and why the rest are not.

How to request indexing of a page

If you have published a new page or made an important change, you can ask Google to crawl it.

  1. Open Search Console and paste the address into the URL Inspection search field at the top.
  2. Read the result. It says whether the page is already on Google and whether anything is preventing indexing.
  3. Click the live URL test if you have fixed an error, to see that Google sees the new version.
  4. Click Request indexing.

The request puts the page in the queue for crawling. There is no promise of when, or whether, it will be indexed. There is also a daily limit on the number of requests, so save the button for pages that matter. For larger changes, such as a new site or many new pages, an updated sitemap and good internal linking are what work.

If the page still does not get indexed, the problem is rarely robots.txt or the sitemap. More often the content is too thin, the page is too similar to another page or it has no links from the rest of the site. Read more about the bigger picture in our guide to technical SEO and in what SEO is.

Quick check of your site

  • Open yourdomain.com/robots.txt and look for Disallow: / under User-agent: *.
  • Check that robots.txt has a Sitemap line pointing to the right address.
  • Open the sitemap and spot-check a few addresses. They should load without redirecting.
  • Compare the number of addresses in the sitemap with the number of indexed pages in Search Console.
  • Go through pages with noindex and check that it is intentional.

Want help?

Solva goes through robots.txt, sitemaps and indexing as part of our SEO work. If you first want to see what your site looks like today, you can order a free SEO audit.

Frequently asked questions

Where is robots.txt located?

Robots.txt is always in the root of the domain, for example yourdomain.com/robots.txt. The file has to be exactly there for search engines to find it, and each subdomain needs its own file.

Can I remove a page from Google with robots.txt?

No. Robots.txt stops crawling but does not remove the address from the index. Use a noindex tag instead and make sure the page is not also blocked in robots.txt, so that Google can read the instruction.

Which pages should be in an XML sitemap?

Only pages you want to appear on Google: canonical addresses that can be indexed and that return status code 200. Redirects, 404 pages, noindex pages and duplicates should not be included.

How do I get Google to index my page?

Make sure the page can be crawled, does not have noindex, is in your sitemap and is linked from other pages on the site. Then paste the address into URL Inspection in Google Search Console and click Request indexing.

Does a small site need a sitemap?

A small site with good internal linking usually gets found anyway, but a sitemap does no harm and makes it easier to follow indexing in Search Console. Most publishing systems create it automatically.

Sources

Written by Solva. We are a web agency in Uppsala that builds websites and handles SEO and ads for businesses across Sweden. More about us

Want to be more visible on Google?

We go through your site and show you what is holding it back. The review is free.