Page indexing report · Search Console
Indexed, though blocked by robots.txt
Your robots.txt stops Google from reading the page, but Google indexed the URL anyway because other pages link to it.1 It’s a warning, and often harmless. Here’s when it matters, and the right way to get a page out of Google, which isn’t robots.txt.
Check the URL first
The checker tests the URL against your live robots.txt and looks for a noindex. It can’t tell whether Google still has the URL indexed. URL Inspection in Search Console can.
In 30 seconds
- robots.txt controls crawling, not indexing. Google calls it “not a mechanism for keeping a web page out of Google”.2
- The URL can appear in results, but without a description.2
- To remove it, unblock the page and add noindex. A blocked page’s noindex is never seen.3
- For cart, search and filter URLs, it’s usually safe to leave.6
What it means
Google’s explanation: it “always respects robots.txt”, but that “doesn’t necessarily prevent indexing if someone else links to your page”.1 Googlebot never fetched the page. It indexed the address from links pointing at it.
The block applies at the crawl stage, but this status is about what Google shows in results.
Discover
Google learns that a URL exists, mostly from links on pages it already knows and from XML sitemaps.
What goes wrong: Nothing links to the page and it isn’t in a sitemap, so Google never hears about it.
Report statuses at this stage
- URL is unknown to Google (shown in URL Inspection, not in the report)
Crawl
Googlebot checks robots.txt, then requests the URL. It spaces requests out so it doesn’t overload the server.
What goes wrong: The URL is blocked, returns an error, or the server is slow, so Google backs off and crawls less.
Report statuses at this stage
Render
Google runs the page’s JavaScript in a recent version of Chrome to see the finished page.
What goes wrong: Content only appears after a click or scroll, or the scripts fail, so the rendered page is nearly empty.
Report statuses at this stage
- No status of its own. Problems show up at the next stage as thin pages or soft 404s.
Index
Google decides which version of the page to keep and whether it is worth storing at all.
- Group duplicates. Pages with the same or very similar content go into one group.
- Pick a canonical. One URL represents each group. Canonical tags, redirects and sitemaps are hints, not commands.
- Select for the index. Google decides whether the canonical page is worth keeping. Google says this largely depends on quality.
What goes wrong: The page duplicates another, is marked noindex, or Google decides it adds too little to keep.
Report statuses at this stage
Serve
Indexed pages can appear in search results. Being indexed is not the same as ranking; an indexed page can still get no clicks.
What goes wrong: The page is indexed but doesn’t match what people search for, or other pages answer it better.
Report statuses at this stage
- Page is indexed
- Indexed, though blocked by robots.txt (this page)
- Page indexed without content
Because Google can’t read the page, the result has no description.2 Google explains the missing snippet the same way: a rule such as robots.txt stopped Googlebot from reading the content.4 John Mueller has said Google then has to rely mainly on links to judge whether the URL is worth showing, and that it will prefer pages it can crawl.5 He has also said the average user won’t see such URLs.6
Is it a problem?
Usually fine
- Cart, checkout and add-to-cart URLs
- Internal search results and filter or sort parameters
- URLs that only show up in a site: search
Worth fixing
- A page you want to rank, blocked by mistake
- Private or sensitive URLs that must not be in Google
- Pages with a noindex that robots.txt hides
- Blocked URLs listed in your XML sitemap
John Mueller in 2026, on add-to-cart URLs: “Blocking them with robots.txt is fine.”7
Find the cause
Answer a few questions to pick the right fix. Each cause is explained below.
Links to blocked URLs
The underlying cause. Internal links, other sites or your own sitemap point at a blocked URL, so Google knows it exists.1 Cart buttons, search forms and filter links create these URLs at scale.
A rule that’s broader than intended
robots.txt rules match from the start of the path, so a short rule can block far more than you meant. The tester below shows which rule matches.
noindex hidden by robots.txt
Adding noindex and a robots.txt block together, to be safe, backfires. Google says that if the page is blocked, the crawler “will never see the noindex rule”, and the page can still appear in results.3
Platform defaults
Shopify’s default robots.txt blocks cart, checkout, account, and filtered or sorted collection URLs.11 Those landing here is expected.
Find the rule and pick the fix
Paste your robots.txt and a URL from the report to see which rule blocks it.
Blocked
Matched Disallow: /*?sort= in the User-agent: * group. It’s the longest rule that matches.
Runs in your browser with Google’s matching rules: the most specific user-agent group applies, the longest matching path wins, and Allow wins a tie.
Then pick what you want to happen to the URL:
It should rank
- Find the rule that blocks the URL with the tester above.
- Remove it, narrow it, or add an
Allowfor the path. - In Search Console’s robots.txt report, request a recrawl of the file.9 Google otherwise caches it for up to 24 hours.8
- Inspect the URL and click Request indexing.
# Before: blocks /products and /pricing too
User-agent: *
Disallow: /p
# After: blocks only /p/...
User-agent: *
Disallow: /p/It should be removed
- Add a noindex to the page, or return 404 or 410 if the page is gone.
- Remove the robots.txt rule. For noindex to work, the page “must not be blocked by a robots.txt file”.3
- Let Google recrawl the URL. It drops the page once it sees the noindex.
- Keep the page crawlable. If you block it again, Google can no longer see the noindex.
For an HTML page:
<meta name="robots" content="noindex">For a PDF or other file, send an HTTP header instead:3
# Apache
<Files ~ "\.pdf$">
Header set X-Robots-Tag "noindex"
</Files>It must go now
It’s harmless junk
Cart, search and filter URLs blocked by robots.txt are fine to leave. John Mueller has said the average user won’t see them and “I wouldn’t fuss over it”.6 If you’d rather they weren’t indexed at all, use noindex without a robots.txt block; Mueller said that’s fine too, though the URLs will then be crawled.6
How to fix it
- Export the URLs and sort them: should rank, should be removed, or harmless.
- Remove any of them from your XML sitemap that you don’t want indexed.
- Should rank: remove or narrow the robots.txt rule, request a robots.txt recrawl, then request indexing.
- Should be removed: add noindex (or delete the page), then remove the robots.txt rule so Google can see it. Google says robots.txt “is not the correct mechanism” for this.1
- Harmless: leave them.
Myths
Some top-ranking guides say these. They’re wrong.
| Often said | What’s actually true |
|---|---|
| Blocking a page in robots.txt removes it from Google | robots.txt isn’t a way to keep a page out of Google; blocked URLs can still appear in results.2 |
| Shopify doesn’t let you edit robots.txt | You can, by adding a robots.txt.liquid template to your theme. Shopify calls it an unsupported customisation.11 |
How long it takes
- Up to 24 hoursfor Google to pick up a robots.txt change, as it usually caches the file that long.8
- About 6 monthsis how long a temporary removal hides a URL.10
- No fixed timefor Google to recrawl an unblocked URL and see its noindex. Google gives none.
- Up to about 2 weeksfor Search Console to validate a fix, sometimes much longer.1
Indexed, though blocked vs Blocked by robots.txt
| Indexed, though blocked by robots.txt | Blocked by robots.txt | |
|---|---|---|
| Can Google crawl it? | No | No |
| Is it indexed? | Yes, from links alone1 | No |
| Report section | Indexed, flagged as a warning | Why pages aren’t indexed |
| To get it indexed | Remove the robots.txt rule | Remove the robots.txt rule |
| To keep it out | Unblock and add noindex | Nothing, unless it starts getting indexed |
Questions
Is Indexed, though blocked by robots.txt bad?
Usually not. It’s a warning, not an error. For cart, search and filter URLs that nobody searches for, Google’s John Mueller has said not to worry about it. It matters when the URL is a page you want to rank, or one you need out of Google.
How do I remove a page that is indexed but blocked by robots.txt?
Remove the robots.txt block and add a noindex tag, or delete the page so it returns 404. Google has to crawl the page to see either. For urgent cases, the Removals tool hides it for about six months while you do that.
Why did Google index a page that robots.txt blocks?
robots.txt controls crawling, not indexing. If other pages link to a blocked URL, Google can index the URL from those links alone, without reading the page.
Can I use noindex and robots.txt together?
Not on the same URL. If robots.txt blocks the page, Google never sees the noindex, so the URL can stay indexed. Use one or the other: noindex to keep a page out of the index, robots.txt to stop crawling.
How do I fix this on Shopify?
Shopify’s default robots.txt blocks cart, checkout, account and some filtered collection URLs, and those are usually fine to leave. To change the rules, add a robots.txt.liquid template to your theme. Shopify Support doesn’t help with edits, so change only what you need.
What is the difference from Blocked by robots.txt?
Both mean robots.txt stops Google from crawling the URL. Blocked by robots.txt means the URL isn’t indexed. Indexed, though blocked means Google indexed it anyway, from links, and shows it without a description.
Sources
- Google, “Page indexing report”, Search Console Help
- Google, “Introduction to robots.txt”, Search Central, updated December 2025
- Google, “Block Search indexing with noindex”, Search Central, updated December 2025
- Google, “No page information in search results”, Search Console Help
- John Mueller (Google), Webmaster Central hangout, July 2019, reported by Search Engine Journal
- John Mueller (Google), LinkedIn, September 2024, reported by Search Engine Journal
- John Mueller (Google), Reddit, June 2026, reported by Search Engine Journal
- Google, “How Google interprets the robots.txt specification”, Search Central, updated August 2026
- Google, “robots.txt report”, Search Console Help
- Google, “Removals and SafeSearch reports tool”, Search Console Help
- Shopify, “Editing robots.txt.liquid”, Shopify Help Center