Page indexing report · Search Console
Blocked by robots.txt
A rule in your robots.txt file stops Google from crawling the URL. Google’s Help page calls this “URL blocked by robots.txt”.1 For carts, accounts and search pages, that’s the point. For pages you want in search, it’s a rule to find and change, and the tester on this page shows you which one.
Check the URL first
The checker fetches your live robots.txt and tells you whether it blocks Googlebot from this URL. Google may still be using a copy from up to a day ago, and the report shows its last check.
In 30 seconds
What the status means
Before Googlebot fetches a URL, it checks the site’s robots.txt. If a rule disallows the URL, Google doesn’t request it at all. The report lists it here and notes that this doesn’t guarantee the page stays out of the index: if Google finds information about it elsewhere, there’s a small chance it gets indexed anyway.1
This is the one status in this group that’s set at the crawl stage, before Google has seen anything on the page. Select a stage to see what happens there.
Discover
Google learns that a URL exists, mostly from links on pages it already knows and from XML sitemaps.
What goes wrong: Nothing links to the page and it isn’t in a sitemap, so Google never hears about it.
Report statuses at this stage
- URL is unknown to Google (shown in URL Inspection, not in the report)
Crawl
Googlebot checks robots.txt, then requests the URL. It spaces requests out so it doesn’t overload the server.
What goes wrong: The URL is blocked, returns an error, or the server is slow, so Google backs off and crawls less.
Report statuses at this stage
Render
Google runs the page’s JavaScript in a recent version of Chrome to see the finished page.
What goes wrong: Content only appears after a click or scroll, or the scripts fail, so the rendered page is nearly empty.
Report statuses at this stage
- No status of its own. Problems show up at the next stage as thin pages or soft 404s.
Index
Google decides which version of the page to keep and whether it is worth storing at all.
- Group duplicates. Pages with the same or very similar content go into one group.
- Pick a canonical. One URL represents each group. Canonical tags, redirects and sitemaps are hints, not commands.
- Select for the index. Google decides whether the canonical page is worth keeping. Google says this largely depends on quality.
What goes wrong: The page duplicates another, is marked noindex, or Google decides it adds too little to keep.
Report statuses at this stage
Serve
Indexed pages can appear in search results. Being indexed is not the same as ranking; an indexed page can still get no clicks.
What goes wrong: The page is indexed but doesn’t match what people search for, or other pages answer it better.
Report statuses at this stage
Google’s John Mueller has described robots.txt as close to absolute: if Google can parse the file, it follows the rules.9 So this status is never Google being unreliable. A rule matched.
Is it a problem?
Usually fine
- Cart, checkout, account and order pages
- Internal search results
- Sorting, filter and session parameter URLs
- Admin areas and API endpoints
- Platform defaults, such as Shopify’s
Worth fixing
- Pages you want in search: products, articles, landing pages
- The whole site, after a launch or migration
- CSS, JavaScript or API files that your pages need to render
- Pages you blocked to get them out of Google
- Blocked URLs listed in your sitemap
Blocked page resources matter more than they look. If Google can’t load scripts or styles a page needs, it can get a blank or broken page, which can lead to a soft 404.7
Find the cause
Answer a few questions. Each cause is explained below.
A deliberate block
robots.txt exists to manage which URLs crawlers request, mainly to avoid overloading the site.2 Blocking carts, accounts and infinite filter combinations is normal. Shopify’s default file blocks admin, cart, account, order and sorted collection pages for exactly that reason.11
A staging file or site-wide setting went live
A Disallow: / meant for a test site is the classic launch-day mistake. Older WordPress versions added it when “Discourage search engines” was ticked. Since version 5.3, that setting uses a noindex tag instead.12
A rule that matches more than you meant
Rules match from the start of the path, paths are case-sensitive, and * and $ are wildcards. When rules conflict, the longest matching path wins.3 So Disallow: /p blocks /products, and Disallow: /Blog/ doesn’t block /blog/.
Blocking pages to remove them
Google says robots.txt is not a mechanism for keeping a page out of Google. A blocked URL can still appear, without a description.2 And if the page carries a noindex, Google never sees it, because it can’t crawl the page.5 A Noindex: line in robots.txt doesn’t help: Mueller has called it an unsupported directive that does nothing.10
robots.txt is unreachable
If the file returns a server error, Google pauses crawling for 12 hours and then relies on its cached copy for up to 30 days.3 Firewalls and bot protection that block Googlebot from the file can do the same. Check the robots.txt report in Search Console for fetch errors.4
Test your robots.txt
Paste your robots.txt and a URL path to see which rule applies, using Google’s matching rules. Copy the file from https://yoursite.com/robots.txt.
Blocked
Matched Disallow: /*?sort= in the User-agent: * group. It’s the longest rule that matches.
Runs in your browser with Google’s matching rules: the most specific user-agent group applies, the longest matching path wins, and Allow wins a tie.
How Google reads robots.txt
The rules most often behind a surprise block. Pick one for details.
Caching
Google generally caches robots.txt for up to 24 hours, longer if it can’t fetch a fresh copy. A Cache-Control: max-age header can change that.3 To speed things up after a fix, use Request a recrawl in Search Console’s robots.txt report.4
4xx responses
If robots.txt returns a 4xx error other than 429, Google crawls as if there were no robots.txt at all.3 A missing file means everything is allowed.
5xx responses
A server error pauses crawling for 12 hours. For the next 30 days Google uses the last cached copy while it retries. After that, if the rest of the site is reachable, it acts as if there were no robots.txt; if the site is generally unavailable, it stops crawling.3
File size
Google reads the first 500 KiB and ignores anything after it.3
Which rule wins
Google applies the most specific rule, measured by the length of the path. If an Allow and a Disallow match equally, the less restrictive one wins.3
Disallow: /shop/
Allow: /shop/sale/ # longer, so /shop/sale/ is crawlableCase sensitivity
Field names like user-agent and disallow are case-insensitive. Paths are case-sensitive.3
Disallow: /Private/ # does not block /private/Wildcards
* matches any sequence of characters, including none. $ marks the end of the URL. Rules match from the start of the path.3
Disallow: /*?sort= # any URL with ?sort=
Disallow: /*.pdf$ # URLs ending in .pdfRedirects
Google follows at least five redirect hops when fetching robots.txt. Beyond that, it treats the file as a 404.3
How to fix it
- Decide whether each blocked URL should be crawled. Most sites have some that shouldn’t.
- For the ones that should, find the matching rule with the tester above. Make sure you’re reading the robots.txt for the right host and protocol.
- Remove or narrow the rule, or add a longer Allow rule for the path you want crawled.3
- In Search Console, open the robots.txt report and request a recrawl of the file.4
- Request indexing for your most important URLs, and validate the fix in the report.
- For pages you blocked to remove them: unblock them and add noindex instead, or return 404 or 410.6
What other guides get wrong
Several popular guides, including some from 2025, still give this advice:
| Often said | What’s actually true |
|---|---|
| Check the rule with the robots.txt Tester in Search Console | That tool has been replaced by the robots.txt report, which shows the files Google found, when it crawled them, any errors, and lets you request a recrawl.4 Google’s Help page for this status still mentions the tester. |
How long it takes
- Up to 24 hoursfor Google to use a changed robots.txt.3 You can ask for a sooner recrawl in the robots.txt report.4
- Up to about 2 weeksfor Search Console to validate a fix, sometimes longer.1
- About 6 monthsfor a temporary removal in the Removals tool, if you need a URL hidden while a noindex takes effect.8
Google gives no timeline for recrawling and indexing the unblocked pages themselves.
Blocked vs Indexed, though blocked by robots.txt
| Blocked by robots.txt | Indexed, though blocked by robots.txt | |
|---|---|---|
| Crawled? | No | No |
| Indexed? | No | Yes, from links to it, without reading the page1 |
| Report section | Not indexed | Indexed, with a warning |
| If you want it out of Google | Unblock it and add noindex | Unblock it and add noindex1 |
Questions
What does “Blocked by robots.txt” mean in Search Console?
A rule in your robots.txt file stops Googlebot from crawling the URL, so Google can’t read the page. Google’s Help page calls it “URL blocked by robots.txt”. It’s fine for pages you don’t want crawled.
What is the difference between Blocked by robots.txt and Indexed, though blocked by robots.txt?
Both mean robots.txt stops Google crawling the URL. In the first, the URL isn’t indexed. In the second, Google indexed the URL anyway, from links pointing to it, without reading the page.
How do I unblock a page in robots.txt?
Find the Disallow rule that matches the URL and remove or narrow it, or add a longer Allow rule. Then request a recrawl of robots.txt in Search Console’s robots.txt report and request indexing for the page.
Should I use robots.txt to remove pages from Google?
No. Google says robots.txt doesn’t keep pages out of Google, and a blocked page can still be indexed from links. Use a noindex tag, a login or a 404 or 410 instead, and leave the page crawlable so Google sees it.
Where is the robots.txt tester in Search Console?
The old robots.txt Tester was replaced by the robots.txt report. It shows the robots.txt files Google found, when it last crawled them, and any errors, and lets you request a recrawl. To test a URL against your rules, use the tester on this page.
Why is my Shopify store blocked by robots.txt?
Shopify’s default robots.txt blocks admin, cart, account, order and sorted collection pages, which shouldn’t be in search anyway. If a page you want indexed is blocked, you can edit the robots.txt.liquid template, though Shopify Support doesn’t help with those edits.
Sources
- Google, “Page indexing report”, Search Console Help
- Google, “Introduction to robots.txt”, Search Central, updated December 2025
- Google, “How Google interprets the robots.txt specification”, Search Central, updated August 2026
- Google, “robots.txt report”, Search Console Help
- Google, “Robots meta tag, data-nosnippet, and X-Robots-Tag specifications”, Search Central, updated March 2026
- Google, “Block Search indexing with noindex”, Search Central, updated December 2025
- Google, “Troubleshoot Google Search crawling errors”, Search Central, updated December 2025
- Google, “Removals and SafeSearch reports tool”, Search Console Help
- John Mueller (Google), LinkedIn, September 2024, reported by PPC Land
- John Mueller (Google), LinkedIn, March 2024, reported by Search Engine Journal
- Shopify, “Editing robots.txt.liquid”, Shopify Help Center
- Peter Wilson, “Changes to prevent search engines indexing sites”, Make WordPress Core, September 2019