Page indexing report · Search Console
Blocked due to unauthorized request (401)
Google asked for the page and your server replied that it needs a login. Googlebot doesn’t log in, so the page isn’t indexed.1 That’s correct for private pages and a problem for public ones. Here’s how to tell which, and how to fix it.
Check the URL first
The checker shows whether the URL asks an ordinary visitor for a login. It can’t see what Google’s own requests get. A rule that targets Google’s IP addresses or user agent only shows up in URL Inspection and your server logs.
In 30 seconds
- The server answered Google with a 401, a request for credentials. Google doesn’t index 4xx URLs and removes ones that were indexed.2
- On staging sites, account pages and private areas it’s the right outcome. Just keep those URLs out of sitemaps and public links.
- On public pages, remove the login, or for gated content let verified Google crawlers in.1
- Verify Google by reverse and forward DNS or its published IP ranges, never by user agent alone.4
What the status means
A 401 response means “authenticate first”. Google’s report describes it as the page being blocked to Googlebot by a request for authorization, and suggests checking it by visiting the page in incognito mode.1
Google handles 401 like most other 4xx codes. It ignores any content in the response, doesn’t index the URL, and removes it from the index if it was there. It then crawls the URL less and less often.2 A 401 doesn’t slow crawling of the rest of the site.2
Where it happens
The 401 stops Google at the crawl stage, before it sees any content. Select a stage to see what happens there.
Discover
Google learns that a URL exists, mostly from links on pages it already knows and from XML sitemaps.
What goes wrong: Nothing links to the page and it isn’t in a sitemap, so Google never hears about it.
Report statuses at this stage
- URL is unknown to Google (shown in URL Inspection, not in the report)
Crawl
Googlebot checks robots.txt, then requests the URL. It spaces requests out so it doesn’t overload the server.
What goes wrong: The URL is blocked, returns an error, or the server is slow, so Google backs off and crawls less.
Report statuses at this stage
- Discovered – currently not indexed
- Blocked by robots.txt
- Server error (5xx)
- Not found (404)
- Redirect error
- Blocked due to unauthorized request (401) (this page)
- Blocked due to access forbidden (403)
- Blocked due to other 4xx issue
Render
Google runs the page’s JavaScript in a recent version of Chrome to see the finished page.
What goes wrong: Content only appears after a click or scroll, or the scripts fail, so the rendered page is nearly empty.
Report statuses at this stage
- No status of its own. Problems show up at the next stage as thin pages or soft 404s.
Index
Google decides which version of the page to keep and whether it is worth storing at all.
- Group duplicates. Pages with the same or very similar content go into one group.
- Pick a canonical. One URL represents each group. Canonical tags, redirects and sitemaps are hints, not commands.
- Select for the index. Google decides whether the canonical page is worth keeping. Google says this largely depends on quality.
What goes wrong: The page duplicates another, is marked noindex, or Google decides it adds too little to keep.
Report statuses at this stage
Serve
Indexed pages can appear in search results. Being indexed is not the same as ranking; an indexed page can still get no clicks.
What goes wrong: The page is indexed but doesn’t match what people search for, or other pages answer it better.
Report statuses at this stage
Is it a problem?
The 401 comes from your own server, so it’s yours to change. Google notes that in general you can only fix issues whose Source column says Website.1
Usually fine
- Staging, preview and development sites
- Account, checkout and admin pages
- Private documents and intranet pages
- Old URLs nobody links to any more
Worth fixing
- Public pages you want in search
- URLs listed in your XML sitemap
- A sudden jump in the count after a launch or deploy
- Members-only content you want indexed
Google itself says confidential content should be password protected.7 A login blocks crawlers and people alike.6
Find the cause
Answer a few questions. Each cause is explained below.
A private URL leaked into public view
Google finds URLs through links and sitemaps. If a staging host or a private folder shows up here, something public points at it. Mueller’s advice on staging sites starts with the same point: don’t link to them.8 Keep the login and remove the references.
A login was left on the live site
Sites launched from a password-protected staging setup sometimes keep the password on some folders. Plugins that protect members’ content can also cover more URLs than intended. An incognito window shows it straight away.1
Only Google gets the 401
If the page loads for you but URL Inspection’s live test still reports a 401, a rule is treating Google’s requests differently. The live test shows the failure in the Page fetch field.3 Confirm it in your access logs by finding requests from verified Google IPs and the status they got.
Gated content you want indexed
Google’s report says to either remove the login or let Googlebot through after verifying its identity.1 Serving the full page to Google while visitors see a paywall isn’t treated as cloaking, as long as Google sees what a subscriber sees.11 The paywall markup makes that explicit.10
Where 401s come from
Pick a source to see how to recognise it. The last tab shows how to check a request really came from Google.
HTTP Basic or Digest auth
The browser shows its own grey username and password box, not a page from your site. The response is a 401 with a WWW-Authenticate header. It’s set in server config (Apache .htaccess, nginx auth_basic) or a hosting panel’s “password protect directory” option.
curl -sI https://example.com/page | grep -iE '^HTTP|www-authenticate'
# HTTP/2 401
# www-authenticate: Basic realm="Restricted"Application login wall
A login form inside your site’s design. Only some apps answer with a 401. Many redirect to a login page instead, or show the form with a 200. Those won’t land in this row: a redirect is reported as a redirect, and a login form served with 200 can be crawled like any other page.
If the report says 401, look for the code path that sends it: API routes, middleware, or a framework’s auth guard returning 401 for anyone without a session.
Staging and preview sites
Protecting staging with a login is what Google’s John Mueller recommends: he called server-side authentication “generally the best approach for staging servers”.8 Some hosting platforms can also put preview deployments behind a login.
When staging URLs show up in this report, the login is fine. Something on the public site points at them: an absolute link, a canonical tag, hreflang, or a sitemap generated on the wrong host.
Members-only or paid content
If you want gated pages in search, Google’s rule is that Googlebot must be able to access the page.10 Google doesn’t treat a paywall as cloaking if it can see the full content behind it, as anyone with access would.11
Mark the gated part with isAccessibleForFree set to false and a hasPart that points at it with a cssSelector.10
robots.txt behind the login
If the whole host is behind a login, /robots.txt returns a 401 too. Google treats a robots.txt served with a 4xx status as if it didn’t exist.9 Any disallow rules in it aren’t read, so Google may try every URL it knows and get a 401 on each.
That isn’t harmful on a staging host. It explains why the row can be large.
Is it really Googlebot?
Anyone can send a Googlebot user agent. Google’s documented checks use DNS or its published IP ranges.4 The reverse lookup must end in googlebot.com, google.com or googleusercontent.com, and the forward lookup must return the original IP.4
# 1. Reverse DNS: the IP from your access log
host 66.249.66.1
# 1.66.249.66.in-addr.arpa domain name pointer crawl-66-249-66-1.googlebot.com.
# 2. Forward DNS: the name must point back to the same IP
host crawl-66-249-66-1.googlebot.com
# crawl-66-249-66-1.googlebot.com has address 66.249.66.1For many requests, match IPs against Google’s JSON files instead:4
# Googlebot's IP ranges (CIDR), from Google's published file
curl -s https://developers.google.com/static/crawling/ipranges/common-crawlers.json \
| jq -r '.prefixes[] | .ipv4Prefix // .ipv6Prefix'URL Inspection’s live test uses a different user agent, Google-InspectionTool.5 A rule that matches the word “Googlebot” can treat the two differently.
How to fix it
- Open the row in the Page indexing report and switch the filter to “All submitted pages”. Those are the URLs you asked Google to index.1
- Inspect one URL and run the live test. Check the Page fetch result and the last crawl date.3
- Open the URL in a private window to see what a logged-out visitor gets.1
- Private URL: remove it from sitemaps and public links, and leave the login in place.
- Public URL: remove the login rule. Gated URL: let verified Google crawlers in and add paywall markup.10
- Run the live test again. If Page fetch succeeds, click Request indexing for key URLs, then Validate fix in the report.3
How long it takes
- A day or sofor a URL you request, by Google’s estimate, though it can take much longer.3
- Up to about 2 weeksfor Search Console to validate a fix, sometimes much longer.1
- No set timefor private URLs to drop out of the report. Google crawls 4xx URLs gradually less often.2
401 vs 403
| Blocked due to unauthorized request (401) | Blocked due to access forbidden (403) | |
|---|---|---|
| What the server says | Log in first | You’re not allowed |
| Google’s view | Blocked by a request for authorization1 | Googlebot never sends credentials, so the server is returning it incorrectly1 |
| Usual cause | Password protection, staging, members areas | Firewall, CDN or security plugin rules |
| Effect on indexing | The same: not indexed, and removed if indexed before2 | |
Questions
What does “Blocked due to unauthorized request (401)” mean?
Google requested the URL and the server answered 401, asking for credentials. Googlebot doesn’t log in, so it can’t see the page and won’t index it. Pages that were indexed and start returning 401 are removed.
Is a 401 in Search Console bad?
Not if the page is meant to be private. Logins are the right way to keep staging sites and account pages out of Google. It matters when a public page or a sitemap URL is in the row.
How do I let Googlebot past a login?
Remove the login for pages that should be public. For gated content you want indexed, serve the full page only to requests verified as Google by reverse and forward DNS or Google’s published IP ranges, and add paywall structured data. Never trust the user agent alone.
Why is my staging site showing 401 errors in Search Console?
Google found staging URLs somewhere public, often a sitemap, canonical tag or absolute link that points at the staging host. The login is working. Remove the references and Google will crawl those URLs less over time.
Should I use robots.txt instead of a password on staging?
No. robots.txt stops crawling, not indexing, and it’s easy to copy to the live site by mistake. Google’s John Mueller has recommended server-side authentication for staging servers.
What is the difference between 401 and 403 in Search Console?
401 means the server asked for credentials. 403 means it refused the request outright. Google notes that Googlebot never sends credentials, so a 403 is the server’s choice and often comes from a firewall or CDN rule.
Sources
- Google, “Page indexing report”, Search Console Help
- Google, “How HTTP status codes affect Google’s crawlers”, Google Crawling Infrastructure, updated February 2026
- Google, “URL Inspection tool”, Search Console Help
- Google, “Verify requests from Google crawlers and fetchers”, Google Crawling Infrastructure, updated March 2026
- Google, “Google’s common crawlers”, Google Crawling Infrastructure, updated July 2026
- Google, “What is Googlebot”, Search Central, updated February 2026
- Google, “Control the content you share on Search”, Search Central, updated December 2025
- John Mueller (Google), Webmaster Central hangout, reported by Search Engine Journal, September 2019
- Gary Illyes (Google), “Don’t use 403s or 404s for rate limiting”, Search Central Blog, February 2023
- Google, “Subscription and paywalled content markup”, Search Central, updated September 2026
- Google, “Spam policies for Google web search”, Search Central, updated August 2026