Page indexing report · Search Console
Blocked due to access forbidden (403)
Google requested the page and something on your side refused it. Googlebot never sends credentials, so Google says the server is returning this error incorrectly, and the page won’t be indexed.1 Here’s how to find which layer sent the 403 and let Google back in.
Check the URL first
The checker shows whether the URL forbids an ordinary request. A pass doesn’t clear it: firewalls and CDNs often block by Google’s IP addresses, which no outside test can reproduce. URL Inspection and your logs can.
In 30 seconds
- Google doesn’t index URLs that return 403, and removes ones that were indexed.2
- The usual cause is a CDN, firewall or security rule that blocks Google’s IPs, often added automatically.3
- Since September 2026, Cloudflare’s Block setting for AI training also blocks Googlebot.12
- Never use 403 to slow crawling. Use 429 or 503, briefly.4
What the status means
HTTP 403 means the request was understood and refused. Google’s report points out that Googlebot never provides credentials, so a 403 can’t mean “wrong password”. It means a rule decided to turn Google away. The fix Google suggests is to admit visitors who aren’t signed in, or to allow Googlebot after verifying its identity.1
Google treats 403 like most other 4xx codes: content in the response is ignored, the URL isn’t indexed, and if it was indexed it’s removed. Crawling of the URL slows gradually.2
Where it happens
The 403 stops Google at the crawl stage. Select a stage to see what happens there.
Discover
Google learns that a URL exists, mostly from links on pages it already knows and from XML sitemaps.
What goes wrong: Nothing links to the page and it isn’t in a sitemap, so Google never hears about it.
Report statuses at this stage
- URL is unknown to Google (shown in URL Inspection, not in the report)
Crawl
Googlebot checks robots.txt, then requests the URL. It spaces requests out so it doesn’t overload the server.
What goes wrong: The URL is blocked, returns an error, or the server is slow, so Google backs off and crawls less.
Report statuses at this stage
- Discovered – currently not indexed
- Blocked by robots.txt
- Server error (5xx)
- Not found (404)
- Redirect error
- Blocked due to unauthorized request (401)
- Blocked due to access forbidden (403) (this page)
- Blocked due to other 4xx issue
Render
Google runs the page’s JavaScript in a recent version of Chrome to see the finished page.
What goes wrong: Content only appears after a click or scroll, or the scripts fail, so the rendered page is nearly empty.
Report statuses at this stage
- No status of its own. Problems show up at the next stage as thin pages or soft 404s.
Index
Google decides which version of the page to keep and whether it is worth storing at all.
- Group duplicates. Pages with the same or very similar content go into one group.
- Pick a canonical. One URL represents each group. Canonical tags, redirects and sitemaps are hints, not commands.
- Select for the index. Google decides whether the canonical page is worth keeping. Google says this largely depends on quality.
What goes wrong: The page duplicates another, is marked noindex, or Google decides it adds too little to keep.
Report statuses at this stage
Serve
Indexed pages can appear in search results. Being indexed is not the same as ranking; an indexed page can still get no clicks.
What goes wrong: The page is indexed but doesn’t match what people search for, or other pages answer it better.
Report statuses at this stage
Is it a problem?
It’s a problem more often than a 401, because a 403 on a public page is rarely deliberate. Google notes that in general you can only fix issues whose Source column says Website, and this one comes from your site.1
Usually fine
- Admin, login and account URLs
- Directories you deliberately don’t list
- Private files nobody links to
Worth fixing
- Any public page or sitemap URL
- A sudden jump after a CDN, security or plugin change
- Many unrelated URLs at once, which points to a site-wide rule
- robots.txt itself returning 403
The last one matters. Google treats a robots.txt served with a 4xx status as if it didn’t exist, so your disallow rules stop applying.4
Find the cause
Answer a few questions. Each cause is explained below.
Your CDN or firewall is blocking Google
Google’s Martin Splitt and Gary Illyes wrote that flood protection can put wanted crawlers on a CDN’s blocklist, that it can happen automatically, and that it can be hard or impossible to control because it happens on the CDN’s side.3 They recommend checking blocklists every now and then. On Cloudflare, block and challenge rules both produce a 403.10
An AI-crawler setting
Cloudflare classes Googlebot as a mixed-use crawler, one whose crawl serves both search and AI training. It now applies its Block setting for AI training to these crawlers.12 If you turned on AI blocking and the 403s started then, this is the first thing to check.
Rate limiting with the wrong status
A 403 doesn’t tell Google to slow down. Google says 4xx codes other than 429 have no effect on crawl rate.2 If the server needs relief, Google recommends 500, 503 or 429, and only for a couple of hours to a day or two.9
Country blocking
Google says Googlebot’s default IP addresses appear to be in the US, and that it should be treated like any visitor from there.8 A rule that blocks US traffic, or traffic from “unknown” locations, blocks most of Google’s crawling.
The origin server forbids it
When browsers get the 403 too, the origin is the likely source: deny rules, mod_security, or IP blocklists.10 File permissions are the other common culprit, especially for uploads.
Where 403s come from
Pick a source to see how to check for it and what to change. The last tab shows how to prove what Google received.
CDN or firewall bot rules
Google says crawlers you want can “end up in your CDN’s blocklist”, usually in the web application firewall, and that the block may happen automatically.3 Cloudflare lists WAF rules with a block or challenge action, the Security Level setting, DDoS protection and Browser Integrity Check among the features that produce a 403.10
Check: the CDN’s security event log, filtered to Google’s IP ranges.6 Fix: allow verified Google crawlers. Google lists where Cloudflare, Akamai, Fastly, F5 and Google Cloud document this.3
Bot challenge pages
A “checking your browser” page stops Googlebot, which doesn’t pretend to be human. Google strongly recommends answering automated clients with a 503 instead, so the content isn’t removed from the index.3
Check: the live test’s screenshot, if one is available. Google suggests that if it shows a bot challenge, talk to your CDN.3 Fix: exempt verified crawlers from challenges.
AI-crawler blocking
Cloudflare announced that from September 15, 2026, its Block option for AI training applies to mixed-use crawlers, naming Googlebot among them.12 Its new Disallow AI Training option publishes a no-training preference in robots.txt and keeps those crawlers allowed for search.12
Before the change, one site owner reported Googlebot getting 403s on a sitemap with AI Training set to Block. John Mueller asked for details; the cause wasn’t established.13
Check: your AI crawler settings and the date you changed them. Fix: use the Disallow option, or opt out of training through Google-Extended in robots.txt.12
Rate limiting with 403
Gary Illyes wrote that site owners and some CDNs were using 4xx errors to slow Googlebot, and asked them to stop: “please don’t do that”.4 A 403 has no effect on crawl rate and removes the URL from search.2
Check: rate-limit rules and the status they return. Cloudflare’s rate limiting rules return 429 by default, while most of its other security features return 403.11 Fix: answer with 429 or 503.9
Country blocking
Googlebot’s default IP addresses appear to be in the US, though it also crawls from other countries. Google says to treat it like any other visitor from the country it appears to come from.8
Check: geo-blocking rules at the CDN, host and app. Fix: if the content can legally be shown in the US, lift the block there.
Server rules and plugins
A 403 without CDN branding usually comes from the origin server: permission rules such as an Apache .htaccess file, mod_security rules, or IP deny rules.10 WordPress security plugins can add the same kind of rule.
Check: server error logs and the security plugin’s blocked-request log for Google’s IPs. Fix: remove or narrow the rule.
File and folder permissions
Files the web server user can’t read, or folders without an index file and with directory listing off, answer 403 to everyone. This shows up most with uploaded PDFs and images.
Check: open the file in a private window. Fix: correct the permissions, or remove the link if the file shouldn’t be public.
Prove what Googlebot got
A test from your own computer can’t show what Google receives, because many blocks key on Google’s IP addresses. Your access logs can. Verify the IPs with Google’s DNS method or its published IP ranges.6
# Requests claiming to be Googlebot that got a 403
# (combined log format: IP is field 1, status is field 9)
grep 'Googlebot' access.log | awk '$9 == 403 {print $1}' | sort | uniq -c | sort -rn | head
# Then verify each IP really is Google: reverse, then forward DNS
host 66.249.66.1
host crawl-66-249-66-1.googlebot.comThe live test runs as Google-InspectionTool, not Googlebot.7 A rule that matches only the word “Googlebot” can block one and not the other.
How to fix it
- Open the row and note the “First detected” date. Compare it with recent CDN, security and plugin changes.
- Inspect one URL and run the live test. A 403 shows as a failed Page fetch.5
- Find the layer that answered: CDN security events first, then server and plugin logs, filtered to verified Google IPs.6
- Allow verified Google crawlers there, by IP range or verified-bot setting. Don’t allow by user agent alone.
- Replace any 403 used for rate limiting or bot checks with 429 or 503.3
- Run the live test again. When it passes, request indexing for key URLs and click Validate fix.5
- Check the CDN blocklist again in a few weeks. Google warns that IPs can end up there automatically.3
How long it takes
- A day or sofor a URL you request, by Google’s estimate, though it can take much longer.5
- Up to about 2 weeksfor Search Console to validate a fix, sometimes much longer.1
- 1–2 daysat most for returning 429 or 503 to slow Google. Longer may affect how the site appears in Google.9
Google gives no figure for how quickly a 403 removes pages that were indexed. Treat it as urgent.
403 vs 503 or 429
If you have to turn Google away for a while, the status code decides what happens to your pages.
| 403 Forbidden | 503 or 429 | |
|---|---|---|
| Google reads it as | The content doesn’t exist2 | A server problem; 429 counts as one2 |
| Indexed pages | Removed2 | Kept for now, dropped if it persists2 |
| Crawl rate | No effect2 | Google slows down temporarily2 |
| Use it for | Content that should never be public | Overload and bot checks, briefly3 |
Questions
What does “Blocked due to access forbidden (403)” mean?
Google requested the URL and your server, CDN or firewall answered 403 Forbidden. Google doesn’t index URLs that return 403 and removes ones that were indexed. Googlebot never sends credentials, so the 403 is a rule on your side.
Why does my page load in a browser but Google gets a 403?
Most likely a firewall, CDN or security plugin rule matches Google’s IP addresses or user agent. Bot protection can add Google’s IPs to a blocklist automatically. Check your CDN security events and access logs for verified Google requests.
Is Cloudflare blocking Googlebot?
It can. Firewall rules, bot challenges and security settings can all return 403 to crawlers, and since September 15, 2026 Cloudflare’s Block option for AI training also applies to Googlebot. Check Cloudflare’s security events and your AI crawler settings.
How do I allow Googlebot through my firewall?
Allow Google’s published IP ranges or your provider’s verified-bot category. Don’t allow by user agent alone, because anyone can fake it. Google documents a reverse and forward DNS check for single IPs.
Can I use 403 to slow Google down?
No. Google says 403 has no effect on crawl rate and gets URLs removed from search. Return 429 or 503 for a short time instead, or fix the server load.
How long until pages come back after fixing a 403?
Google says a requested URL is typically indexed in a day or so, though it can take longer. Validating the fix in the report typically takes up to about two weeks.
Sources
- Google, “Page indexing report”, Search Console Help
- Google, “How HTTP status codes affect Google’s crawlers”, Google Crawling Infrastructure, updated February 2026
- Martin Splitt and Gary Illyes (Google), “Crawling December: CDNs and crawling”, Search Central Blog, December 2024
- Gary Illyes (Google), “Don’t use 403s or 404s for rate limiting”, Search Central Blog, February 2023
- Google, “URL Inspection tool”, Search Console Help
- Google, “Verify requests from Google crawlers and fetchers”, Google Crawling Infrastructure, updated March 2026
- Google, “Google’s common crawlers”, Google Crawling Infrastructure, updated July 2026
- Google, “How Google crawls locale-adaptive pages”, Search Central, updated December 2025
- Google, “Reduce Google crawl rate”, Google Crawling Infrastructure, updated October 2026
- Cloudflare, “Error 403”, Cloudflare Support docs
- Cloudflare, “Actions reference”, Cloudflare Ruleset Engine docs
- Cloudflare, “Have it both ways: stay discoverable in search while disallowing AI training”, Cloudflare blog, September 2026
- Report on r/TechSEO, with a reply from John Mueller (Google), reported by Search Engine Journal, August 2026