Baidu and Google crawling issues after CDN integration: How do I check status codes and blocking rules?
After a website is put behind a CDN, regular users can access it normally, but Baidu Search Resource Platform or Google Search Console begins reporting crawl errors.
Common symptoms include:
服务器错误(5xx)
网页无法访问
已发现但尚未抓取
抓取频率下降
robots.txt无法获取
网址检查返回403
百度抓取诊断失败
索引页面逐渐减少Many site owners first check the homepage, see that it opens normally in a browser, and assume search engine crawling should be fine as well.
But a normal browser experience does not mean Baidu and Google see the same result.
CDNs, WAFs, and bot protection systems may return different content based on the visitor's IP, region, User-Agent, request frequency, cookies, and JavaScript execution capability. Regular users may get a 200, while search engine crawlers may get a 403, 429, or 503—or even a CAPTCHA page with a 200 status code.
So what really needs to be checked is:
When Baidu and Google send requests, what status code and content does the CDN ultimately return?
Which layers does search engine crawling pass through?
After a site is placed behind a CDN, search engine crawling typically passes through:
Googlebot或Baiduspider
↓
DNS解析与CDN调度
↓
CDN边缘节点
↓
Bot管理、WAF、频控和地区规则
↓
缓存
↓
源站Web服务器
↓
网站程序Any layer can cause crawling to fail.
For example:
DNS directs the crawler to an abnormal node
The CDN blocks access from the crawler's region
The WAF identifies high-frequency crawling as an attack
Bot management requires JavaScript verification
Rate limiting returns 429
The CDN fails to fetch from origin and returns 502 or 504
The cache retains an old 403 page
robots.txt is blocked by a CDN rule
The origin server returns different content based on User-Agent
Therefore, “not indexed after moving to a CDN” does not directly prove that the CDN is affecting SEO. You must first use status codes, response content, and logs to identify the specific failure.
Check the most important URLs first
Do not check only the homepage. Test at least the following addresses:
https://www.example.com/
https://www.example.com/robots.txt
https://www.example.com/sitemap.xml
一篇已经收录的文章
一篇新发布的文章
一个出现抓取异常的具体URL
页面引用的主要CSS和JavaScriptFirst check the status code for a regular request:
curl -sS -D - -o /dev/null \
https://www.example.com/articleThen save the full response:
curl -sS -D response-headers.txt \
-o response-body.html \
https://www.example.com/articleYou need to check more than the status code on the first line. Also check:
Location
Retry-After
Server
Via
Age
X-Cache
X-Robots-Tag
Content-Type
CDN请求ID
WAF事件IDIf the status code is 200, also openresponse-body.htmlto view the content and confirm that it returns the normal article, not:
Access Denied
Checking your browser
Please enable JavaScript
Human Verification
Request blocked
网站维护中A normal status code with an error page as the content can still affect crawling and indexing.
What do different status codes mean for crawling?
Status code | What search engines see | Check first |
|---|---|---|
200 | Content fetched successfully, but it may still be a CAPTCHA or soft 404 | Page body, CAPTCHA, |
301/308 | Permanent redirect | Target address and redirect chain |
302/307 | Temporary redirect | Whether it incorrectly redirects to the homepage, login page, or verification page |
401 | Authentication required | CDN access control, origin authentication |
403 | Access explicitly denied | WAF, geographic blocking, IP rules, bot rules |
404/410 | Page does not exist | URL, cache, routing, and publish status |
429 | Too many requests | CDN rate limiting, bot rate limiting, origin rate limiting |
500 | Origin server internal error | Application and server logs |
502 | CDN received an abnormal upstream response | Origin fetch protocol, port, and connection |
503 | Service temporarily unavailable | Overload, maintenance, no healthy origin server |
504 | CDN timed out waiting for the origin server | Application, database, and origin fetch timeouts |
Google states that sustained 5xx and 429 responses cause its crawling systems to temporarily slow down crawling; URLs that persistently return server errors may eventually drop out of the index.
Network timeouts, connection resets, and DNS errors have no HTTP status code, but Google still treats them as serious availability problems.
A 200 status code can also indicate a crawl error
The most insidious case is not a 403, but when the CDN returns:
HTTP/2 200 OK
Content-Type: text/htmlwhile the page content is:
正在验证您的浏览器……
请完成验证码……
访问过于频繁……For regular users, the browser may execute JavaScript and write a cookie, after which they automatically enter the site; but the search engine may only crawl the verification page.
This kind of response may also be judged as a soft 404 or erroneous content.
Google has specifically warned that when a CDN returns a random error page with a 200 status code, it makes it difficult for search systems to correctly understand the page status.
Therefore, when checking, you must compare:
HTTP status code
Page title
Page body
Content-Type
canonical
robots meta
X-Robots-TagContent differences between regular user and crawler requests
For example, search for characteristics in the response content:
curl -sS https://www.example.com/article |
grep -Ei 'captcha|access denied|checking your browser|noindex'If the page should be an article but the returned content is only a few KB of verification code, you cannot consider crawling normal just because the status code is 200.
Perform an initial comparison using User-Agent
You can simulate the User-Agents for Googlebot and Baiduspider to see whether the CDN returns different results.
Test a regular request:
curl -sS -D normal-headers.txt \
-o normal-body.html \
https://www.example.com/articleSimulate Googlebot:
curl -sS -D googlebot-headers.txt \
-o googlebot-body.html \
-A 'Mozilla/5.0 (compatible; Googlebot/2.1; +http://www.google.com/bot.html)' \
https://www.example.com/articleSimulate Baiduspider:
curl -sS -D baiduspider-headers.txt \
-o baiduspider-body.html \
-A 'Mozilla/5.0 (compatible; Baiduspider/2.0; +http://www.baidu.com/search/spider.html)' \
https://www.example.com/articleCompare response headers:
diff normal-headers.txt googlebot-headers.txt
diff normal-headers.txt baiduspider-headers.txtCompare response bodies:
diff normal-body.html googlebot-body.html
diff normal-body.html baiduspider-body.htmlIf a regular request returns 200 but the simulated crawler returns a 403 or verification page, check the rules configured for User-Agent.
But there is an important limitation:
Changing the User-Agent can only reveal User-Agent-based differences; it cannot simulate a search engine's real source IP, ASN, region, or access behavior.
Therefore,curl -A Googlebota normal request does not prove that the real Googlebot is not blocked by IP rules.
Do not allow search engines solely by User-Agent
User-Agent is easy to spoof.
Any request can claim to be:
Googlebot
BaiduspiderIf the WAF bypasses all security rules simply because it sees this string, attackers can exploit that rule as well.
The correct approach is to verify the source IP.
How to verify Googlebot
Google officially recommends using one of the following methods:
Match the visiting IP against Google's published crawler IP ranges
Perform a reverse DNS lookup on the source IP, then a forward DNS lookup on the resulting hostname, and confirm that it resolves back to the original IP
The verification process provided by Google is:
host 66.249.66.1Get a result similar to:
crawl-66-249-66-1.googlebot.comThen perform a forward lookup on this hostname:
host crawl-66-249-66-1.googlebot.comConfirm that the result matches the original IP.
Google states that checking only the User-Agent is unreliable; use IP ranges or forward and reverse DNS verification.
How to verify Baiduspider
After finding the source IP in server or CDN logs, you can run:
host 123.125.66.120Baidu's official documentation states that Baiduspider hostnames usually use*.baidu.comor*.baidu.jpformat, and recommends using reverse DNS lookup to determine the crawl source.
To reduce the risk of forged reverse DNS, when actually configuring allowlists you should also confirm that the resolved result matches the original IP, rather than matching only the hostname string.
If the CDN provides “Verified Search Engine Bot” or “Verified Bot” capabilities, prefer the vendor-maintained verification mechanism over manually maintaining IP lists that go years without updates.
The seven most common blocking rules after moving to a CDN
1. Bot management requires JavaScript verification
Some CDNs return a JavaScript challenge or CAPTCHA to suspicious requests.
Regular browsers can execute JavaScript and save cookies, but search engine crawlers may not be able to pass this challenge.
Check the following in the CDN console:
Bot管理
浏览器完整性检查
JavaScript Challenge
Managed Challenge
验证码规则
人机验证For verified Googlebot and Baiduspider requests, avoid returning interactive verification pages.
2. The WAF identifies high-frequency crawling as a scanning attack
Search engines may access multiple pages in rapid succession.
The WAF may trigger rules due to the following characteristics:
A single IP accesses too many pages
Requests are spaced very close together
A large number of different URLs are accessed
No cookies
Images are not loaded
Request headers differ from a regular browser
Continuously accesses new pages in the sitemap
The logs usually show:
403
429
WAF Block
Rate Limit
Bot Score
Request Rate ExceededDo not simply turn off the entire WAF. First identify the specific rule, then adjust its action for verified search engine requests.
3. Geographic blocking mistakenly affects Googlebot
Some websites target only users in mainland China, so they block overseas IPs directly in the CDN.
But Googlebot crawl requests may come from regions where the website is not open. If geographic rules run before search engine verification, Google crawling may receive a 403.
Similarly, some websites targeting overseas users block mainland China IPs, which can affect Baiduspider.
Check:
国家和地区访问控制
海外访问限制
仅允许中国大陆
仅允许指定国家
数据中心IP封禁
ASN封禁
代理和云主机封禁Search engine crawlers often use data center networks. If rules directly block all cloud provider or data center IPs, they may also mistakenly block legitimate crawlers.
4. Rate limiting returns 403 instead of 429
If the server genuinely needs to temporarily limit crawling due to load, the more appropriate status code is usually:
429 Too Many Requestsor:
503 Service Unavailable
Retry-After: 120Google explicitly recommends that when a CDN needs to temporarily block crawling, it can use 429 or 503 to let crawling systems know the condition is temporary.
Do not use 403 to mean “too busy right now, please try again later.” A 403 is more like an explicit denial of access and is not suitable for controlling crawl frequency.
5. IPv6 crawlers are not in the allowlist
If a search engine accesses over IPv6 but the WAF allowlist contains only IPv4, the following may occur:
IPv4抓取正常
IPv6抓取403或超时Confirm:
Whether CDN logs record IPv6 sources
Whether the allowlist supports both IPv4 and IPv6
Whether the origin server incorrectly restricts CDN IPv6 origin fetch addresses
Whether the firewall allows only the old IPv4 node ranges
Whether the search engine verification logic supports IPv6
Do not disable all IPv6 just for convenience. First use logs to confirm the address type used by the anomalous requests.
6. HEAD requests or special methods are blocked
Some search tools, monitoring programs, or resource verification requests may use HEAD.
If the CDN allows only GET and POST, HEAD requests may receive a 403 or 405.
Test:
curl -I https://www.example.com/articleand compare with GET:
curl -sS -D - -o /dev/null \
https://www.example.com/articleIf GET returns 200 and HEAD returns 403 or 405, determine whether the website intentionally restricts it or a CDN rule is mistakenly blocking it.
7. The CDN cached an old 403 or 5xx response
When a site first moves to a CDN, if the origin firewall has not yet allowed origin fetch requests, CDN nodes may initially receive a 403 or 502.
Later, the origin configuration is fixed, but nodes still return the old error page.
Check:
curl -I https://www.example.com/articleFocus on:
Age
X-Cache
Cache-Status
CDN厂商特有缓存字段If 403, 404, or 5xx responses show a cache hit, check the error status code caching rules and purge the affected URLs.
Do not purge the entire site at first. First purge one problem page androbots.txtto verify whether the cache is indeed the cause.
robots.txt must be checked separately
When search engines access a website, they usually first fetch:
https://www.example.com/robots.txtCheck directly:
curl -sS -D - \
https://www.example.com/robots.txtNormally you should see:
HTTP/2 200
Content-Type: text/plainand the expected rules, for example:
User-agent: *
Disallow:
Sitemap: https://www.example.com/sitemap.xmlWatch out for:
robots.txt returns 403
robots.txt returns 429 or 5xx
robots.txt redirects to a login page
robots.txt returns a CAPTCHA page
The CDN cached an old version of robots.txt
Different regions return different robots.txt files
Misconfigured
Disallow: /The HTTP and HTTPS versions have different content
With
wwwand withoutwwwuse different rules
Baidu Search Resource Platform states that Baiduspider first accesses robots.txt in the site root directory and determines its crawl scope according to the rules in that file.
Google also emphasizes that robots.txt is mainly used to manage crawling, not to reliably prevent URLs from appearing in search results.
Therefore, it is not enough for robots.txt to return 200; its content must also be correct.
Check whether pages have been accidentally given a noindex
CDN response header rules may add the following to an entire directory or even the whole site:
X-Robots-Tag: noindexThe website template may also contain:
<meta name="robots" content="noindex,nofollow">Check response headers:
curl -sS -D - -o /dev/null \
https://www.example.com/article |
grep -i 'x-robots-tag'Check the HTML:
curl -sS https://www.example.com/article |
grep -i 'robots'In particular, check:
Whether the production site inherited test environment configuration
Whether the CDN adds it uniformly for a certain path
X-Robots-TagWhether the mobile and desktop versions differ
Whether error pages are cached as normal pages
Whether different User-Agents receive different robots tags
Successful crawling and permission to index are not the same thing.
The page returns 200 but containsnoindex, and search engines may still not index it.
Bypass the CDN to compare origin results
If you suspect the CDN has changed status codes or content, access the origin server directly, bypassing the CDN.
Assume:
网站域名:www.example.com
源站IP:203.0.113.10Request through the CDN:
curl -sS -D cdn-headers.txt \
-o cdn-body.html \
https://www.example.com/articleRequest the origin server while bypassing the CDN, keeping the correct Host and HTTPS SNI:
curl -sS -D origin-headers.txt \
-o origin-body.html \
--resolve www.example.com:443:203.0.113.10 \
https://www.example.com/articleThen compare:
diff cdn-headers.txt origin-headers.txt
diff cdn-body.html origin-body.htmlThe results can be interpreted as follows:
CDN result | Origin result | Investigate first |
|---|---|---|
403 | 200 | CDN WAF, bot management, geographic, or rate limiting rules |
429 | 200 | CDN rate limiting |
503 | 200 | CDN nodes, health checks, or edge rules |
Verification page 200 | Normal article 200 | JS challenge or bot protection |
robots.txt disallows crawling | Origin allows crawling | CDN cached an old robots.txt |
Both return 403 | Origin WAF, application permissions, or access control | |
Both return 5xx | Origin and application failure | |
CDN and origin content differ | Caching, edge rewrites, or content returned by UA |
If crawling errors occur only in certain regions, you can also use CdnChart'swebsite speed test toolto check status codes and access results in different regions.
If you are not sure which CDN the domain currently uses or which nodes it resolves to, you can first useCDN Checker toolto check CNAME and node information.
What should you look for in logs?
Looking only at origin access logs may not be enough.
If a request is blocked at the CDN edge, the origin server never receives it. Therefore, check all of the following:
CDN access logs
CDN security event logs
WAF block logs
Bot management logs
Rate limiting logs
Origin access logs
Origin error logs
Application logs
It is recommended to filter by the following fields:
时间
完整URL
User-Agent
来源IP
国家和地区
ASN
HTTP状态码
命中的安全规则
执行动作
CDN节点
请求ID
回源状态码
回源耗时If origin logs record only the CDN origin fetch IP, use CDN logs or a trusted client IP field to view the true source. Do not fully trust client-submitted values without proper configuration.X-Forwarded-For, otherwise logs and security rules may be spoofed.
Recheck with Google and Baidu official tools
Google Search Console
After fixing, you can use:
URL Inspection
Test Live URL
Page Indexing report
Crawl Stats
HTTPS report
robots.txt-related checks
Google recommends using Crawl Stats to view Googlebot's crawl history and site availability issues.
Focus on:
主机状态
按响应类型统计的抓取请求
5xx数量
抓取响应时间
Googlebot类型
异常开始时间Baidu Search Resource Platform
You can use the Crawl Diagnostics in Baidu Search Resource Platform to view from Baiduspider's perspective:
Whether it can connect to the website
The returned status
The crawled page content
Whether the crawl IP is correct
Whether the page matches what regular users see
Baidu's description of the Crawl Diagnostics tool states that it can be used to check the connection between the website and Baidu and to view what Baiduspider actually crawls.
A single successful request from an official tool does not mean all nodes have fully recovered, but it does show that at least the current diagnostic request can access the site normally.
Do not immediately conclude that “indexing has recovered” after fixing
After changing CDN rules, it is recommended to verify in this order:
Purge the erroneous cache for affected URLs, robots.txt, and sitemaps.
Test status codes and response bodies with a regular User-Agent.
Perform an initial comparison using Googlebot and Baiduspider User-Agents.
Check CDN security logs to confirm that blocking is no longer triggered.
Use Google URL Inspection to test the live URL.
Use Baidu Crawl Diagnostics to recrawl.
Monitor crawl stats, server logs, and indexing reports.
Check access results for different regions and for IPv4 and IPv6.
Do not repeatedly submit thousands of URLs after fixing the rules, and do not expect indexing volume to fully recover the same day.
Search engines need to recrawl pages and reprocess content. Recovery speed depends on how long the errors lasted, the number of affected URLs, the site's crawl frequency, and page quality; it cannot be determined by a fixed number of days.
Frequently Asked Questions
The browser opens normally, so why does Google still fail to crawl?
The CDN may return different results based on IP, region, User-Agent, request frequency, or cookies. A normal browser experience only shows that your access conditions did not trigger a rule; it does not mean Googlebot receives the same response.
If simulating Googlebot returns 200, does that mean Google crawling is normal?
No.curl -A GooglebotYou can simulate only the User-Agent, not Google's real source IP, network, or access behavior. You still need to check CDN logs, verify the real Googlebot IP, and test with Search Console.
Should Googlebot and Baiduspider be added to the allowlist?
You can reduce unnecessary bot challenges and rate limiting for verified search engine crawlers, but do not allow them based solely on User-Agent, and do not let the allowlist bypass all high-risk security rules.
When the website is under heavy load, can I use 403 to limit crawlers?
It is not recommended. For temporary overload, it is more appropriate to return 429 or 503, and provide as appropriateRetry-After. A 403 means access denied and is not suitable for controlling temporary crawl pressure.
Will a 404 for robots.txt affect crawling?
Different search engines may handle an unavailable robots.txt differently. The safest approach is to make it consistently return 200 with correct plain text content, and ensure CDN verification pages, redirect chains, or WAF blocking do not affect the file.
Why does only Baidu fail to crawl while Google is normal?
This may be due to mainland China nodes, Baiduspider source IPs, geographic rules, or different WAF policies for Baiduspider. It is also possible that Baidu and Google are directed to different CDN nodes; compare their logs, nodes, and response content.
Why does only Google fail to crawl while Baidu is normal?
First check overseas region blocking, data center IP restrictions, IPv6 rules, and bot challenges. Googlebot's access source differs from Baidu's crawler, so even when accessing the same URL, they may hit completely different CDN policies.
Will a CDN returning 403 immediately cause a page to drop out of the index?
It may not happen immediately, but continued inability to crawl can affect search engines' ability to update page content and maintain the index. The sooner you identify the scope of blocking, fix the rules, and restore stable 200 responses, the better.
- Googlebot crawl failure
- Baiduspider crawl errors
- CDN impact on Baidu indexing
- CDN impact on Google indexing
- Spider access 403
- Search engine crawl 429 errors
- CDN blocks crawlers
- WAF blocks Googlebot