Back to blog

Baidu and Google crawling issues after CDN integration: How do I check status codes and blocking rules?

CdnChart Technical TeamPublished on 2026-10-0314 min read
Baidu and Google crawling issues after CDN integration: How do I check status codes and blocking rules?

After a website is put behind a CDN, regular users can access it normally, but Baidu Search Resource Platform or Google Search Console begins reporting crawl errors.

Common symptoms include:

服务器错误(5xx)
网页无法访问
已发现但尚未抓取
抓取频率下降
robots.txt无法获取
网址检查返回403
百度抓取诊断失败
索引页面逐渐减少

Many site owners first check the homepage, see that it opens normally in a browser, and assume search engine crawling should be fine as well.

But a normal browser experience does not mean Baidu and Google see the same result.

CDNs, WAFs, and bot protection systems may return different content based on the visitor's IP, region, User-Agent, request frequency, cookies, and JavaScript execution capability. Regular users may get a 200, while search engine crawlers may get a 403, 429, or 503—or even a CAPTCHA page with a 200 status code.

So what really needs to be checked is:

When Baidu and Google send requests, what status code and content does the CDN ultimately return?

Which layers does search engine crawling pass through?

After a site is placed behind a CDN, search engine crawling typically passes through:

Googlebot或Baiduspider
        ↓
DNS解析与CDN调度
        ↓
CDN边缘节点
        ↓
Bot管理、WAF、频控和地区规则
        ↓
缓存
        ↓
源站Web服务器
        ↓
网站程序

Any layer can cause crawling to fail.

For example:

  • DNS directs the crawler to an abnormal node

  • The CDN blocks access from the crawler's region

  • The WAF identifies high-frequency crawling as an attack

  • Bot management requires JavaScript verification

  • Rate limiting returns 429

  • The CDN fails to fetch from origin and returns 502 or 504

  • The cache retains an old 403 page

  • robots.txt is blocked by a CDN rule

  • The origin server returns different content based on User-Agent

Therefore, “not indexed after moving to a CDN” does not directly prove that the CDN is affecting SEO. You must first use status codes, response content, and logs to identify the specific failure.

Check the most important URLs first

Do not check only the homepage. Test at least the following addresses:

https://www.example.com/
https://www.example.com/robots.txt
https://www.example.com/sitemap.xml
一篇已经收录的文章
一篇新发布的文章
一个出现抓取异常的具体URL
页面引用的主要CSS和JavaScript

First check the status code for a regular request:

curl -sS -D - -o /dev/null \
  https://www.example.com/article

Then save the full response:

curl -sS -D response-headers.txt \
  -o response-body.html \
  https://www.example.com/article

You need to check more than the status code on the first line. Also check:

Location
Retry-After
Server
Via
Age
X-Cache
X-Robots-Tag
Content-Type
CDN请求ID
WAF事件ID

If the status code is 200, also openresponse-body.htmlto view the content and confirm that it returns the normal article, not:

Access Denied
Checking your browser
Please enable JavaScript
Human Verification
Request blocked
网站维护中

A normal status code with an error page as the content can still affect crawling and indexing.

What do different status codes mean for crawling?

Status code

What search engines see

Check first

200

Content fetched successfully, but it may still be a CAPTCHA or soft 404

Page body, CAPTCHA,noindex

301/308

Permanent redirect

Target address and redirect chain

302/307

Temporary redirect

Whether it incorrectly redirects to the homepage, login page, or verification page

401

Authentication required

CDN access control, origin authentication

403

Access explicitly denied

WAF, geographic blocking, IP rules, bot rules

404/410

Page does not exist

URL, cache, routing, and publish status

429

Too many requests

CDN rate limiting, bot rate limiting, origin rate limiting

500

Origin server internal error

Application and server logs

502

CDN received an abnormal upstream response

Origin fetch protocol, port, and connection

503

Service temporarily unavailable

Overload, maintenance, no healthy origin server

504

CDN timed out waiting for the origin server

Application, database, and origin fetch timeouts

Google states that sustained 5xx and 429 responses cause its crawling systems to temporarily slow down crawling; URLs that persistently return server errors may eventually drop out of the index.

Network timeouts, connection resets, and DNS errors have no HTTP status code, but Google still treats them as serious availability problems.

A 200 status code can also indicate a crawl error

The most insidious case is not a 403, but when the CDN returns:

HTTP/2 200 OK
Content-Type: text/html

while the page content is:

正在验证您的浏览器……
请完成验证码……
访问过于频繁……

For regular users, the browser may execute JavaScript and write a cookie, after which they automatically enter the site; but the search engine may only crawl the verification page.

This kind of response may also be judged as a soft 404 or erroneous content.

Google has specifically warned that when a CDN returns a random error page with a 200 status code, it makes it difficult for search systems to correctly understand the page status.

Therefore, when checking, you must compare:

  • HTTP status code

  • Page title

  • Page body

  • Content-Type

  • canonical

  • robots meta

  • X-Robots-Tag

  • Content differences between regular user and crawler requests

For example, search for characteristics in the response content:

curl -sS https://www.example.com/article |
  grep -Ei 'captcha|access denied|checking your browser|noindex'

If the page should be an article but the returned content is only a few KB of verification code, you cannot consider crawling normal just because the status code is 200.

Perform an initial comparison using User-Agent

You can simulate the User-Agents for Googlebot and Baiduspider to see whether the CDN returns different results.

Test a regular request:

curl -sS -D normal-headers.txt \
  -o normal-body.html \
  https://www.example.com/article

Simulate Googlebot:

curl -sS -D googlebot-headers.txt \
  -o googlebot-body.html \
  -A 'Mozilla/5.0 (compatible; Googlebot/2.1; +http://www.google.com/bot.html)' \
  https://www.example.com/article

Simulate Baiduspider:

curl -sS -D baiduspider-headers.txt \
  -o baiduspider-body.html \
  -A 'Mozilla/5.0 (compatible; Baiduspider/2.0; +http://www.baidu.com/search/spider.html)' \
  https://www.example.com/article

Compare response headers:

diff normal-headers.txt googlebot-headers.txt
diff normal-headers.txt baiduspider-headers.txt

Compare response bodies:

diff normal-body.html googlebot-body.html
diff normal-body.html baiduspider-body.html

If a regular request returns 200 but the simulated crawler returns a 403 or verification page, check the rules configured for User-Agent.

But there is an important limitation:

Changing the User-Agent can only reveal User-Agent-based differences; it cannot simulate a search engine's real source IP, ASN, region, or access behavior.

Therefore,curl -A Googlebota normal request does not prove that the real Googlebot is not blocked by IP rules.

Do not allow search engines solely by User-Agent

User-Agent is easy to spoof.

Any request can claim to be:

Googlebot
Baiduspider

If the WAF bypasses all security rules simply because it sees this string, attackers can exploit that rule as well.

The correct approach is to verify the source IP.

How to verify Googlebot

Google officially recommends using one of the following methods:

  • Match the visiting IP against Google's published crawler IP ranges

  • Perform a reverse DNS lookup on the source IP, then a forward DNS lookup on the resulting hostname, and confirm that it resolves back to the original IP

The verification process provided by Google is:

host 66.249.66.1

Get a result similar to:

crawl-66-249-66-1.googlebot.com

Then perform a forward lookup on this hostname:

host crawl-66-249-66-1.googlebot.com

Confirm that the result matches the original IP.

Google states that checking only the User-Agent is unreliable; use IP ranges or forward and reverse DNS verification.

How to verify Baiduspider

After finding the source IP in server or CDN logs, you can run:

host 123.125.66.120

Baidu's official documentation states that Baiduspider hostnames usually use*.baidu.comor*.baidu.jpformat, and recommends using reverse DNS lookup to determine the crawl source.

To reduce the risk of forged reverse DNS, when actually configuring allowlists you should also confirm that the resolved result matches the original IP, rather than matching only the hostname string.

If the CDN provides “Verified Search Engine Bot” or “Verified Bot” capabilities, prefer the vendor-maintained verification mechanism over manually maintaining IP lists that go years without updates.

The seven most common blocking rules after moving to a CDN

1. Bot management requires JavaScript verification

Some CDNs return a JavaScript challenge or CAPTCHA to suspicious requests.

Regular browsers can execute JavaScript and save cookies, but search engine crawlers may not be able to pass this challenge.

Check the following in the CDN console:

Bot管理
浏览器完整性检查
JavaScript Challenge
Managed Challenge
验证码规则
人机验证

For verified Googlebot and Baiduspider requests, avoid returning interactive verification pages.

2. The WAF identifies high-frequency crawling as a scanning attack

Search engines may access multiple pages in rapid succession.

The WAF may trigger rules due to the following characteristics:

  • A single IP accesses too many pages

  • Requests are spaced very close together

  • A large number of different URLs are accessed

  • No cookies

  • Images are not loaded

  • Request headers differ from a regular browser

  • Continuously accesses new pages in the sitemap

The logs usually show:

403
429
WAF Block
Rate Limit
Bot Score
Request Rate Exceeded

Do not simply turn off the entire WAF. First identify the specific rule, then adjust its action for verified search engine requests.

3. Geographic blocking mistakenly affects Googlebot

Some websites target only users in mainland China, so they block overseas IPs directly in the CDN.

But Googlebot crawl requests may come from regions where the website is not open. If geographic rules run before search engine verification, Google crawling may receive a 403.

Similarly, some websites targeting overseas users block mainland China IPs, which can affect Baiduspider.

Check:

国家和地区访问控制
海外访问限制
仅允许中国大陆
仅允许指定国家
数据中心IP封禁
ASN封禁
代理和云主机封禁

Search engine crawlers often use data center networks. If rules directly block all cloud provider or data center IPs, they may also mistakenly block legitimate crawlers.

4. Rate limiting returns 403 instead of 429

If the server genuinely needs to temporarily limit crawling due to load, the more appropriate status code is usually:

429 Too Many Requests

or:

503 Service Unavailable
Retry-After: 120

Google explicitly recommends that when a CDN needs to temporarily block crawling, it can use 429 or 503 to let crawling systems know the condition is temporary.

Do not use 403 to mean “too busy right now, please try again later.” A 403 is more like an explicit denial of access and is not suitable for controlling crawl frequency.

5. IPv6 crawlers are not in the allowlist

If a search engine accesses over IPv6 but the WAF allowlist contains only IPv4, the following may occur:

IPv4抓取正常
IPv6抓取403或超时

Confirm:

  • Whether CDN logs record IPv6 sources

  • Whether the allowlist supports both IPv4 and IPv6

  • Whether the origin server incorrectly restricts CDN IPv6 origin fetch addresses

  • Whether the firewall allows only the old IPv4 node ranges

  • Whether the search engine verification logic supports IPv6

Do not disable all IPv6 just for convenience. First use logs to confirm the address type used by the anomalous requests.

6. HEAD requests or special methods are blocked

Some search tools, monitoring programs, or resource verification requests may use HEAD.

If the CDN allows only GET and POST, HEAD requests may receive a 403 or 405.

Test:

curl -I https://www.example.com/article

and compare with GET:

curl -sS -D - -o /dev/null \
  https://www.example.com/article

If GET returns 200 and HEAD returns 403 or 405, determine whether the website intentionally restricts it or a CDN rule is mistakenly blocking it.

7. The CDN cached an old 403 or 5xx response

When a site first moves to a CDN, if the origin firewall has not yet allowed origin fetch requests, CDN nodes may initially receive a 403 or 502.

Later, the origin configuration is fixed, but nodes still return the old error page.

Check:

curl -I https://www.example.com/article

Focus on:

Age
X-Cache
Cache-Status
CDN厂商特有缓存字段

If 403, 404, or 5xx responses show a cache hit, check the error status code caching rules and purge the affected URLs.

Do not purge the entire site at first. First purge one problem page androbots.txtto verify whether the cache is indeed the cause.

robots.txt must be checked separately

When search engines access a website, they usually first fetch:

https://www.example.com/robots.txt

Check directly:

curl -sS -D - \
  https://www.example.com/robots.txt

Normally you should see:

HTTP/2 200
Content-Type: text/plain

and the expected rules, for example:

User-agent: *
Disallow:

Sitemap: https://www.example.com/sitemap.xml

Watch out for:

  • robots.txt returns 403

  • robots.txt returns 429 or 5xx

  • robots.txt redirects to a login page

  • robots.txt returns a CAPTCHA page

  • The CDN cached an old version of robots.txt

  • Different regions return different robots.txt files

  • MisconfiguredDisallow: /

  • The HTTP and HTTPS versions have different content

  • Withwwwand withoutwwwuse different rules

Baidu Search Resource Platform states that Baiduspider first accesses robots.txt in the site root directory and determines its crawl scope according to the rules in that file.

Google also emphasizes that robots.txt is mainly used to manage crawling, not to reliably prevent URLs from appearing in search results.

Therefore, it is not enough for robots.txt to return 200; its content must also be correct.

Check whether pages have been accidentally given a noindex

CDN response header rules may add the following to an entire directory or even the whole site:

X-Robots-Tag: noindex

The website template may also contain:

<meta name="robots" content="noindex,nofollow">

Check response headers:

curl -sS -D - -o /dev/null \
  https://www.example.com/article |
  grep -i 'x-robots-tag'

Check the HTML:

curl -sS https://www.example.com/article |
  grep -i 'robots'

In particular, check:

  • Whether the production site inherited test environment configuration

  • Whether the CDN adds it uniformly for a certain pathX-Robots-Tag

  • Whether the mobile and desktop versions differ

  • Whether error pages are cached as normal pages

  • Whether different User-Agents receive different robots tags

Successful crawling and permission to index are not the same thing.

The page returns 200 but containsnoindex, and search engines may still not index it.

Bypass the CDN to compare origin results

If you suspect the CDN has changed status codes or content, access the origin server directly, bypassing the CDN.

Assume:

网站域名:www.example.com
源站IP:203.0.113.10

Request through the CDN:

curl -sS -D cdn-headers.txt \
  -o cdn-body.html \
  https://www.example.com/article

Request the origin server while bypassing the CDN, keeping the correct Host and HTTPS SNI:

curl -sS -D origin-headers.txt \
  -o origin-body.html \
  --resolve www.example.com:443:203.0.113.10 \
  https://www.example.com/article

Then compare:

diff cdn-headers.txt origin-headers.txt
diff cdn-body.html origin-body.html

The results can be interpreted as follows:

CDN result

Origin result

Investigate first

403

200

CDN WAF, bot management, geographic, or rate limiting rules

429

200

CDN rate limiting

503

200

CDN nodes, health checks, or edge rules

Verification page 200

Normal article 200

JS challenge or bot protection

robots.txt disallows crawling

Origin allows crawling

CDN cached an old robots.txt

Both return 403

Origin WAF, application permissions, or access control

Both return 5xx

Origin and application failure

CDN and origin content differ

Caching, edge rewrites, or content returned by UA

If crawling errors occur only in certain regions, you can also use CdnChart'swebsite speed test toolto check status codes and access results in different regions.

If you are not sure which CDN the domain currently uses or which nodes it resolves to, you can first useCDN Checker toolto check CNAME and node information.

What should you look for in logs?

Looking only at origin access logs may not be enough.

If a request is blocked at the CDN edge, the origin server never receives it. Therefore, check all of the following:

  • CDN access logs

  • CDN security event logs

  • WAF block logs

  • Bot management logs

  • Rate limiting logs

  • Origin access logs

  • Origin error logs

  • Application logs

It is recommended to filter by the following fields:

时间
完整URL
User-Agent
来源IP
国家和地区
ASN
HTTP状态码
命中的安全规则
执行动作
CDN节点
请求ID
回源状态码
回源耗时

If origin logs record only the CDN origin fetch IP, use CDN logs or a trusted client IP field to view the true source. Do not fully trust client-submitted values without proper configuration.X-Forwarded-For, otherwise logs and security rules may be spoofed.

Recheck with Google and Baidu official tools

Google Search Console

After fixing, you can use:

  • URL Inspection

  • Test Live URL

  • Page Indexing report

  • Crawl Stats

  • HTTPS report

  • robots.txt-related checks

Google recommends using Crawl Stats to view Googlebot's crawl history and site availability issues.

Focus on:

主机状态
按响应类型统计的抓取请求
5xx数量
抓取响应时间
Googlebot类型
异常开始时间

Baidu Search Resource Platform

You can use the Crawl Diagnostics in Baidu Search Resource Platform to view from Baiduspider's perspective:

  • Whether it can connect to the website

  • The returned status

  • The crawled page content

  • Whether the crawl IP is correct

  • Whether the page matches what regular users see

Baidu's description of the Crawl Diagnostics tool states that it can be used to check the connection between the website and Baidu and to view what Baiduspider actually crawls.

A single successful request from an official tool does not mean all nodes have fully recovered, but it does show that at least the current diagnostic request can access the site normally.

Do not immediately conclude that “indexing has recovered” after fixing

After changing CDN rules, it is recommended to verify in this order:

  1. Purge the erroneous cache for affected URLs, robots.txt, and sitemaps.

  2. Test status codes and response bodies with a regular User-Agent.

  3. Perform an initial comparison using Googlebot and Baiduspider User-Agents.

  4. Check CDN security logs to confirm that blocking is no longer triggered.

  5. Use Google URL Inspection to test the live URL.

  6. Use Baidu Crawl Diagnostics to recrawl.

  7. Monitor crawl stats, server logs, and indexing reports.

  8. Check access results for different regions and for IPv4 and IPv6.

Do not repeatedly submit thousands of URLs after fixing the rules, and do not expect indexing volume to fully recover the same day.

Search engines need to recrawl pages and reprocess content. Recovery speed depends on how long the errors lasted, the number of affected URLs, the site's crawl frequency, and page quality; it cannot be determined by a fixed number of days.

Frequently Asked Questions

The browser opens normally, so why does Google still fail to crawl?

The CDN may return different results based on IP, region, User-Agent, request frequency, or cookies. A normal browser experience only shows that your access conditions did not trigger a rule; it does not mean Googlebot receives the same response.

If simulating Googlebot returns 200, does that mean Google crawling is normal?

No.curl -A GooglebotYou can simulate only the User-Agent, not Google's real source IP, network, or access behavior. You still need to check CDN logs, verify the real Googlebot IP, and test with Search Console.

Should Googlebot and Baiduspider be added to the allowlist?

You can reduce unnecessary bot challenges and rate limiting for verified search engine crawlers, but do not allow them based solely on User-Agent, and do not let the allowlist bypass all high-risk security rules.

When the website is under heavy load, can I use 403 to limit crawlers?

It is not recommended. For temporary overload, it is more appropriate to return 429 or 503, and provide as appropriateRetry-After. A 403 means access denied and is not suitable for controlling temporary crawl pressure.

Will a 404 for robots.txt affect crawling?

Different search engines may handle an unavailable robots.txt differently. The safest approach is to make it consistently return 200 with correct plain text content, and ensure CDN verification pages, redirect chains, or WAF blocking do not affect the file.

Why does only Baidu fail to crawl while Google is normal?

This may be due to mainland China nodes, Baiduspider source IPs, geographic rules, or different WAF policies for Baiduspider. It is also possible that Baidu and Google are directed to different CDN nodes; compare their logs, nodes, and response content.

Why does only Google fail to crawl while Baidu is normal?

First check overseas region blocking, data center IP restrictions, IPv6 rules, and bot challenges. Googlebot's access source differs from Baidu's crawler, so even when accessing the same URL, they may hit completely different CDN policies.

Will a CDN returning 403 immediately cause a page to drop out of the index?

It may not happen immediately, but continued inability to crawl can affect search engines' ability to update page content and maintain the index. The sooner you identify the scope of blocking, fix the rules, and restore stable 200 responses, the better.

  • Googlebot crawl failure
  • Baiduspider crawl errors
  • CDN impact on Baidu indexing
  • CDN impact on Google indexing
  • Spider access 403
  • Search engine crawl 429 errors
  • CDN blocks crawlers
  • WAF blocks Googlebot

Related posts

Frontend API Requests Blocked by CORS? Here's How to Check CDN Response Headers and Preflight Requests

Frontend API Requests Blocked by CORS? Here's How to Check CDN Response Headers and Preflight Requests

When a frontend call to an API returns a CORS error, the cause may lie in the origin server, CDN caching, the OPTIONS preflight request, a WAF, or duplicate response headers. This article covers browser and curl checks to help you determine at which layer the cross-origin response headers are being dropped.

12 min read
Some users cannot access the site after enabling IPv6: Is it a DNS issue or a CDN node issue?

Some users cannot access the site after enabling IPv6: Is it a DNS issue or a CDN node issue?

After IPv6 is enabled on a website, users in some regions or on certain carriers may be unable to reach it. The cause may lie in the AAAA record, the user's IPv6 network, a CDN node, TLS, MTU, or IPv6 origin fetch. This article covers DNS, curl, and route testing methods to help determine which layer the failure occurs at.

12 min read
Origin Server Accessible but CDN Returns 404? How to Troubleshoot Origin Host and Cache

Origin Server Accessible but CDN Returns 404? How to Troubleshoot Origin Host and Cache

Origin server responding normally, but you get a 404 after putting the CDN in front of it? This article shows you how to tell whether the 404 comes from a CDN node or from your origin, then walks you step by step through troubleshooting origin Host header, port, protocol, caching, URL rewriting, and multi-origin configuration issues.

10 min read
Are CDNs Right for Dynamic APIs? Distinguish Caching From Dynamic Acceleration

Are CDNs Right for Dynamic APIs? Distinguish Caching From Dynamic Acceleration

Can dynamic APIs use a CDN? This article explains the differences between API caching, dynamic acceleration, and edge computing, analyzes caching strategies for public endpoints, login endpoints, and GET versus POST requests, and covers Cache-Control, cache keys, security checks, and multi-region speed testing.

20 min read