Operazioni Cloudflare Pages: contenuti duplicati, redirect e deploy più rapidi
Search Console ha segnalato pagine duplicate e ho subito pensato che il codice fosse sbagliato. Le misurazioni hanno mostrato la stessa home che restituisce 200 su tre indirizzi con md5 identico. Il codice è corretto; la configurazione di deploy no.
Three addresses, one document
$ curl -s -o /dev/null -w "%{http_code} → %{redirect_url}\n" https://fusi.si/
200 →
$ curl -s -o /dev/null -w "%{http_code} → %{redirect_url}\n" https://www.fusi.si/
200 → # ← 应该 301 到主域
$ curl -s https://fusi.si/ | md5sum
a5ab014c... # www 版本一字不差
The three entry points are the apex domain, the www custom domain, and Cloudflare's auto-assigned pages.dev address. As long as all three return 200, Google treats them as duplicate content.
The good news: the canonical tag already points at the apex, so Google's chosen canonical is correct — it simply found two extra copies beyond the one you declared.
Three mechanisms, three jobs
- canonical — declares which URL is authoritative. Fully under your control; all 180 pages already point at the apex.
- 301 redirect — permanently sends www to the apex. The cleanest solution, but it lives in Cloudflare's Redirect Rules; code cannot do it.
- robots block — for addresses like pages.dev that cannot be redirected. It stops crawling but cannot fully stop link equity, which is why it pairs with canonical.
Trap: never block www in robots.txt. That blocks the redirect too, and the www host genuinely disappears. www must stay Allowed.
User-agent: *
Allow: /
Disallow: /fusi-si.pages.dev/ # 仅屏蔽无法跳转的部署地址
Sitemap: https://fusi.si/sitemap.xml
The .html 308 redirect
Pages enables pretty URLs by default: /tools/chatgpt.html 308-redirects to /tools/chatgpt. That is fine as long as internal links, canonical and sitemap all use the short form consistently.
This site reported 34 "page with redirect" entries — exactly the old .html addresses from earlier deploys. A scan of all 181 internal link types found zero .html links, confirming these come from off-site and historical sources, which only Google's recrawl can resolve.
The fastest way to diagnose this class of problem is to sweep internal links rather than inspect pages one by one:
grep -ro 'href="[^"]*\.html[^"]*"' --include="*.html" . | wc -l
# 0 = 站内干净,剩下的就是站外来的
Cut deploy time from minutes to seconds
npx wrangler re-resolves the package on every invocation; two attempts both ran past five minutes without finishing. npx had already downloaded it — it just kept re-resolving.
# 直接调用 npx 缓存里已下载的二进制
WR="$LOCALAPPDATA/npm-cache/_npx/<hash>/node_modules/.bin/wrangler"
$WR pages deploy _deploy --project-name=fusi-si # 几秒完成
- Assemble a separate deploy directory with static output only — no scripts, no source snapshots, no caches
- Keep credentials in a gitignored file rather than asking each time; write logs to a file instead of piping, since buffering makes the process look hung
- Chain build, self-check and deploy into one script so a failed check blocks deployment