Multilingual Static Sites: The Source-Output Trap
This site has 180 pages across five languages. One day the English pages had English navigation but entirely Chinese body text — and it had been live for days. The cause wasn't missing translations; it was build order.
Symptom: only the nav was translated
The dropdown correctly showed EN and the nav was in English, but the hero, cards and lists were all Chinese. This half-right, half-wrong shape is easy to misread as missing translation data.
The real diagnosis is not the page but the source: count how many language spans remain in each page's body.
# 统计 </nav> 之后的多语 span 数量
import re
s = open("src5/tools/chatgpt.html", encoding="utf-8").read()
body = s[s.find("</nav>"):]
print(len(re.findall(r'class="(?:zh|en|id|sl|it)"', body)))
The answer was 0. All 36 source pages had zero body spans; only the 25 in the header nav survived, because a separate script had re-added those later.
Root cause: source and output in one directory
The pipeline worked like this: generators wrote pages with five-language spans to the root directory, then the build step folded those spans into per-language pages and wrote the Chinese version back to the same path, since deployment needs Chinese at the root.
That step overwrites the source with an already-folded single-language page. The next snapshot then captures pages that have no spans left, so folding has nothing to choose from and all five languages come out Chinese.
Key lesson: whenever source and output share a directory and the build writes output back, you must snapshot the multilingual source elsewhere before folding. Order isn't a habit here — it's correctness.
Fix: make the order executable
Memorising the order does not work: conversations forget it, people get it wrong. Put the order in a script so it becomes part of the code, and make a missing step fail the build.
# build_all.py —— 唯一构建入口
STEPS = [
("i18n_lint.py", "翻译数据自检"),
("gen_content.py", "生成内容页(带 5 语 span)"),
# ...
("inject_seo.py", "注入 SEO 头(仅根页)"),
("snapshot_src5.py", "快照多语源 ← 必须在折叠之前"),
("build_i18n.py", "折叠成 180 个单语页"),
("audit_site.py", "全站自检,不过则不部署"),
]
- The snapshot step self-audits: any page with zero body spans or a duplicated footer exits with code 1
- The injection step is limited to root pages, skipping src5 and the language subdirectories
Verify: don't eyeball the pages
Eyeballing only covers the page you happened to open, and it cannot see time-order problems such as "source already degraded while output still looks fine". Let the script report the numbers.
# 构建末尾断言
assert span_count(body) > 0, "正文无多语 span(会退化成单语)"
assert footer_count(page) == 1, "页脚重复"
assert len(re.findall(r'rel="alternate" hreflang', head)) == 6
The last line came from a real misdiagnosis: counting hreflang across the whole page also counted the language menu's <a hreflang> attributes, wrongly reporting duplicates. The <head> actually had exactly six. When a number looks wrong, verify the measurement scope before changing product code.