Autocontrollo delle traduzioni: lasciate che la macchina colga gli errori
Scrivendo traduzioni in cinque lingue ho continuato a fare lo stesso errore: un verbo tedesco nell'indonesiano, testo inglese in uno slot cinese, uno slot in meno. Quasi invisibile all'occhio, ma compare subito in produzione.
First, understand the slot layout
Five languages are not always a 5-tuple. Real projects mix at least four layouts; a script that assumes "everything from index 1 is foreign" will flag perfectly legal content.
# 5 = (zh, en, id, sl, it)
# 10 = 问答按语交替 (zh_q, zh_a, en_q, en_a, ...)
# 6 = (name, zh, en, id, sl, it) 前两位是名称与中文
# 4 = (zh_q, zh_a, en_q, en_a)
# 3 = (id, sl, it) 根本没有中文槽位
The 3-tuple line is the trap: asserting "index 0 must be Chinese" flags perfectly fine Indonesian-only data as broken.
Three classes of error to catch
- 外语槽位含 CJK:印尼语里混进「訳」这样的日文汉字,或中文漏进英文槽Foreign slot containing CJK: a Japanese kanji like 訳 slipping into Indonesian, or Chinese leaking into an English slot
- 中英混排:一句里两种文字拼在一起,比如「如此Jdahit」「特征CORE功能」Mixed scripts mid-sentence: two writing systems fused, like 「如此Jdahit」 or 「特征CORE功能」
- Misaligned or empty slots: a tuple short by one element, or a language left blank
The third is the sneakiest: an empty slot silently falls back to a neighbouring language, so the page looks filled while the language is wrong.
Infer from structure, not position
def latin_indices(n):
"""给定槽位总数,返回「应该是拉丁语」的下标集合。"""
return {1: {1, 2, 3, 4},
2: {1},
3: {1, 2},
4: {2, 3},
5: {1, 2, 3, 4},
6: {2, 3, 4, 5},
9: {2, 3, 4, 5, 6, 7, 8},
10: {2, 3, 4, 5, 6, 7, 8, 9}}.get(n, set())
With that table the check becomes project-agnostic, and reusable for the next set of pages without rewriting it.
配套做法:校验数据时把断言跑在脚本里,不要跑在脑子里。我写完 181 条译文后靠脚本抓出了「如此Jdahit」「menjadikannya一套完整」这类错误——全是我肉眼读不出来的。Companion practice: run assertions in a script, not in your head. After writing 181 translations the script caught things like 「如此Jdahit」 that I could not see by reading.
Don't hand-edit data files
When appending large translation blocks via a shell script I produced a duplicated robots.txt group; editing the JSON by hand created nested lists and duplicated columns.
Conclusion: data files are only ever written by scripts, and one-off scripts must validate before writing — abort when validation fails.
if bad:
print("发现 %d 处问题,未写入" % bad)
return 1 # 中止,不碰文件
write_file(path, data)