Chinese has to be cut before it can be read中文得先切开才能读
I wanted to build a small Chinese decision model — the shape Jev and Laya have: text in, one option out of a list. Before any of that, something has to cut the sentence into pieces, because Chinese does not come with spaces.
净利润同比下降48.7%
By character: 净 / 利 / 润 / 同 / 比 / 下 / 降. By word: 净利润 / 同比 / 下降 — “net profit” / “year on year” / “fell”. The model never sees the sentence. It sees the pieces. So this choice fixes what the model is even able to notice, before a single weight is trained.
The question I set out to answer was whether characters are better than words. The answer is yes, and the interesting part is that almost none of the margin is where I expected it.
我想做一个小的中文决策模型,就是 Jev 和 Laya 那个形状:进去一段文本,从若干选项里挑一个出来。在这之前,必须先有一步把句子切成小块——中文不带空格。
净利润同比下降48.7%
按字切是 净 / 利 / 润 / 同 / 比 / 下 / 降;按词切是 净利润 / 同比 / 下降。模型永远看不到句子,它只看得到这些块。所以这一步在任何一个权重被训练之前,就已经决定了模型有没有可能注意到某些东西。
我要回答的问题是:字比词好吗?答案是好。有意思的是,这个”好”几乎完全不在我预期的地方。
The model I meant to change already does it我想改的那个模型,本来就是这么切的
Before building anything, I measured what Laya’s tokenizer — a 256,000-entry Gemma-style vocabulary, inherited from mmBERT — actually does to Chinese financial text. On 20,000 filing titles:
动手之前,先量一下 Laya 现在的 tokenizer——一个 256,000 词条的 Gemma 式词表,从 mmBERT 继承来的——在中文财经文本上到底切成什么样。两万条公告标题上:
| Measure | Value |
|---|---|
| Filing titles | 23.8 chars → 16.6 tokens |
| Chinese characters per token | 1.15 |
| Pure-CJK tokens in vocab | 21,754 (48.3% multi-character) |
| Pure-digit tokens in vocab | 337 one-digit, 31 two, 5 three, 1 four |
| Nearly one character per token already. Switching to pure character granularity moves 1.15 to 1.00. | |
Here is what that looks like in practice:
净利润同比下降48.7%,营业收入3.05亿元
_ 净 | 利润 | 同比 | 下降 | 4 | 8 | . | 7 | %, | 营业 | 收入 | 3 | . | 0 | 5 | 亿元
Two things settle immediately. Digits are already split one by one — 193.5% becomes 1|9|3|.|5|% and 1234567 costs seven tokens. Whatever advantage a model has at reading financial magnitudes, character granularity cannot add to it, because it is already there. And Chinese is already at 1.15 characters per token, so the change buys 0.15 and costs roughly 15% more tokens for the same sentence.
For a decision model that stuffs its candidate answers into the same sequence as the text, token budget is not an abstraction. It is the thing that breaks first.
实际切出来长这样:
净利润同比下降48.7%,营业收入3.05亿元
_ 净 | 利润 | 同比 | 下降 | 4 | 8 | . | 7 | %, | 营业 | 收入 | 3 | . | 0 | 5 | 亿元
两件事当场就定了。数字本来就是逐位切的——193.5% 是 1|9|3|.|5|%,1234567 要七个 token。模型读财经数字量级的本事不管来自哪,都不会因为改成字级而增加,因为它已经在字级了。汉字也已经是 1.15 字/token,改成纯字级只赚 0.15,代价是同一句话多花约 15% 的 token。
对一个把候选选项和正文塞进同一条序列的决策模型来说,token 预算不是抽象概念,它是最先崩掉的那个东西。
What word-splitting does to 193.5%按词切对 193.5% 做了什么
I built five tokenizers on the same corpus — 2.09 M filing titles plus 766 k newswire items, 86.0 M characters total, nothing borrowed from outside. Characters (4,724 entries, 100% coverage), BPE at 8k and 32k, jieba word segmentation with a 50k vocabulary and [UNK] for anything outside it, and the same word segmentation with unknown words spelled out in characters instead.
我在同一份语料上建了五个 tokenizer——209 万条公告标题加 76.6 万条快讯,合计 8600 万字,不引入任何外部词表。字级(4,724 个,100% 覆盖)、BPE 8k、BPE 32k、jieba 分词 + 五万词表(未登录词 → [UNK])、以及同样的分词但未登录词拆成字。
| Tokenizer | Vocab | Chars / token | UNK | Numbers → UNK |
|---|---|---|---|---|
| char | 4,724 | 1.000 | 0.00% | 0.00% |
| BPE 8k | 8,000 | 1.979 | 0.00% | 0.00% |
| BPE 32k | 32,000 | 2.678 | 0.00% | 0.00% |
| jieba words, UNK | 50,005 | 2.097 | 1.43% | 32.82% |
| jieba words, char fallback | 52,606 | 2.025 | 0.00% | 0.00% |
| The 32.82% is counted over every numeric string in a sample of real newswire headlines, not over constructed examples. | ||||
The failure is specific and total:
jieba words, UNK: 同比 | 增长 | [UNK] <- 193.5%, 19.35% and 1935% are all this token
char fallback: 同比 | 增长 | 1 9 3 . 5 %
char: 同 比 增 长 1 9 3 . 5 %
A word segmenter treats 193.5% as a word. There are unboundedly many numbers, so most of them miss the vocabulary and land on the same [UNK]. Three quantities two orders of magnitude apart become literally indistinguishable — in a corpus of earnings announcements.
I expected this to be the whole story. It is not. Hold onto it; section 07 comes back and prices it.
失败方式非常具体,而且是彻底的:
jieba + UNK: 同比 | 增长 | [UNK] <- 193.5%、19.35%、1935% 全是这一个 token
未登录拆字: 同比 | 增长 | 1 9 3 . 5 %
字级: 同 比 增 长 1 9 3 . 5 %
分词器把 193.5% 当成一个词。数字有无穷多种写法,绝大多数进不了词表,于是全落在同一个 [UNK] 上。三个差着两个数量级的数,在一份业绩公告语料里,变得字面意义上无法区分。
我本以为这就是全部答案。不是。先记住这个数,第 07 节会回来给它标价。
The cleanest cut: statistics, not neural networks最干净的一刀:不用神经网络
Before training anything, the cheapest possible test: TF-IDF features into logistic regression, four tasks, the only variable being whether the features are character n-grams or word n-grams. No pretraining, no architecture, no token budget. If characters win here, it is about the cutting, not about the model.
Regularisation strength was tuned separately for every arm on a held-out slice of the training set — a discipline earned the hard way on an earlier project, where reusing one task’s hyperparameter turned “significantly ahead” into “dead level”.
在训练任何东西之前,先做最便宜的那个测试:TF-IDF 特征 + 逻辑回归,四个任务,唯一变量是特征用字 n-gram 还是词 n-gram。没有预训练、没有架构、没有 token 预算。如果字级在这里就赢,那这件事跟模型无关,就是切法本身。
正则强度对每个臂在训练集内部的留出集上各自调——这条纪律是上一个项目用血换来的:沿用别的任务的超参数,能把”显著领先”变成”完全打平”。
| Task | char 1–4gram | char 1–2gram | word 1gram | word 1–2gram |
|---|---|---|---|---|
| Forum sentiment, 3-way | 0.7056 | 0.7066 | 0.6718 | 0.6850 |
| FinCUGE sentiment, 3-way | 0.7871 | 0.7856 | 0.7589 | 0.7738 |
| FinCUGE news, 15-way | 0.8314 | 0.8258 | 0.8213 | 0.8133 |
| Newswire → overnight gap | 0.6874 | 0.6880 | 0.6813 | 0.6825 |
| Four tasks, four wins for characters, no losses. The margin shrinks as the training set grows: +0.022 on 8.5k samples, +0.006 on 52k. | ||||
Characters win all four. The margin shrinks with sample size, which is what you would expect if word-splitting is a lossy compression that more data can partly undo.
The dimensionality is worth a second look. On forum sentiment, word 1-grams give 4,790 features and character 1–2grams give 18,888. Segmentation is not a neutral reformatting: it throws away distinctions, and it does so by making a decision the model is better placed to make itself.
One check that matters more than the result: the character score on the overnight-gap task, 0.6880, matches the 0.6863 I measured on the same data in an earlier project with an independently written pipeline. The harness is reproducing known work, not measuring something else.
四个任务字级全赢。差距随样本量增大而收窄——如果分词是一种有损压缩、而更多数据能部分补回来,正应该是这个形状。
维度这一栏值得多看一眼。股吧情感上,词 1gram 只有 4,790 维,字 1-2gram 有 18,888 维。分词不是中性的重新排版,它扔掉了区分度,而且是通过替模型做了一个模型自己更擅长做的决定。
有一个检验比结果本身更重要:跳空任务上字级的 0.6880,和我在上一个项目里用另一套独立写的流程在同一份数据上测到的 0.6863 对得上。这说明这套 harness 在复现已知结果,而不是在量别的东西。
The Chinese encoders already chose中文编码器早就选完了
Same corpus, six existing tokenizers, measuring how much Chinese each piece holds and how much of the CJK vocabulary is multi-character:
同一份语料,六个现成的 tokenizer,量每块装多少中文、汉字词表里多字词占多少:
| Model | Vocab | Chars / token | Multi-char CJK |
|---|---|---|---|
| bert-base-chinese | 21,128 | 1.13 | 0.0% |
| chinese-roberta-wwm-ext | 21,128 | 1.09 | 0.0% |
| mBERT multilingual | 119,547 | 1.08 | 0.0% |
| XLM-R base | 250,002 | 1.63 | 54.9% |
| Qwen3-0.6B | 151,669 | 1.59 | byte-level BPE |
| mmBERT-base (Laya's base) | 256,000 | 1.38 | 48.3% |
The two most-used Chinese-specific encoders are 100% single characters — not a single multi-character Chinese token between them. The models that do carry words are the ones serving 100+ languages, where a fixed vocabulary has to be rationed across all of them. That is a budgeting compromise, not a finding about Chinese.
So the direction was never really in doubt. What was in doubt is the size: going from 1.38 characters per token to 1.00, what does it actually buy, and what does it cost?
两个最常用的中文专用编码器是 100% 单字,一个多字汉字 token 都没有。带词的都是服务 100 多种语言的模型,固定大小的词表要在所有语言之间配给。那是预算妥协,不是关于中文的结论。
所以方向其实从来没什么悬念。有悬念的是幅度:从 1.38 字/token 压到 1.00,到底换来多少,又花掉多少?
Five models, one corpus, one variable五个模型,一份语料,一个变量
Every arm gets an identical 6-layer, 384-wide encoder (10.65 M non-embedding parameters, the same in all of them) and a decision head copied from laya.common.build_sequence — instruction, then one [MASK] per candidate answer, then the text, with each option scored from the hidden state at its own mask. Masked-language pretraining on the same 86 M characters. The only thing that differs is the tokenizer.
Each arm processes 150 M tokens. One extra arm, char 2×, gets 300 M — a control I will explain in the next section, and which turned out to be the most valuable thing in the run.
Four downstream tasks, three fine-tuning seeds each, context budget equalised in tokens across arms.
每个臂用完全相同的 6 层、宽 384 的编码器(非 embedding 参数 10.65 M,各臂一致),加一个照搬 laya.common.build_sequence 的决策头——指令,然后每个候选选项前放一个 [MASK],然后是正文,每个选项的得分取它自己那个 mask 位的隐状态。在同一份 8600 万字上做掩码语言预训练。唯一的差异是 tokenizer。
每个臂处理 1.5 亿 token。额外加一个 char 2× 臂,处理 3 亿——这是个对照组,下一节解释,它最后成了整轮实验里最值钱的东西。
四个下游任务,每个三个微调种子,上下文预算按 token 数对齐。
| Arm | Vocab | Params | Forum sentiment | FinCUGE sentiment | News 15-way | Overnight gap (AUC) | Train time |
|---|---|---|---|---|---|---|---|
| char | 4,724 | 12.7 M | 0.6466 | 0.6921 | 0.7643 | 0.6907 | 14 min |
| char 2× (control) | 4,724 | 12.7 M | 0.6510 | 0.7008 | 0.8250 | 0.6922 | 28 min |
| BPE 8k | 8,000 | 13.9 M | 0.6347 | 0.7003 | 0.8209 | 0.6853 | 16 min |
| BPE 32k | 32,000 | 23.1 M | 0.6479 | 0.6941 | 0.7519 | 0.6732 | 31 min |
| words + char fallback | 52,606 | 31.1 M | 0.6385 | 0.7097 | 0.8179 | 0.6879 | 47 min |
| words, UNK | 50,005 | 30.1 M | 0.6000 | 0.6744 | 0.8017 | 0.6876 | 47 min |
| Seed standard deviations: ≤0.0035 on overnight gap, up to ±0.038 on the 15-way task. Only the gap column supports ranking by small margins. | |||||||
On the one column that supports it, granularity orders cleanly, and coarser is worse at every step:
char 2× 0.6922 > char 0.6907 > words+fb 0.6879 ≈ words/UNK 0.6876 > BPE 8k 0.6853 > BPE 32k 0.6732
Best minus worst is +0.019, more than five times the seed spread. And that task is the one that turns on reading financial magnitudes — exactly where granularity should matter if it matters anywhere.
在唯一一列支持这么读的数据上,粒度排得很干净,而且每粗一档就差一点:
char 2× 0.6922 > char 0.6907 > 词+拆字 0.6879 ≈ 词/UNK 0.6876 > BPE 8k 0.6853 > BPE 32k 0.6732
最好减最差是 +0.019,是种子波动的五倍以上。而且这个任务正是靠读财经数字量级的那个——如果粒度在哪里该起作用,就是这里。
The six-point gap that wasn’t about tokenizers那个跟 tokenizer 无关的 6 个点
Partway through the run, one number contradicted everything else. On the 15-way news task, BPE 8k scored 0.8209 against characters’ 0.7643 — a six-point gap, surviving a doubled context budget, so not the truncation effect I had been warning about.
I had two candidate explanations, and no way to pick between them from the numbers in front of me. Either long texts are genuinely harder in character form (187 pieces versus 146), or the word arms had simply seen the corpus more times.
That second one is a hole I dug myself. Equalising tokens processed is not the same as equalising data seen, because the whole point of a coarser tokenizer is that it needs fewer tokens for the same text:
跑到一半,有一个数和其他所有结果相反。在 15 类新闻任务上,BPE 8k 拿到 0.8209,字级只有 0.7643——差了 6 个点,而且把上下文预算翻倍之后差距还在,所以不是我一直在提醒的那个截断问题。
我有两个候选解释,但手上的数字没法区分:要么长文本在字级形态下确实更难(187 块 vs 146 块),要么词级那几个臂只是把语料多读了几遍。
后一个是我自己挖的坑。按 token 对齐不等于按数据对齐——粗粒度 tokenizer 的全部意义就在于同样的文本需要更少的 token:
| Arm | Corpus in tokens | Passes over the corpus |
|---|---|---|
| char | 89.6 M | 1.67 |
| BPE 8k | 44.0 M | 3.41 |
| BPE 32k | 33.5 M | 4.47 |
| words, UNK | 44.6 M | 3.36 |
| char 2× (control) | 89.6 M | 3.35 |
| The control exists to make exactly one comparison: char at 3.35 passes against BPE 8k at 3.41. | ||
The control settles it. Give characters the same number of passes the word arms got, and the six-point gap disappears entirely:
对照组把这件事定死了。给字级和词级臂同样的遍数,6 个点的差距整个消失:
| Task | char | char 2× | Change |
|---|---|---|---|
| News, 15-way | 0.7643 | 0.8250 | +0.061 |
| FinCUGE sentiment | 0.6921 | 0.7008 | +0.009 |
| Forum sentiment | 0.6466 | 0.6510 | +0.004 |
| Overnight gap | 0.6907 | 0.6922 | +0.002 |
| Doubling the training helps the long-text task enormously and the short-text tasks almost not at all. | |||
Words are fine. Blanks are not.”词”没问题,“看不懂”才有问题
Back to the 32.8%. Two arms use the same jieba segmenter and the same 50k vocabulary. The only difference is what happens to a word that is not in it: [UNK], or spelled out in characters.
回到那个 32.8%。有两个臂用同一个 jieba 分词器、同一个五万词表,唯一的区别是词表里没有的词怎么办:变成 [UNK],还是拆成字。
| Task | words → UNK | words → characters | Change |
|---|---|---|---|
| Forum sentiment | 0.6000 | 0.6385 | +0.039 |
| FinCUGE sentiment | 0.6744 | 0.7097 | +0.035 |
| News, 15-way | 0.8017 | 0.8179 | +0.016 |
| Overnight gap | 0.6876 | 0.6879 | +0.000 |
| The fallback version takes first place outright on FinCUGE sentiment — ahead of every character and subword arm. | |||
So “word-level tokenization is bad” needs correcting to “word-level tokenization with an unknown token is bad.” Give it a character fallback and it becomes competitive everywhere and best-in-run on one task.
And now the honest pricing of the headline number. That 32.8% of destroyed numbers, on the task that actually depends on reading earnings figures, is worth 0.0003 — the difference between 0.6876 and 0.6879. Effectively nothing.
Why so little? Because of something I had measured on this same data in an earlier project and then failed to apply: the number-reading advantage lives in the 16% of samples that actually contain an earnings figure. Across all 81,000 events, deleting the numbers barely moves the average. It would still matter on that slice, and it would matter a great deal for a model whose job is to compare magnitudes. But as an argument against word-splitting in general, I was overstating it, and the measurement says so.
The real case against the UNK is not the numbers. It is the 3.9 points it costs on short text, where every token is load-bearing.
所以”词级不行”这句话要改成 “带未登录标记的词级不行”。补上字级兜底,它在所有任务上都有竞争力,在一个任务上还是全场第一。
现在给标题那个数字一个诚实的标价。那 32.8% 被毁掉的数字,在真正依赖读业绩数字的那个任务上,值 0.0003——0.6876 和 0.6879 之差。基本等于零。
为什么这么小?因为一件我在同一份数据上、在上一个项目里已经量过却没拿来用的事实:读数字的优势只存在于真正含业绩数字的那 16% 样本上。摊到全部 8 万条事件里,删掉数字几乎不动平均值。在那 16% 上它仍然重要,对一个工作就是比较量级的模型更是重要。但作为一个反对按词切的通用论据,我把它说大了,实测就是这么说的。
反对 [UNK] 的真正理由不是数字,是它在短文本上要花掉的 3.9 个点——短文本里每一个 token 都是承重的。
What a big vocabulary costs a small model大词表对小模型收多少税
The original training times were contaminated — the queue overlapped with other jobs on the same GPU — so I re-measured every arm alone on an idle machine, 60 steps each after warm-up.
原来那批训练耗时是污染的——队列和别的任务抢同一块 GPU——所以我在空闲机器上把每个臂单独重测了一遍,预热后各跑 60 步。
| Arm | Vocab | s / step | k tokens / s | Output layer's share of forward FLOPs | 150 M tokens |
|---|---|---|---|---|---|
| char | 4,724 | 0.373 | 175.7 | 14.6% | 14.2 min |
| BPE 8k | 8,000 | 0.419 | 156.5 | 22.4% | 16.0 min |
| BPE 32k | 32,000 | 0.807 | 81.2 | 53.6% | 30.8 min |
| words, UNK | 50,005 | 1.240 | 52.9 | 64.3% | 47.3 min |
| words + fallback | 52,606 | 1.220 | 53.7 | 65.5% | 46.5 min |
| At 52k entries, two thirds of the forward pass is scoring the vocabulary rather than understanding the text. | |||||
Every training step scores every vocabulary entry. That cost scales linearly with vocabulary size and does not care how small your model is — so on a model with 10.65 M non-embedding parameters, a 52,000-entry vocabulary spends two thirds of the forward pass on the output projection.
This is what makes the headline claim work. Characters are 3.3× faster per token, so the control arm — characters with double the token budget, which won or tied every task in the run — cost 28 minutes against word-splitting’s 47. Characters need more tokens; each token is much cheaper; the two do not cancel, characters win both.
The effect fades as models grow: put the same 52k vocabulary on a 1B-parameter model and the output layer is a rounding error. But the whole premise here is a small model.
每一个训练步都要给词表里的每一项打分。这笔开销随词表大小线性增长,而且完全不管你的模型多小——所以在一个非 embedding 参数只有 10.65 M 的模型上,五万两千词条的词表会把前向计算的三分之二花在输出层投影上。
这就是标题那个论断成立的原因。字级每 token 快 3.3 倍,所以那个对照臂——字级 + 双倍 token 预算、四个任务全部第一或并列第一——只花了 28 分钟,而按词切要 47 分钟。字级需要更多 token,但每个 token 便宜得多;两者不是抵消,是字级两头都赢。
模型变大这个效应会消失:把同样的五万词表放到一个 1B 参数的模型上,输出层就是个舍入误差。但这整篇文章的前提就是小模型。
What characters cost字级的代价
Same text, 2.7× as many tokens as a 32k vocabulary. In a decision model that is not a footnote, because the candidate answers ride in the same sequence as the text.
On the 15-way news task — filings averaging 107 characters, 230 at the 90th percentile — with an identical token budget across arms:
同一段文本,字级要用 32k 词表 2.7 倍的 token。在决策模型里这不是脚注,因为候选选项和正文挤在同一条序列里。
在 15 类新闻任务上——文本平均 107 字,90 分位 230 字——给各臂完全相同的 token 预算:
| Arm | Mean tokens | Truncated | Head cost of 15 options |
|---|---|---|---|
| char | 174 | 22.6% | 48 tokens |
| BPE 8k | 143 | 7.5% | 34 tokens |
| BPE 32k | 115 | 1.1% | 30 tokens |
| words + fallback | 126 | 1.4% | 31 tokens |
| Doubling the budget to 512 recovered characters from 0.7511 to 0.7643 — real, and smaller than the training-passes effect in section 07. | |||
An earlier project of mine watched this same budget break a decision model outright: with 77 candidate options, the header budget went negative and every option was hard-clipped to a mask plus three tokens, so direct_debit_payment_not_recognised became “direct debit payment” — the negation destroyed before the forward pass. Characters make that ceiling arrive sooner. Plan the sequence length for the tokenizer you picked, not the one you are used to.
我上一个项目里见过这个预算把一个决策模型整个搞崩:77 个候选选项时头部预算变成负数,每个选项被硬截成一个 mask 加三个 token,于是 direct_debit_payment_not_recognised 变成了 “direct debit payment”——否定在前向传播之前就被删掉了。字级会让这个天花板来得更早。按你选的 tokenizer 去规划序列长度,不要按你习惯的那个。
You cannot change granularity afterwards粒度是改不了的
One arm left: take the real Laya, already pretrained on subwords, and force it to character granularity by inserting a space between adjacent characters. Same fine-tuning recipe on both sides.
还剩一个臂:拿真正的 Laya——已经用子词预训练好的——在相邻字符之间插空格,逼它按字切。两边用完全相同的微调配方。
| Task | Native subword | Forced character | Mean tokens |
|---|---|---|---|
| Forum sentiment | 0.7249 | 0.7127 | 94 → 102 (+8.5%) |
| Overnight gap (AUC) | 0.6925 | 0.6939 | 73 → 85 (+15.9%) |
| One loss, one tie, 8–16% more tokens. Neither gap clears the run-to-run noise, so read it as "no difference", not "worse". | |||
The native numbers reproduce earlier measurements on the same data (0.7324 and 0.6961, from a different fine-tuning recipe), so the comparison is honest on both sides.
The more interesting part is that the retrofit does not even cut cleanly:
native: ▁ 净 | 利润 | 同比 | 下降 | 4 8 . 7 % 17.3 tokens / 25 chars
forced: ▁ 净 | ▁利 | ▁ | 润 | ▁同 | ▁比 | ▁下 | ▁降 30.3 tokens / 25 chars (+75%)
A SentencePiece vocabulary stores word-initial variants — ▁利 exists, ▁润 does not. Every character without one costs an extra orphan ▁. The tokenizer is not a layer you can swap; it is baked into what the embeddings mean. If you want a different granularity, you pretrain again.
And for Laya specifically the whole exercise is moot, because section 02 already showed it sits at 1.15 characters per token with digits split individually. There is nothing to convert it to.
原生那两个数和同一份数据上更早的测量对得上(0.7324 和 0.6961,来自另一套微调配方),所以两边的对比是诚实的。
更有意思的是,这种事后改造根本切不干净:
原生: ▁ 净 | 利润 | 同比 | 下降 | 4 8 . 7 % 17.3 token / 25 字
强制: ▁ 净 | ▁利 | ▁ | 润 | ▁同 | ▁比 | ▁下 | ▁降 30.3 token / 25 字 (+75%)
SentencePiece 的词表里存的是”词首变体”——▁利 有,▁润 没有。每个没有词首变体的字都要白赔一个孤立的 ▁。tokenizer 不是一个能换的层,它烧进了 embedding 的含义里。想换粒度,就得重新预训练。
而对 Laya 来说这件事本身就没意义,因为第 02 节已经证明它就在 1.15 字/token、数字逐位切。没有什么可以让它”改成”的。
Characters — and the margin is mostly the bill用字,而且赢的主要是账单
| Dimension | Verdict |
|---|---|
| Quality vs a small subword vocab (8k) | Roughly level. +0.005 AUC on the one trustworthy task |
| Quality vs a large subword vocab (32k) | Characters, +0.019 AUC |
| Quality vs word segmentation + UNK | Characters, +0.047 on short text |
| Cost | Characters, 3.3× faster per token |
| Robustness | Characters. Zero UNK, numbers intact, no coverage question |
| Sequence length | Characters, 2.7× the tokens for the same text |
| Value for an existing Laya | None. It is already there, and it cannot be converted |
Five things I would do, building this from scratch:
- Characters, vocabulary around 5k. 4,724 covered this corpus completely.
- Give characters roughly double the token budget. It is the single largest improvement in the run (+0.061 on long text) and it is still the cheaper option — 28 minutes against 47.
- If you want subword, stay at or below 8k, and always fall back to characters for unknowns. 32k was a pure loss on a small model: −0.019 AUC and double the compute.
- Never use a word segmenter with an unknown token. It is the only arm that came last on two separate tasks, and the fix costs nothing.
- Recompute your sequence length. 2.7× the tokens for the same text, with the candidate options sharing the budget.
The thing I would hand to someone else, though, is not on that list. It is that the most dangerous number in this run was the one that agreed with a clean story — six points for subword on long text, exactly the kind of result that gets written up — and the only reason it did not become a published conclusion is that a control had been queued before the first model started training. The controls you regret are the ones you decided to add later.
如果从零做这件事,我会做五件事:
- 字级,词表约 5,000。 4,724 个字 100% 覆盖了这份语料。
- 给字级大约双倍的 token 预算。 这是整轮实验里最大的单项提升(长文本 +0.061),而且它仍然是更便宜的选择——28 分钟对 47 分钟。
- 要用子词就别超过 8k,而且未登录词一定要回退到字。 32k 在小模型上是纯亏:AUC −0.019,算力翻倍。
- 永远不要用”分词 + 未登录标记”。 它是唯一一个在两个独立任务上垫底的方案,而修复它不花钱。
- 重算序列长度。 同样的文本要 2.7 倍 token,而且候选选项还要分掉这个预算。
不过真正想交给别人的那条不在上面这个清单里。这轮实验里最危险的一个数字,恰恰是那个符合一个干净故事的数字——子词在长文本上领先 6 个点,正是那种会被写成文章的结果——而它没有变成一个发表出去的结论,唯一的原因是那个对照组在第一个模型开始训练之前就已经排进队列了。你会后悔的对照组,都是那些你打算”之后再加”的。