Word-splitting deletes a third of the numbers. That is not why it loses. 按词切会删掉三分之一的数字。但它输不是输在这里
Five Chinese tokenizers over one 86-million-character corpus, each feeding an identical small decision model. Characters win — on price, not accuracy — and the single result that pointed the other way turned out to be a confound I built myself. 同一份 8600 万字的 A 股语料,五种中文切分方案,各喂一个完全相同的小决策模型。字级赢了,但赢在成本不在准确率;而唯一一个看起来相反的结果,是我自己埋的坑。
Two fighters learned to flank. Ten fighters learned nothing. 两个人学会了包夹,十个人什么都没学到
Sword-and-shield humanoids in a GPU physics engine, trained by competitive self-play from 1 v 1 up to 5 v 5, to see whether teamwork appears on its own. It did, exactly once, in the one place where nobody should be surprised. 在 DGX Spark 上用竞争式自博弈训练物理仿真的剑盾人形,从 1 打 1 一路推到 5 打 5,看会不会自己长出没人写过的配合。配合只出现了一次,而且出现在最不该让人意外的地方。
The video model didn't fit. The text encoder was the problem. 装不下的不是视频模型,是文本编码器
A 33B video model shipped with a 32B text encoder, on a box with 119 GB of unified memory and no OOM killer to catch you. One device flag moved 27 GB off the peak for free. MiniMax H3 是个 33B 的视频模型,还配了个 32B 的文本编码器。在只有 119 GB 统一内存、且没有 OOM killer 兜底的机器上,一个 device 参数把峰值削掉了 27 GB,而且不要钱。
Five stars all the way down: what 1.1 million recipe reviews know that their ratings don't 评分全是五星,等于没有评分——好在评论会说人话
72% of the ratings are five stars. The reviews underneath them are where the truth lives: too salty, went dry, kids devoured it, halve the salt next time. One local machine read all 1.12 million of them into structured facts, and a 27B model trained on those facts now reads a recipe and tells you whether people will like it — and roughly why. Food.com 上七成二的评分是满分五星,这样的标签基本没有信息量。可是评论正文里藏着真话:太咸了、烤干了、孩子抢着吃、下次盐减半。让一个本地模型把一百一十二万条评论逐条读成结构化的表格,再拿这张表教一个 27B 模型看菜谱打分——它赢过了所有基线,包括一个懂做菜但不懂排序的零样本大模型。
The formula was real. My ruler was wrong. 公式是真的,不对的是我那把尺子
《相声的有限元》 contains a real formula for how long an audience will laugh, and a rather elegant one: its three optimality criteria compose so that the quality coefficient becomes exactly 1. Its output is seconds of laughter from a live audience; I fed it silent written jokes. Every claim came back null — and then I discovered my measuring instrument had a discrimination of 1.09 over chance, rebuilt it, and watched the answer move — into a place that is worse for the book, not better. 《相声的有限元》里有一个真实存在、而且相当自洽的公式,用来预测观众会笑多久——它的质量系数在三条判据同时满足时,会干干净净地等于 1。这个公式的输出单位是「现场观众笑了几秒」,而我拿 4046 条没有舞台、没有观众的书面笑话去量它,四条主张一条没过。后来我发现自己那把尺子的区分度只有 1.09(随机是 1.00),于是重造了一把——答案跟着变了。
A model that stopped reading in 1930 says today's Democrats are the Republican party 一个 1930 年就停止阅读的模型说:今天的民主党,是共和党
Take a model that has read nothing written after 1930 and use it as an instrument. Feed it only conduct — who votes for a party, what it does with tariffs, land, unions, churches, foreigners — with every name and every ideological label stripped out, and ask it who this is. It passes its own 1930 controls 32 times out of 34. Then it identifies today's Democratic coalition as the Republican party. 找一个只读过 1931 年之前文本的模型,拿它当一把尺子来量。只告诉它一个机构在做什么事,不给名字,不给年份,不给任何主义标签,让它猜这是谁。它先做对了 34 道 1930 年的常识题里的 32 道。然后它把今天的民主党纲领判成了共和党,把今天的共和党纲领判成了禁酒党。
Your local model is slow. Your machine isn't — you're using one lane of a 128-lane road. 你的本地模型不慢,是你只开了一条道——一条 128 车道的路
Single-stream generation on a DGX Spark is bandwidth-bound and unfixable. Everything else about it is width, and width is nearly free. Here is what that buys, measured — including the two places it buys nothing. DGX Spark 的单流速度卡在带宽上,改不了。但它剩下的全是宽度,而宽度基本白送。这篇记录我拿宽度换到了什么,也包括换不到东西的那两个地方。
The DGX Spark isn't slow. It's bandwidth-starved — and speculative decoding helps. DGX Spark 不慢,是带宽饿着了——而投机解码能喂饱它
The DGX Spark isn't slow. It's bandwidth-starved, and that's fixable. Five speculative decoding strategies, eight tasks, one binary. DGX Spark 不是慢,是带宽不够,而这个能治。五种投机解码策略、八项任务、一个二进制。
Britannica drew a map of all human knowledge in 1974. An LLM is what it was trying to be. 1974 年,大英百科画了一张人类知识的全图。它想成为的东西,今天叫大语言模型。
In 1974 Britannica bet its 15th edition on a philosopher's outline of all human knowledge. Libraries never used it. Fifty years later, a neural network learned the same map from data — and lost the citations. 1974 年,大英百科把第 15 版押在一位哲学家画的知识全图上。图书馆从来没用过它。五十年后,一个神经网络从数据里把同一张图学了出来——却把出处弄丢了。