A label that is almost always the same value满屏的五星
I want an app that reads a recipe — just the ingredients and the steps — and predicts whether people will actually like it. The obvious training data is Food.com: 231,637 recipes, 1,132,367 user interactions, eighteen years of home cooks reporting back. The obvious label is the star rating.
The obvious label is broken. 72.1% of all ratings are five stars, another 16.6% are four, and 5.3% of interactions carry no star at all. People rate the recipes they chose to cook, cooked successfully, and came back to praise. A supervised model trained on this learns one lesson — everything is delicious — and it learns it very fast.
But underneath each star is a paragraph, and the paragraphs do not suffer from grade inflation. “Way too salty, I’d cut the butter in half.” “Mine needed 40 minutes, not 25.” “Bland until I doubled the spices.” “Even my picky eater asked for seconds.” The rating tells you that someone was happy; the review tells you why — and “why” is the thing an app should actually predict. So the plan became: first turn a million free-text reviews into structured facts, then teach a model to predict those facts from the recipe alone.
我想做一个 app:给它一份菜谱——只有食材和步骤——它告诉你这道菜大家会不会喜欢。训练数据是现成的,Food.com 上有 231,637 份菜谱、1,132,367 条用户互动,十八年的家庭厨房实战记录。标签似乎也是现成的:星级评分。
可这个标签是坏的。72.1% 的评分是五星,再加 16.6% 的四星,另有 5.3% 的互动干脆没打星。道理不难想:人们只给自己挑中的菜谱做饭,做成了才回来夸。所谓王婆卖瓜,这里是买瓜的人替王婆吆喝。拿这样的数据去训练,模型只会学到一句话——「什么都好吃」——而且学得飞快。
好在每颗星星底下都压着一段话,而这些话没有通货膨胀。「太咸了,黄油下次减半。」「说好 25 分钟,我烤了 40。」「淡得没味,香料加倍才对。」「我家那个挑食的都要了第二碗。」星级只告诉你「满意」,评论告诉你「为什么」——而「为什么」才是一个 app 真正该学会预测的东西。于是计划变成两步:先把一百多万条自由文本的评论读成结构化的事实,再教一个模型只看菜谱、预测这些事实。
A million small confessions, one schema一百万条口供,一张表
Each review goes through a local 35B mixture-of-experts model (Qwen3.6-35B-A3B, served by vLLM on the same desk) with a fixed JSON schema and five worked examples. The schema forces the distinctions that matter later:
- Taste axes — salt, sweet, sour, spice each as too little / fine / too much; “richness” as bland / fine / rich. A field is
nullunless the review actually mentions it: “didn’t mention salt” and “salt was fine” are different facts. - Modifications — every change, split into what the reviewer did versus what they suggest, with the ingredient and the amount. “I used half the sugar and it was perfect” is the single most valuable sentence a recipe can receive.
- made_as_written, would_make_again, texture issues, timing, difficulty, audience (kids, guests, meal-prep…) — each nullable, each with a one-line rule for when it may be asserted.
A cheap regex prefilter skips pure applause (“Yummy, thanks!!”), and grammar-constrained decoding guarantees every output parses. The run covered 1,123,604 of 1,123,812 reviews — 99.98% — with 208 permanent failures, and produced 1,043,300 reviews carrying at least one concrete fact. Nothing was sent to any API; the reviews never left the machine.
具体做法:每条评论过一遍本地的 35B 混合专家模型(Qwen3.6-35B-A3B,vLLM 伺服,就在书桌上这台机器里),配一个固定的 JSON schema 和五个手写示例。这张表强迫模型分清几件后面性命攸关的事:
- 口味轴——咸、甜、酸、辣各分「不够/正好/过头」,另设一个「层次」轴分「寡淡/正常/浓郁」。评论没提的字段一律填
null:「没提咸淡」和「咸淡正好」是两个不同的事实,混为一谈整张表就废了。 - 改动——评论里的每一处改动都要记下来,并且分清是「已经这么做了」还是「建议下次这么做」,连食材带用量。「糖减半,完美」是一份菜谱能收到的最值钱的一句话。
- 另有是否照原样做、愿不愿再做、口感问题、时间准不准、难度、适合谁吃(孩子、待客、便当……)——全部可空,全部配一行「什么情况下才准断言」的规矩。
一层正则预过滤先把纯鼓掌的评论(「Yummy,谢谢楼主!!」)拦在门外,语法约束解码保证每一条输出都能解析。最后跑完 1,123,812 条里的 1,123,604 条,99.98%,永久失败 208 条;其中 1,043,300 条评论至少贡献了一条具体事实。全程没有调用任何外部 API,一条评论都没有出过这台机器。
From 0.9 to 10 requests per second, the hard way从每秒 0.9 条到 10 条,全是学费
A million LLM calls on one small box is a throughput problem, and the tuning turned up things I have not seen documented anywhere.
The prefix cache has a 1,072-token step function. On hybrid GDN/attention models (the Qwen3.5/3.6/3.8 family), vLLM sets its cache block to 1,072 tokens so attention pages can hold the recurrent state. Only whole blocks hit. My shared prompt prefix was ~975 tokens — one token short of the cliff edge, in effect — and the cache hit rate was exactly 0%. Padding the prompt past 1,072 tokens took hit rate to 86% and nearly doubled throughput. The fix for a cold cache was to make the prompt longer.
Speculative decoding made everything worse, twice. MTP looks free until you read the scheduler: with EAGLE-style speculation vLLM silently drops the last matched cache block, so every request re-prefills ≥1,072 tokens — a net loss for short outputs. And n-gram speculation produced corrupted tokens on this hybrid architecture ("would_make_again":null":null — that is a real output). Both went back in the box.
// vllm/v1/core/single_type_kv_cache_manager.py
if use_eagle and computed_blocks[0]:
computed.pop() // your last 1,072 cached tokens, gone per request
Submit in waves, not streams. Filling the server to its 512-sequence ceiling and letting it drain batches prefill together and keeps decode steps pure; a continuous stream mixes them and loses ~30%. Sustained rate on real reviews: 5.7/s, about 490k reviews a day, roughly sixty hours end to end for the corpus.
在一台小机器上发一百万次模型调用,本质是个吞吐工程。调优路上撞见了几件我在任何文档里都没读到过的事,学费不能白交,记在这里。
前缀缓存是一个 1,072 token 的台阶函数。在混合线性注意力架构上(Qwen3.5/3.6/3.8 一族),vLLM 会把缓存块设成 1,072 个 token,好让注意力页装得下递归状态。缓存只认整块。我的共享前缀约 975 个 token——就差那么一步没跨上台阶——命中率因此是精确的零。把提示词垫长、跨过 1,072 之后,命中率到了 86%,吞吐几乎翻倍。给冷缓存开的药方,竟然是把提示词写得更长。
**投机解码帮了两次倒忙。**MTP 看着是白捡的加速,直到你去读调度器源码:开着 EAGLE 类投机时,vLLM 会悄悄丢掉命中的最后一个缓存块,于是每个请求都要白算至少 1,072 个 token 的 prefill——对短输出的任务是净亏损。换 n-gram 投机,这个混合架构直接吐出损坏的 token("would_make_again":null":null,这是真实输出,不是我打错了)。两样都请回了盒子里。
// vllm/v1/core/single_type_kv_cache_manager.py
if use_eagle and computed_blocks[0]:
computed.pop() // 每个请求,最后 1,072 个缓存 token 就这么没了
**按波次投喂,不要流式投喂。**一次把服务器灌满到 512 条的上限、等它排空再灌下一波,prefill 会挤在一起做,解码步保持纯净;细水长流式的持续投喂会把两者搅在一起,白丢三成吞吐。真实评论上的持续速度:每秒 5.7 条,一天约 49 万条,整个语料库跑完差不多六十个小时。
The crowd agrees on what ruins a dish众口其实不难调
Aggregate the labels to recipe level and correlate each aspect rate with the recipe’s (Bayesian-shrunk) mean rating, over the 19,425 recipes with at least ten informative reviews:
| aspect rate方面比例 | r | reading读法 |
|---|---|---|
| would_make_again | +0.63 | the strongest positive signal there is最强的正面信号 |
| called bland被说寡淡 | −0.57 | the strongest complaint by far杀伤力最大的抱怨 |
| soggy / dry / watery / tough湿软/干柴/汤稀/肉老 | −0.32 … −0.19 | every texture failure costs stars口感翻车条条扣分 |
| called hard to make被说费劲 | −0.15 | effort is a tax on ratings麻烦本身就要扣分 |
| too sweet太甜 | +0.02 | in a dessert-heavy corpus, “a bit sweet” is not a complaint在甜品占半壁江山的语料里,「有点甜」不算控诉 |
The most common modifications across the corpus read like a folk theory of American home cooking: substitute the sugar, add garlic, reduce the sugar, substitute the milk, substitute the water, substitute the butter. People fight sugar and reach for garlic. None of this is surprising — which is exactly the point. If the extracted labels recovered obvious culinary truth from a million noisy paragraphs, they are probably carrying real signal.
把标签聚合到菜谱层面,对每个「方面比例」和菜谱的(贝叶斯收缩后的)平均分算相关,样本是 19,425 个有十条以上有效评论的菜谱。结果见上表,这里只说读法。
「愿意再做」是最强的正面信号(+0.63),「被说寡淡」是杀伤力最大的抱怨(−0.57);湿软、干柴、汤太稀、肉太老,每一种口感翻车都明码标价地扣分;连「做起来费劲」本身都要扣 0.15——麻烦是评分上的一种税。最妙的是「太甜」几乎不扣分:在一个甜品占了半壁江山的语料库里,「有点甜」根本不算控诉,算描述。
全库最常见的改动排下来,活脱脱一部美式家常菜的民间药方:换掉糖、加大蒜、减糖、换牛奶、换水、换黄油。众人跟糖搏斗,向大蒜求援。这些结论一点都不出奇——而这正是要点所在:如果从一百万段嘈杂的文字里抽出来的标签,恰好复原了人尽皆知的厨房常识,那它八成装着真信号。
A ridge regression and a well-read amateur一条岭回归,一位读过万卷菜谱的票友
Before training anything expensive, two baselines on a held-out test set of 7,616 recipes (target: Bayesian mean rating; metrics: Spearman, and pairwise accuracy on recipe pairs whose true ratings differ by ≥0.3 stars):
| model模型 | Spearman | pairwise成对准确率 |
|---|---|---|
| Ridge on bag-of-ingredients + tags岭回归·食材+标签词袋 | +0.275 | 71.0% |
| Gradient boosting, numeric features only梯度提升·仅数值特征 | +0.168 | 60.2% |
| Zero-shot LLM judge (35B, 0–100 tastiness)零样本大模型评委(35B,0–100 打分) | +0.125 | 61.5% |
Two things worth staring at. First, the ridge regression’s most negative learned features are a diet in disguise: sugar substitute, fat-free powdered milk, tuna in water, frozen tater tots, cola. The single strongest pattern in eighteen years of reviews is that health-substitutions get punished. Second — and this is the finding that justifies everything downstream — the zero-shot judge, a model that demonstrably knows cooking, is the worst ranker of the three. It scored half of all recipes in the 70–80 band. Knowing what good food is and knowing how 7,616 recipes rank against each other are different skills, and only one of them can be learned from a book.
在动用任何昂贵的东西之前,先在 7,616 个从未见过的测试菜谱上立两条基线(目标是贝叶斯平均分;指标是 Spearman 相关,外加真实分差 ≥0.3 星的菜谱对上的成对准确率)。数字见上表。
有两处值得盯着看。第一,岭回归学到的最负面的特征,摆出来像一张减肥食谱:代糖、脱脂奶粉、水浸金枪鱼、冷冻薯饼、可乐。十八年的评论数据里最强的一条规律竟然是:**用健康替代品换掉本尊的菜谱,会被狠狠扣分。**荤素本无罪,偷梁换柱有罪。
第二件事,是整个后续工作的立身之本:那位零样本大模型评委——一个明明白白懂做菜的模型——是三者中最差的排序者,两千个菜谱有一半被它打在 70 到 80 分之间,好坏不分。所谓纸上得来终觉浅:知道什么是好菜,和知道这 7,616 份菜谱彼此之间谁高谁低,是两种本事;前一种可以从书上学来,后一种只能从数据里练出来。这也正是那一百万条标注的价值所在。
Teaching the 27B, and freezing the machine twice上灶,也当了两回机
The scorer is Qwen3.8-27B with a LoRA (r=32, 176.7M trainable parameters, 0.68%) and a small regression head on the last token: one output for the rating, seven auxiliary outputs for the aspect risks — bland, dry, soggy, tough, watery, would-make-again, hard — each masked out where the aggregate had no data. The aux heads are simultaneously a regulariser and the app’s future explanation channel.
The war stories, in the order they cost me:
- Unified memory bites differently. On a DGX Spark, GPU memory is system memory, and CUDA’s pages are unswappable. An unguarded training run filled all 119GB, the kernel livelocked in reclaim with the OOM killer never firing, and the desktop froze — twice, two hard reboots. The guardrails that ended it:
torch.cuda.set_per_process_memory_fraction(0.75)(which fails loudly instead of eating the host) plus a MemAvailable watchdog. from_pretrainedreturns the model in eval mode, and HF layers skip gradient checkpointing unlessself.trainingis true. One missingmodel.train()was the difference between 58GB and 119GB of activations.- Keep LoRA off the recurrent layers. Adapting the GDN in-projections destabilised training (loss 0.2 → 41 in four steps); attention + MLP only, plus a LayerNorm in front of the head, and the curve went monotone.
After that, 984 optimiser steps at ~89 seconds each — about a day. Validation Spearman climbed 0.134 → 0.236 → 0.301 → 0.335 without a single reversal.
打分器的结构:Qwen3.8-27B 挂一个 LoRA(r=32,可训练参数 1.77 亿,占 0.68%),末位 token 接一个小回归头——一路输出预测评分,七路辅助输出预测各项风险:寡淡、干柴、湿软、肉老、汤稀、愿意再做、费劲。聚合数据里缺哪项,哪项就用掩码跳过。这七路辅助头一物二用:训练时是正则化,上线后是 app 解释「为什么」的出口。
学费清单,按肉疼程度排序:
- **统一内存的咬法不一样。**DGX Spark 上显存就是内存,而 CUDA 的页面不可交换。一次没设防的训练把 119GB 全部吃光,内核在内存回收里活活锁死,OOM killer 全程没开过枪,桌面直接冻住——两次,两次硬重启。最后靠两道护栏了结:
torch.cuda.set_per_process_memory_fraction(0.75)(超了就大声报错,而不是把宿主机吃掉)加一个盯着 MemAvailable 的看门狗。 from_pretrained回来的模型默认是 eval 模式,而 HF 的层只在self.training为真时才走梯度检查点。一行没写的model.train(),就是激活内存 58GB 和 119GB 的区别。- **LoRA 别碰递归层。**把 LoRA 挂到线性注意力的投影上,训练四步 loss 从 0.2 蹿到 41;只挂注意力和 MLP,回归头前面再垫一层 LayerNorm,曲线立刻服帖。
此后就是 984 个优化步、每步约 89 秒——差不多一天。验证集 Spearman 一路 0.134 → 0.236 → 0.301 → 0.335,中途没有回过一次头。
It beats every baseline, and it can say why赢了所有基线,还说得出道理
On the same 7,616 held-out recipes:
| model模型 | Spearman | pairwise成对准确率 |
|---|---|---|
| Qwen3.8-27B LoRA scorer | +0.326 | 75.2% |
| Ridge baseline岭回归基线 | +0.275 | 71.0% |
| Zero-shot judge零样本评委 | +0.125 | 61.5% |
The auxiliary heads generalise too: on test recipes, predicted bland-risk correlates +0.315 with the actual bland-rate, soggy +0.227, dry +0.194. Those numbers are what let the app say “this one runs a bland risk” instead of just emitting a float.
My favourite check is a pair the model has never seen. A garlic-butter herb roast chicken scores 4.76 with bland-risk 0.13. The same request rewritten as a deliberately sad “fat-free microwave alfredo” — fat-free evaporated milk, artificial butter sprinkles, processed cheese slices — scores 4.51 with the bland-risk more than doubled to 0.29. The gap looks small until you remember the whole corpus lives between 4.3 and 4.9: on this ruler, a quarter star is the distance between mid-table and the top. Warm inference is 0.4 seconds on the same desk.
同一份 7,616 个菜谱的考卷上,成绩见上表:Spearman +0.326、成对准确率 75.2%,全面超过岭回归的 +0.275 和 71.0%,把零样本评委远远甩在身后。
辅助头也没白训:在测试集上,预测的「寡淡风险」和真实寡淡比例相关 +0.315,湿软 +0.227,干柴 +0.194。有了这几个数,app 才有资格说一句「这道菜有寡淡的风险」,而不是只吐出一个光秃秃的分数。
我最喜欢的一组检验,是两份模型从未见过的菜谱。一份正经的蒜香黄油烤鸡:4.76 分,寡淡风险 0.13。另一份是故意写来找骂的「脱脂微波炉阿尔弗雷多面」——脱脂淡奶、人造黄油粉、再制芝士片——4.51 分,寡淡风险翻了一倍还多,0.29。分差看着不大?别忘了整个语料库都挤在 4.3 到 4.9 之间:在这把尺子上,四分之一颗星就是中游和头部的距离。顺带一句,热推理只要 0.4 秒,就在同一张书桌上。
Grading my own labels, ruthlessly给自己的标签打分,不留情面
Every conclusion above stands on machine-written labels, so I hand-checked 283 of them against the original reviews, field by field. The honest table:
| field字段 | precision精确率 | verdict判词 |
|---|---|---|
| informative / difficulty难度 / made_as_written | 93–97% | trust可信 |
| bland寡淡 / modifications改动 | 90–92% | trust — conveniently, the two most valuable fields可信——恰好是最值钱的两个字段 |
| would_make_again / audience受众 | ~70% | over-asserted on gushing five-star prose在狂夸型五星评论上被高估 |
| “rich”「浓郁」 | ~50% | actually a praise-intensity proxy实为「夸奖强度」的代理变量 |
| texture issues口感问题 | ~36% | praise gets inverted (“so moist” → dry), categories blur夸奖被反转(「特别润」标成干)、类目互串 |
The failure modes are systematic, not random: the extractor turns enthusiasm into “rich”, flips texture praise into texture complaints, and occasionally reverses the direction of timing feedback. Aggregate correlations survive this noise — bland at −0.57 is built on a 92%-precision field — but any per-recipe explanation shown to a user should be drawn from the trustworthy rows only. Auditing your own pipeline is cheaper than shipping its mistakes.
上面的每一条结论都站在机器写的标签上,所以我回头人工核对了其中 283 条,逐字段对着评论原文打分。诚实的账目见上表。
错得并不随机,而是成体系:抽取器会把「热情洋溢」误读成「浓郁」,会把口感上的夸奖反转成口感投诉(「特别润」标成了干柴),偶尔还把时间反馈的方向搞反(「烤太久了」记成「时间不够」)。聚合层面的相关性扛得住这些噪声——寡淡的 −0.57 建立在一个精确率 92% 的字段上——但凡是要展示给用户看的单条解释,只能取表里「可信」的那几行。给自己的流水线挑错,总比把错误发给用户便宜。
顺带一个教训:本想拿一个公开的中文菜谱数据集做迁移,下载后发现它 demo 版的评分是占位随机数——均匀分布在 3 到 5 之间,和收藏数相关性为零。任何模型都学不了随机标签,两条基线在它上面双双得零分,反倒证明了评测流程是诚实的。数据先体检,再开工。
The second language came almost free第二种语言几乎是白送的
The scorer was trained on English recipes only. To test whether it survives a language switch, I machine-translated 1,000 held-out test recipes into Chinese with the local 35B — real labels, same recipes, different script — and scored both versions:
| input输入 | Spearman vs. truth与真实评分的 Spearman |
|---|---|
| English original英文原文 | +0.338 |
| Chinese translation中文译文 | +0.271 |
| EN–ZH score agreement中英打分一致性 | 0.826 |
Zero-shot, the Chinese input keeps 80% of the ranking power and lands exactly on the English ridge baseline. The model scores the same recipe almost identically in either language (agreement 0.826) — the taste judgement seems to live below the language layer, which is what you would hope for from a bilingual backbone. Proper Chinese fine-tuning waits for a Chinese corpus with honest labels; until then, the app takes 红烧肉 as readily as pot roast.
打分器只见过英文菜谱。为了试试它换一种文字还认不认得味道,我把 1,000 个测试菜谱用本地 35B 翻成中文——标签是真的、菜谱是同一批、只是换了文字——然后两个版本各打一遍分。结果见上表。
零样本之下,中文输入保住了八成的排序能力,落点恰好压在英文岭回归基线上;同一份菜谱,中英文两个版本的得分一致性高达 0.826。看起来「对味道的判断」这件事,藏在比语言更深的一层——这正是选双语底座时盼着的事。地道的中文微调,要等一份标签诚实的中文语料;在那之前,这个 app 递给它一份红烧肉,它接得住,就像接得住 pot roast 一样。
Three things I would keep留下三句话
When the label is saturated, mine the text next to it. The stars carried almost nothing; the paragraphs carried everything. Structured extraction with nullable fields and a did/suggests split turned praise-inflated reviews into a training corpus, and the aggregate statistics validated themselves against common sense before anything got trained.
Zero-shot knowledge is not ranking ability. The judge that “knows cooking” lost to a ridge regression from 1982. If your task is ordering things, nothing substitutes for supervision on the actual order — the trained 27B beat the same backbone’s zero-shot cousin by 20 Spearman points.
Audit your own pipeline before you believe it. The label audit found two fields that looked fine and weren’t; the Chinese dataset audit found labels that were literally random. Between them they cost half a day, and either one, unaudited, would have quietly poisoned everything downstream. Sixty hours of extraction, a day of training, half a day of honesty — that was the whole budget, on one machine, with nothing leaving it.
**标签饱和的时候,去挖它旁边的文字。**星星里几乎什么都没有,段落里什么都有。可空字段加「做了/建议」之分的结构化抽取,把一堆通胀的夸奖变成了一份能训练的语料;而聚合统计先对了一遍常识的账,才轮到训练开工。
**零样本的学问不等于排序的本事。**那位「懂做菜」的评委,输给了一条上个世纪就有的岭回归。任务如果是把东西排出个先后,就没有什么能替代「在真实顺序上受过监督」这件事——同一个底座,训练过的比零样本的高出二十个 Spearman 点。
**先审自己的流水线,再相信它。**标签自查揪出了两个看着体面、实则不可信的字段;中文数据体检发现标签干脆是随机数。两次审计加起来花了半天,而其中任何一个漏网,都会不声不响地毒害下游的一切。六十个小时的抽取,一天的训练,半天的诚实——这就是全部预算,在一台机器上,什么都没有离开它。