One division一个除法
A GB10 pairs roughly a petaflop of FP4 compute with 273 GB/s of memory bandwidth. Decoding one token requires reading every weight exactly once, so single-stream speed is not a benchmark result — it is arithmetic:
这台机器的算力是过剩的。GB10 芯片有将近一个 PFLOPS 的 FP4 算力,可它的内存带宽只有 273 GB/s。偏偏模型每吐出一个字,都得把全部权重从头到尾读一遍。
于是单流速度这个数,根本不用测,一道除法就出来了:
273 GB/s ÷ 28.5 GB (FP8 weights) = 9.6 tok/s ceiling
measured, one request .............. 7.4 tok/s
No configuration flag fixes arithmetic. During single-stream decode a GB10 is a memory controller with a very expensive ornament attached — the compute sits idle, waiting.
Batching is how you spend the ornament. Sixteen concurrent requests read the weights once between them, so the marginal cost of request seventeen is arithmetic the machine has spare. That asymmetry — latency is fixed, throughput is nearly free — is the whole subject of this piece. Everything below is an attempt to convert idle width into something useful, and an honest account of where that conversion fails.
这个数没法讲价。你调任何参数,它都在那儿。
一个请求慢慢跑的时候,这台机器的样子挺荒诞:内存总线忙得不可开交,旁边那块昂贵的算力一直闲着,就那么等着。
想把它用起来,只有一个办法——同时让十六个请求进来。它们共用同一次权重读取,等于把最贵的那笔开销摊薄成十六份。到第十七个请求的时候,多花的其实全是机器本来就闲着的算力。
所以这台机器的性格就一句话:一个请求慢,这事没救;但让很多请求一起跑,几乎不要钱。 接下来的内容,都是我在琢磨怎么把这份白送的宽度换成有用的东西——顺便也记下了它在哪些地方换不动。
Four lanes open out of 128128 条车道只开了 4 条
The container that had been serving my model for two days looked reasonable:
这台机器已经给我跑了两天模型。启动命令我当时看了好几遍,没觉得有什么问题:
llama-server -hf unsloth/Qwen3.8-27B-GGUF:UD-Q4_K_XL -ngl 999 -c 32768
No -np flag. llama.cpp defaulted to n_slots = 4. Aggregate throughput flat-lined near 34 tok/s no matter how many requests arrived — everything past four simply queued. Concurrency 16 was slower per request than concurrency 4 and produced no more total tokens.
That is a one-word fix worth 8×, and I suspect it is the single most common reason a Spark gets written off as slow.
问题在于少了个 -np。没有它,llama.cpp 默认只开四个槽位。
所以真相是:我以为自己在压测这台机器,其实第五个请求开始就在门口排队了。总吞吐死死钉在 34 tok/s,我丢多少进去都一样。更让人困惑的是,开到 16 并发时每个请求还变慢了,总产出却一点没多——现在回头看,那就是排队的样子。
加一个参数,八倍。我怀疑网上不少”DGX Spark 好慢”的结论,就是这么来的。
They compose它们可以叠加
| aggregate tok/s聚合 tok/s | c=1 | c=16 | c=32 | c=64 | c=128 |
|---|---|---|---|---|---|
| llama.cpp, stock 4 slotsllama.cpp 默认 4 槽 | 10.3 | ~34 | ~34 ↓ | — | — |
| llama.cpp, 16 slots + draft modelllama.cpp 16 槽 + 草稿模型 | 24.2 | 86.4 | 76.5 ↓ | — | — |
| vLLM FP8 | 7.4 | 101.7 | 153.8 | 259.4 | 343.6 |
| vLLM FP8 + MTP | 14.4 | 164.8 | 253.3 | 351.6 | 368.7 |
| + quantised KV cache+ 量化 KV 缓存 | 16.2 | 149.9 | — | 463.5 | — |
Four things independently buy speed, and none substitutes for another.
Batching buys aggregate throughput and nothing else. Speculative decoding buys single-stream latency, which batching structurally cannot: llama.cpp’s draft model took one request from 10.3 to 24.2 tok/s. vLLM’s built-in MTP head (it ships in the official FP8 repo; --speculative-config '{"method":"mtp","num_speculative_tokens":2}') was the single best flag I found — 1.6–1.9× across the entire concurrency range, and unlike a draft model it does not fade as the batch fills.
Model file size is a latency knob, not just a memory one. Because decode speed is weights ÷ bandwidth, the 17.5 GB Q4 GGUF beats the 28.5 GB FP8 checkpoint on single-stream speed (24.2 vs 14.4 tok/s) even though vLLM crushes llama.cpp on everything else.
Quantising the KV cache turned out to be a throughput optimisation. This one surprised me. The pool was 182k tokens; I assumed that was the hardware. It was two defaults: attention KV stored in bf16, and — less obvious — this hybrid SSM model’s Mamba state silently sitting in float32. Setting --kv-cache-dtype fp8 --mamba-cache-dtype float16 took the pool to 522k tokens, 2.9×. GSM8K held at 92.0% (bf16 baseline 91.0%), needle recall stayed perfect. And it got faster: +13% single-stream, +32% aggregate. On a bandwidth-bound machine, a smaller KV cache means fewer bytes read per attention step. The vLLM DGX Spark write-up warns of an fp8-KV performance cost for some workloads; for this hybrid model the opposite held.
能提速的东西有四样,各管一摊,谁也顶替不了谁。
先说批处理。它只涨总吞吐,对单个请求的快慢一点忙都帮不上——这是它的本分,也是它的局限。
想让单个请求变快,得靠投机解码。llama.cpp 那边挂一个草稿模型,一个请求就从 10.3 跳到 24.2 tok/s。vLLM 更省事,官方 FP8 仓库里自带一个 MTP 头,命令行加一句 --speculative-config '{"method":"mtp","num_speculative_tokens":2}' 就完事了。这大概是我这几天找到的最划算的一个参数:从低并发到高并发,稳定给 1.6 到 1.9 倍,而且不像草稿模型那样,批次一满就没脾气了。
第三样容易被忽略:模型文件的大小本身就是个延迟旋钮,不只是占多少内存的问题。既然解码速度 = 权重 ÷ 带宽,那 17.5 GB 的 Q4 GGUF 在单流上就是能赢 28.5 GB 的 FP8,24.2 对 14.4——哪怕 vLLM 在别的所有方面都把 llama.cpp 按在地上摩擦。
第四样是我完全没料到的。我一直以为 KV 池就 182k token 是硬件定死的,后来才发现是两个默认值在偷偷占地方:注意力的 KV 存成了 bf16,还有个更隐蔽的——这是个混合 SSM 模型,它的 Mamba 状态默默用着 float32。
我加了 --kv-cache-dtype fp8 --mamba-cache-dtype float16,池子当场涨到 522k,2.9 倍。质量呢?GSM8K 从 91.0% 到 92.0%,针刺测试全对,没掉。真正的意外是速度:单流 +13%,总吞吐 +32%。
不过转念一想也合理:这台机器卡在带宽上,KV 缓存小了,每一步注意力要读的字节就少了。有意思的是 vLLM 官方那篇 DGX Spark 的文章还专门提醒过 fp8-KV 在某些场景会掉性能——在这个混合模型上,恰好反过来。
Redundancy is nearly free, so buy accuracy with it冗余几乎免费,那就用它买准确率
If six attempts cost barely more wall clock than one, a 27B model can behave like a much larger one — provided something can pick the right attempt. I ran the same A/B on five benchmarks: one sample versus six parallel samples plus a test-driven repair round.
既然试六次和试一次花的时间差不多,那一个 27B 就有机会干出比自己大得多的模型的活。
前提是,得有人能从这六个答案里把对的那个挑出来。
我在五个基准上做了同一组对照:一边只跑一次,另一边并行跑六次、再加一轮照着测试结果去修。
| benchmark | single单发 | parallel + verification并行 + 验证 | lift提升 |
|---|---|---|---|
| hand-built gauntlet (5 tasks)自制五关(5 题) | 1/5 | 4/5 | 4.0× |
| LiveCodeBench-hard (10) | 1/10 | 4/10 | 4.0× |
| Aider polyglot-py (34) | 26.5% | 44.1% | 1.7× |
| AIME 2025 (30) | 40.0% | 46.7% | 1.2× |
| GSM8K (100) | 91.0% | 94.0% | 1.03× |
The lift shrinks monotonically as the baseline rises, and that is the honest shape of the result: parallel verification converts idle throughput into correctness exactly where correctness is scarce, and does nothing where the model is already right.
Two mechanisms deserve separating, because conflating them is the most common mistake I made:
The verifier is the active ingredient, not the sampling. On the gauntlet, every win came from a sample that already passed — diversity found a correct candidate, but only hidden tests could select it. Where no objective verifier exists and you fall back to majority voting, the gain collapses to a few points (GSM8K +3, AIME +6.7). Worse, voting amplifies systematic bias instead of cancelling it: five parallel samples of an arithmetic word problem cost 1.7 s against 0.9 s for one call — nearly free — and the vote returned the wrong answer at 60% agreement while the single call was right.
Sampling rescues variance, never capability. On LiveCodeBench-hard, four problems scored 0/30 across all nine candidates. Six independent samples, three repairs shown the exact failing cases, and later a run with thinking enabled and six times the token budget — all identical. When the model does not know the algorithm exists, no amount of width finds it. The same lesson arrived in miniature on a glob-matching task where nine of nine candidates failed the same two tests in the same way.
这张表最该看的不是哪个数字最大,而是它的形状:基线越高,提升越小,一路单调下去。
翻译成人话就是——这招是拿闲置的算力去买正确率,而且只在正确率本来就不够的地方买得到。模型本来就会做的题,你花再多算力也涨不了一分。
这里有两件事必须分开说,因为把它们混为一谈是我这几天犯得最多的错。
第一,真正干活的是验证器,不是采样。 我自己出的那五道题,每次成功都是因为六个样本里本来就躺着一个对的。多样性负责把对的答案生出来,但把它认出来的是隐藏测试。要是没有客观的验证手段、只能靠投票,收益立刻塌到几分——GSM8K 涨 3 分,AIME 涨 6.7 分。
更糟的是,投票有时候是帮倒忙的。有道算术应用题,我跑五路并行花了 1.7 秒(单跑 0.9 秒,基本白送),结果投票以 60% 的一致率选中了错的答案,而单跑那次是对的。系统性的错误,投票不会抵消,只会盖章。
第二,采样捞得回运气,捞不回能力。 LiveCodeBench 的难题里有四道,九个候选全军覆没,0/30。我试过六次独立采样、三次把具体错在哪贴给它看再修、后来还开着思考模式给了六倍 token 重跑——结果一模一样。
模型不知道那个算法存在的时候,你试一万次也试不出来。还有个更小的例子更能说明问题:一道 glob 匹配题,九个候选里九个都用完全相同的方式,错在完全相同的两个测试上。
Serial tokens, and nothing else串行 token,仅此而已
Once batching made width cheap, one cost remained, and it explains every subsequent result — positive and negative.
A real map-reduce agent job (summarise source files, then synthesise one note) sped up 3.7× in the map phase and only 2.0× end to end, because the single serial reduce call went from a quarter of the runtime to more than half of it. Amdahl’s law, arriving on schedule.
So I tried to parallelise the reduce with a merge tree: 48 summaries → 6 parallel partial merges → 1 final. The reduce phase got worse, 24 s → 62 s. The reason is exact and worth internalising: the expensive part of a reduce is not reading its inputs — prefill is nearly free now — it is writing its output, token by token. The tree did not shorten the writing; it added an entire extra round of it.
宽度便宜了之后,就只剩一样东西还贵。而这一样,能解释我后面遇到的所有事——成功的和失败的都算。
那次是个真实的任务:先总结一堆源文件,再汇总成一份笔记。总结那部分快了 3.7 倍,可整件事从头到尾只快了 2.0 倍。差在哪?差在最后那次汇总调用——它原来只占总时间的四分之一,现在占了一半还多。
阿姆达尔定律,准时报到。
我不服气,想把汇总也并行掉:48 份摘要先分成 6 组各自合并,再汇成一份。结果汇总阶段从 24 秒变成了 62 秒,更慢了。
想明白之后觉得挺显然的:汇总慢,不是慢在读那些摘要上——prefill 现在便宜得很——而是慢在一个字一个字把结果写出来。 我那棵归并树一个字都没帮它少写,反倒多加了一整轮写。
The rule. Wall clock on this machine is governed by serially generated tokens, and by nothing else. Optimisations that shorten the serial token path (tighter output schemas, fewer pipeline levels, terser final answers) win. Optimisations that add serial generation in order to save parallel work lose.
一句话规律。这台机器上,花掉的时间只取决于一件事:有多少 token 是必须一个接一个生成的。凡是能缩短这条串行链的(输出格式更紧、流水线层数更少、最终答案更短),都赢;凡是为了并行而多加一轮生成的,都输。
3.8× on a real job真实任务上的 3.8 倍
Sampling parallelism multiplies redundant work. The thing actually worth building multiplies distinct work: one planner call splits a long job into independent subtasks, every subtask becomes a tool-using worker, all workers ride the same batch, and one integration call closes it out.
The job: document 23 real Python modules — one notes file per module plus an overview. Identical plan in both arms; the only difference is a semaphore.
前面讲的那种并行,本质是把同一件事做很多遍,赌里面有一遍是对的。
真正值得造的是另一种:把不同的事同时做。先让模型规划一次,把一个大活切成互不相干的小活,每个小活派一个会用工具的 worker 去干,所有 worker 挤在同一批里跑,最后再来一次调用把结果收拢。
我拿真活试了:给 23 个 Python 模块写文档,一个模块一份笔记,外加一份总览。两组用的是同一份计划、同一个模型、同样的循环,唯一的区别是那个信号量开到 1 还是开到 24。
| serial (1 worker at a time)串行(一次一个 worker) | parallel (24 slots)并行(24 槽) | |
|---|---|---|
| wall clock墙钟 | 1185.5 s | 315.2 s |
| execute phase执行阶段 | 1077.6 s | 209.6 s (5.1×) |
| plan / integrate规划 / 整合 | 33.5 / 74.4 s | 33.5 / 72.1 s |
| workers dropped被丢弃的 worker | 1/8 | 0/8 |
| files delivered交付文件 | 19 + overview | 21 + overview |
3.8× end to end on a real deliverable, with more complete output than the serial run. The remaining floor is the dependent tail — plan plus integrate, about 105 s — exactly as the serial-token law predicts.
Getting there took five iterations, and the failures are more transferable than the number. Each is a trap any decomposition agent on a local model will hit:
- Ground the planner in the real file listing. Given only top-level directories, it invented twelve plausible module names that did not exist, and honest workers spent eight minutes proving files absent.
- Cap subtask scope (≤3 files) rather than trusting “natural” splits. Coarse plans made every worker overrun its step budget. Width is free on this machine — err wide.
- Check your tool semantics before blaming the model.
rglobon a file path silently returns nothing, so every file-scoped grep answered “(no matches)” and workers spiralled retrying patterns. One-line fix. - Break loops in the harness, not the prompt. Deterministic tools at low temperature produce identical-call spirals — the same 30k-character file read twelve times. The harness now detects a repeated identical call and answers “the result has not changed; proceed.”
- Force write-as-you-go, and validate deliverables rather than prose. Workers research indefinitely and postpone output past their budget. A strict one-read-one-write procedure plus a half-budget “you have written NOTHING” nudge fixed it; validation now counts files written instead of judging the closing summary.
端到端 3.8 倍,而且并行那版交出来的东西还更全。 剩下那 105 秒是甩不掉的——规划加汇总,两头都是串行的,正好印证前面那条规律。
不过这个数字是我改到第五版才拿到的。前四次的失败比这个数字有意思得多,而且我觉得凡是在本地模型上做分解式 agent 的人,早晚都得挨这几下:
一、规划的时候,一定要把真实的文件清单摆在它面前。 我第一版只给了顶层目录名,它当场编了十二个听起来特别合理、实际上根本不存在的模块名出来。然后一群老实巴交的 worker 花了八分钟,认认真真地证明”这文件真没有”。
二、子任务粒度必须卡死(我卡在 3 个文件以内),别指望它”自然拆分”。 拆太粗的话,每个 worker 都会撞上步数上限。反正这台机器上宽度不要钱,宁可拆碎一点。
三、先怀疑自己的工具,再怀疑模型。 有一版 worker 疯狂换正则重试,我以为是模型犯傻,查了半天发现是 rglob 用在文件路径上会一声不吭地返回空——所有针对单个文件的 grep 都在回”没找到”。改一行代码的事。
四、死循环得在 harness 里断,指望 prompt 没用。 工具是确定性的,温度又低,模型很容易卡在同一个调用上出不来——我见过同一个三万字的文件被读了十二遍。现在 harness 一发现重复调用就直接顶回去:“结果没变,往下走。”
五、逼它边干边写,验收只看产出不看文笔。 worker 特别喜欢无限地”再研究研究”,把写东西一路拖到预算耗尽。我改成”读一个、立刻写一个”的死规矩,再加一句预算过半时的提醒——“你到现在一个字都还没写”——就治好了。验收标准也跟着改成数它写出来几个文件,而不是看它结尾那段总结写得漂不漂亮。
Two negative results, one shape两个负面结果,同一个形状
I tried the same idea on tasks that do not decompose, twice, in two different domains. Both lost, in the same way.
Skeleton-parallel code generation. For a single-file coding exercise: one call plans the interface, every function body is generated concurrently against that contract, assembly is programmatic. Result: 73 s versus 41 s per task, and the pass rate halved (20.6% → 11.8%). Two causes, both predictable in hindsight. The skeleton adds a serial round, and these solutions are only 600–1500 tokens — below the crossover where parallel expansion can repay a planning round. And independently-written function bodies drift apart on shared helpers, so the assembled file breaks at the seams.
Parallel reconnaissance for terminal tasks. For Terminal-Bench — real terminals, real containers, “fix this broken git repo” — I added a concurrent recon phase: three scouts plan read-only probes in parallel, findings are prepended to the agent’s first prompt. Everything downstream untouched, so any delta is attributable to recon alone.
同一个思路,我又在两个完全不同的地方各试了一次——这两次的任务都是拆不开的。两次都输了,而且输的方式一模一样。
第一次是把写代码这件事拆开。 单文件编程题,先让模型规划出接口骨架,各个函数体照着这份契约并发生成,最后程序拼起来。结果每题 73 秒,对比单跑的 41 秒;通过率还腰斩了,20.6% 掉到 11.8%。
回头看,两个原因都挺显然的。一是骨架又多加了一个串行轮次,而这些题的解统共才 600 到 1500 个 token——太短了,并行省下来的那点时间根本抵不上多出来的这一轮。二是各写各的函数体,大家对共享 helper 的想象会慢慢跑偏,拼起来就在接缝处裂开。
第二次是给终端任务加并行侦察。 Terminal-Bench 那种真实终端、真实容器、“把这个坏掉的 git 仓库修好”的活,我在正式开工前加了一轮并发侦察:三个 scout 同时规划各自的只读探查命令,结果塞进 agent 的第一个 prompt 里。后面的流程一个字没改,所以出什么差别都算在侦察头上。
| task | baseline基线 | recon (raw)侦察(原始) | recon (digested)侦察(结论化) |
|---|---|---|---|
| fix-git | PASS 253 s | PASS 371 s | PASS 249 s |
| sqlite-with-gcov | PASS 359 s | timeout | PASS |
| crack-7z-hash | PASS 296 s | PASS 481 s | timeout |
| three others其余三个 | timeout | timeout | timeout |
| total合计 | 3/6 | 2/6 | 2/6 |
The recon phase itself costs 8.5 s. The damage is downstream: its findings enlarge the context of every subsequent turn. All three passing tasks ran 47–62% longer, and sqlite-with-gcov — which the baseline finished in 359 s — was pushed past the 660 s deadline into failure.
So I tested the mechanism directly: compress each angle’s raw output into ≤4 factual bullets, cap the injection at 2,000 characters instead of 14,000. The mechanism was confirmed and the benefit still did not appear. Compression removed the slowdown exactly where predicted — fix-git returned to 249 s, sqlite-with-gcov was recovered — but crack-7z-hash, a brute-force task where environment facts are irrelevant, lost the recon time for nothing and timed out. Two fixed, one broken, net zero.
侦察本身只花 8.5 秒,一点都不慢。问题出在它之后:查到的那些东西会一直挂在后面每一轮的上下文里,谁都甩不掉。
三个本来能做完的任务全都慢了 47% 到 62%。其中 sqlite-with-gcov 最冤——它原本 359 秒就能收工,硬是被拖过了 660 秒的死线,判了个失败。
既然怀疑是上下文太重,那就直接试。我把每个角度的原始输出压成不超过 4 条干货,注入量从 14,000 字砍到 2,000。
猜想被证实了,收益还是没出现。 压缩确实精准地把拖慢消掉了:fix-git 回到 249 秒,sqlite-with-gcov 也救回来了。可 crack-7z-hash 那种纯爆破的任务,环境信息对它一点用都没有,白搭进去的侦察时间反而把它推超时了。
修好俩,赔进去一个。白忙活。
The decision rule, from four task shapes. Switch a task to parallel mode if and only if it contains several genuinely independent items and each item's work is at least one generation round. Below that threshold the planning-round tax exceeds the width gain, and parallel breadth converts into context weight — which converts back into wall clock. Documenting 23 modules is many items. Fixing one git repo, writing one file, solving one contest problem: each is one item, and no agent restructuring changes that.
试了四种任务之后,我的判断标准。一个任务值不值得改成并行,就看两条:里面有没有若干件真正互不相干的事,以及每件事的工作量够不够一整轮生成。不够这条线,规划那一轮的开销就超过了并行省下的时间——而且并行摊开的东西最后都会变成上下文的重量,重量又变回时间。
给 23 个模块写文档,那是很多件事。修一个 git 仓库、写一个文件、解一道竞赛题,那都是一件事——不管你怎么改 agent 的结构,它还是一件事。
Where I was wrong我错的地方
Thinking budgets. I first measured AIME 2025 with a 7,168-token thinking cap while published results use full reasoning budgets. Correcting it moved the score materially, and the correction curve turned out to be its own finding:
第一个错:思考预算给太抠了。 我最早测 AIME 2025 的时候,把思考上限卡在 7,168 token,而外面公开的成绩都是敞开了给的。改过来之后分数差挺多——而且这条修正的曲线本身,反倒成了个有意思的发现:
| setting设置 | accuracy准确率 | wall耗时 |
|---|---|---|
| 7k cap, greedy7k 上限,贪心 | 40.0% | 29 min |
| 24k cap, greedy24k 上限,贪心 | 50.0% | 90 min |
| 60k cap, vendor sampling60k 上限,官方采样 | 56.7% | 193 min |
Three benchmarks reacted to the same knob in three different ways, which is the useful part: AIME +10 points from tokens alone (long-form maths is budget-bound); LiveCodeBench-hard unchanged with thinking on and six times the budget (contest code is capability-bound); Aider polyglot 26.5% → 8.8% with thinking on (reasoning ate the budget and truncated the emitted file — when the output is the deliverable, thinking is hostile).
KV pool sizing. I planned concurrency against worst-case max_tokens (3 concurrent × 60k cap = 180k, “the pool is full”). But vLLM allocates KV on demand: max_tokens is a fuse, not a reservation. Actual chains averaged 12–18k, so the same run could have used 8–10 concurrent and finished in roughly a third of the wall clock. Plan concurrency from typical footprint; let preemption handle the worst case.
同一个旋钮,三个基准给出三种完全不同的反应,这才是真正值得记的地方:
- AIME:光是多给 token,就涨了 10 分。这类长推理数学题,卡的是预算。
- LiveCodeBench 难题:开了思考、预算翻六倍,一分没动。这类题卡的是能力,给多少都没用。
- Aider polyglot:开了思考反而从 26.5% 掉到 8.8%。因为推理把预算吃光了,要交付的代码文件写一半就断了。当输出本身就是交付物的时候,思考是纯帮倒忙。
第二个错:KV 池的账我算错了。 我按最坏情况规划并发——3 个并发乘以 60k 上限等于 180k,“池子满了”。但 vLLM 的 KV 是用多少给多少,max_tokens 是个保险丝,不是提前订座。实际的推理链平均才 12–18k,也就是说那次运行本来可以开 8 到 10 个并发,时间能砍到三分之一。
并发要按平均用量算,最坏情况交给抢占机制去兜。
Absolute numbers, honestly placed绝对成绩,诚实定位
A quantised 27B on a desktop does not compete with GPT-5 on rank, and nobody should expect it to. For an anchor I ran GPQA Diamond — all 198 graduate-level science questions, standard zero-shot CoT, vendor sampling, single sample, 130 minutes.
67.7% (134/198). By domain: Physics 90.7%, Biology 57.9%, Chemistry 48.4%. For context, frontier models sit near 94%, the best open flagship (a 397B MoE) at 88.4%, and GPT-4o — the 2024 frontier — near 50%. As far as I can find, this is also the first public GPQA number for this model; its own card publishes agentic benchmarks and no science scores.
The timing is its own argument: a single question’s reasoning chain takes 3–8 minutes, but 24 concurrent brought the amortised cost to 39 seconds per question. Running the benchmark was itself a demonstration of the thesis.
桌面机上跑个量化 27B,排名上肯定打不过 GPT-5,本来也不该指望它打得过。但总得有个参照,于是我跑了 GPQA Diamond:全部 198 道研究生级别的科学题,标准 zero-shot CoT,官方推荐的采样参数,只采样一次,跑了 130 分钟。
67.7%(134/198)。 分科看:物理 90.7%、生物 57.9%、化学 48.4%。对比一下,前沿模型在 94% 上下,最强的开源旗舰(一个 397B 的 MoE)88.4%,而 2024 年还是前沿的 GPT-4o 大概 50%。另外我查了一圈,这好像还是这个模型 GPQA 的第一个公开数字——它自己的模型卡只报 agentic 类的基准,科学题一个没提。
顺便一提,跑这个基准的时间账本身就是个论据:单独一道题,模型要想 3 到 8 分钟;但 24 路并发一摊薄,平均每题只要 39 秒。
换句话说,我跑这个 benchmark 的过程,本身就是这篇文章想说的事。
What to actually run实际该跑什么
docker run -d --gpus all --ipc=host -p 8090:8090 \
-e HF_HOME=/hf -v /path/to/hf_cache:/hf \
--entrypoint vllm nvcr.io/nvidia/vllm:26.04-py3 serve Qwen/Qwen3.8-27B-FP8 \
--host 0.0.0.0 --port 8090 \
--max-model-len 131072 --max-num-seqs 128 \
--gpu-memory-utilization 0.68 \
--trust-remote-code \
--enable-auto-tool-choice --tool-call-parser qwen3_xml \
--reasoning-parser qwen3 \
--enable-prefix-caching \
--kv-cache-dtype fp8 --mamba-cache-dtype float16 \
--speculative-config '{"method":"mtp","num_speculative_tokens":2}'
Four of those flags are non-obvious and each breaks something silently by its absence:
--tool-call-parser qwen3_xml— this model’s template emits<tool_call><function=..>XML. Without it every coding agent dies on its first call.--reasoning-parser qwen3— without it, raw</think>tags leak into the assistant’s visible answer, which looks like a model quality problem and isn’t.--kv-cache-dtype fp8 --mamba-cache-dtype float16— 2.9× the KV pool, +32% aggregate throughput, no measured quality loss (§03).--speculative-config '{"method":"mtp",…}'— depth 2 is the sweet spot; depth 3 lost 9% single-stream in testing.
For one interactive session, ignore all of it and run llama.cpp with the Q4 GGUF and a draft model: 24.2 tok/s single-stream beats vLLM’s 14.4, because 17.5 GB of weights beats 28.5 GB when the division in §01 is what governs.
这里面有四个参数不太起眼,但每一个少了,都会以某种不吭声的方式给你添堵:
--tool-call-parser qwen3_xml—— 这模型的模板吐的是<tool_call><function=..>这种 XML。不加它,任何 coding agent 第一次调用工具就挂。--reasoning-parser qwen3—— 不加它,</think>标签会直接漏到用户看到的回答里,看着像模型有毛病,其实不是。--kv-cache-dtype fp8 --mamba-cache-dtype float16—— KV 池 2.9 倍,总吞吐 +32%,实测质量没掉(见第 03 节)。--speculative-config '{"method":"mtp",…}'—— 深度 2 最划算;我试过深度 3,单流反而掉 9%。
当然,如果你就是一个人开一个窗口聊天,上面这些全都可以不管——直接上 llama.cpp,配 Q4 GGUF 加个草稿模型,单流 24.2 tok/s,比 vLLM 的 14.4 快不少。
道理还是开头那道除法:17.5 GB 的权重,就是比 28.5 GB 读得快。
Three rules三条规则
- Latency is fixed; width is nearly free. Single-stream speed is weights ÷ bandwidth and no flag changes it. Everything else on this machine — batching, sampling, fan-out — is width, and width costs almost nothing until you oversubscribe the slots.
- Width buys correctness only through a verifier. With objective tests, six parallel attempts turned 1/5 into 4/5 and 1/10 into 4/10. With only majority voting, the same width buys a few points and can amplify systematic error into a confident wrong answer.
- Wall clock is serial tokens. Parallelise items that are genuinely independent and whose work exceeds one generation round; below that line, planning rounds and context weight eat the gain. This single rule explains a 3.8× win and two clean losses.
The machine was never slow. It was one lane of a 128-lane road, and most of what I learned was about which cargo is actually splittable.
一、单个请求的快慢没救,但宽度基本白送。 单流速度就是权重除以带宽,调什么参数都改不了。这台机器上剩下的一切——批处理、多路采样、fan-out——全都是宽度,而宽度在你把槽位撑爆之前,几乎不要钱。
二、光有宽度换不来正确率,得配一个验证器。 有客观测试的时候,六路并行能把 1/5 做成 4/5、1/10 做成 4/10。只能靠投票的时候,同样的宽度只值几分钱,还可能把一个系统性的错误放大成特别自信的错答案。
三、时间花在哪,只看有多少 token 必须一个接一个地生成。 所以只并行那些真正互不相干、而且每份工作量够一整轮生成的东西。不够这条线,规划的开销和上下文的重量会把收益吃得干干净净。就这一条,既解释了那次 3.8 倍,也解释了两次输得干干净净的失败。
说到底,这台机器从来就不慢。它是一条 128 车道的路,而我一直只开着一条道在走。这几天真正学到的,其实不是怎么把车开快,而是——哪些货是真的能拆开、分到不同车道上去的。