Notes
← All posts
2026-08-20 ·Independent benchmark·NVIDIA GB10 / DGX Spark·llama.cpp

The DGX Spark isn't slow. It's bandwidth-starved — and speculative decoding helps. DGX Spark 不慢,是带宽饿着了——而投机解码能喂饱它

The DGX Spark isn't slow. It's bandwidth-starved, and that's fixable. Five speculative decoding strategies, eight tasks, one binary. DGX Spark 不是慢,是带宽不够,而这个能治。五种投机解码策略、八项任务、一个二进制。

208.24tok/s

The reading that started it. A 27B dense model on hardware whose bandwidth caps it near 15 tok/s. Something had to be wrong, and it turned out to be my measurement, not the hardware. 一切的起点,是这个读数。 一个 27B 稠密模型,跑在带宽只够它到 15 tok/s 的硬件上。肯定哪里不对——后来发现不对的是我的测量方法,不是硬件。

01 — The problem问题

Bandwidth, not compute瓶颈是带宽,不是算力

The DGX Spark pairs a GB10 superchip with 128 GB of unified LPDDR5X. Enough room for a 70B model on a desk, but only 273 GB/s to feed it. Running Qwen3.8-27B at Q4_K_XL (17.5 GB of weights), I measured 11.5 tokens per second. Across eight tasks (Chinese and English prose, code generation and editing, JSON, reasoning, translation, a 2,177-token summary) every median landed between 11.47 and 11.66 tok/s. A 1.7% total spread. The workload doesn’t matter; that’s what a machine looks like when it spends its time waiting on memory instead of computing.

Every generated token reads all 17.5 GB of weights. At the published 273 GB/s that caps generation near 15.6 tok/s; measured 11.5 is 74% of the ceiling. No configuration flag fixes arithmetic.

Speculative decoding fits this imbalance precisely: verifying k drafted tokens costs one forward pass (a single 17.5 GB read), so accepted guesses are paid for with idle compute instead of scarce bandwidth. The thing people complain about, the wasted FLOPS, is exactly what makes the fix work. The same reasoning applies to Strix Halo, Apple Silicon, and any unified-memory design.

DGX Spark 给了你 128 GB 统一内存,桌上放得下一个 70B 的模型——但只有 273 GB/s 的带宽去喂它。

我跑 Qwen3.8-27B 的 Q4_K_XL 量化(17.5 GB 权重),测出来 每秒 11.5 个 token。有意思的是接下来这件事:八项任务——中英文散文、代码生成、代码编辑、JSON、推理、翻译、一份 2,177 token 的长摘要——中位数全部落在 11.47 到 11.66 之间,前后差 1.7%。

任务是什么根本不重要。一台机器把时间全花在等内存、而不是算数上的时候,就是长这个样子。

原因不难算:每吐一个 token 都得把 17.5 GB 权重整个读一遍。按 273 GB/s 算,上限就在 15.6 tok/s 附近,我测到的 11.5 已经是天花板的 74%。这个数没法靠调参数改。

而投机解码恰好是冲着这种失衡来的:验证 k 个草稿 token 只需要一次前向传播,也就是一次 17.5 GB 的读取。换句话说,猜对的那些 token 是拿闲置算力换来的,没花带宽。大家平时抱怨的”算力浪费”,在这里恰恰是解药。同样的道理也适用于 Strix Halo、Apple Silicon,以及任何统一内存的设计。

02 — The lie那个假数字

Where 208 tok/s came from208 tok/s 是怎么来的

llama.cpp’s --spec-type ngram-mod drafts from repeated n-grams, no draft model needed. Benchmarked the usual way (same prompt, five runs, median) it reported up to 208 tok/s. But the first request in a fresh process is always ~11.5, and every later one climbs: the n-gram cache persists across requests inside the server process, so by run two it contains the complete answer from run one. The benchmark was measuring the cache replaying its own output. Setting "cache_prompt": false disables the KV prompt cache. It does not touch the n-gram draft cache, and no request-level flag does.

llama.cpp 有个 --spec-type ngram-mod,从重复出现的 n-gram 里凑草稿,连草稿模型都不用。我按常规方式测——同一个 prompt 跑五次取中位数——它报出了 208 tok/s。

在一台理论上限 15.6 的机器上。

真相是:新起的进程里,第一个请求永远是 11.5 左右,之后一次比一次快。因为那个 n-gram 缓存在服务进程里跨请求活着,跑到第二次的时候,它已经装着第一次的完整答案了。我测的根本不是模型,是缓存在回放自己刚说过的话。

顺带一提,"cache_prompt": false 只能关掉 KV prompt 缓存,碰不到 n-gram 那个草稿缓存——而且没有任何请求级的参数能关掉它。

If you benchmark llama.cpp's n-gram speculation by repeating a prompt, your numbers are wrong. Restart the process or vary the prompt. The honest cold-start figure here is 11.5–12.4 tok/s, no measurable benefit at all.如果你通过重复同一个 prompt 来测 llama.cpp 的 n-gram 投机解码,你的数字是错的。请重启进程或更换 prompt。这里诚实的冷启动数字是 11.5–12.4 tok/s,没有任何可测量的收益。
03 — The results结果

Five strategies on one binary同一个二进制上的五种策略

Same model, same quantization, same server binary, idle machine; the only variable is the --spec-type flag and its draft model. Five runs per cell, median reported, spread ±0.1–2.8%.

同一个模型、同样的量化、同一个服务端二进制、机器空着不干别的。唯一变的就是 --spec-type 这个参数和它挂的草稿模型。每格跑五次取中位数,离散度在 ±0.1% 到 2.8% 之间。

Token generation · tok/s Qwen3.8-27B UD-Q4_K_XL · median of 5 · higher is better
TaskBaseline Draft 2BDSparkDFlash2DFlash2 ×
zh-prose11.6614.4218.4322.15 1.90×
en-prose11.5413.2417.1921.09 1.83×
code-gen11.5415.5423.8928.06 2.43×
code-edit11.5412.9618.7123.09 2.00×
json-out11.5419.5820.8323.33 2.02×
reasoning11.5423.3526.6128.33 2.45×
translate11.5324.1124.6829.48 2.56×
long-ctx11.4714.0519.6322.02 1.92×
ngram-mod (cold, honest) is omitted: 11.5–12.4 tok/s, indistinguishable from baseline.

DFlash2 wins every task, at 1.83× to 2.56× the baseline. It is a block-diffusion drafter: it emits a whole block of guesses in one pass and traces a coherent path through the candidates, and it’s a 2 GB file sitting next to a 17.5 GB model. Note that its numbers finally vary by task (21–29 tok/s) while the baseline was flat: the bandwidth ceiling is no longer the binding constraint.

八项任务,DFlash2 全赢,从基线的 1.83 倍到 2.56 倍。

它是个块扩散草稿器:一次前向就吐出一整块猜测,然后在这些候选里找出一条连贯的路。而它本身只是个 2 GB 的文件,安静地躺在那个 17.5 GB 的模型旁边。

还有个细节值得注意:它的成绩终于开始随任务变化了,21 到 29 tok/s 之间摆动,而基线是一条平线。这说明带宽已经不再是那个卡住一切的东西了。

04 — Findings发现

Three things the numbers taught me这些数字教给我的三件事

Acceptance rate is not comparable across drafter architectures接受率无法跨草稿器架构比较

On code generation the sequential 2B drafter accepts 88% of its guesses and delivers 15.5 tok/s; DFlash2 accepts 75% and delivers 28.1. Ranking by acceptance rate, the metric every paper reports, picks the slower system by 1.8×. A rejected sequential guess is a whole wasted forward pass; a rejected block position costs almost nothing. When comparing sequential against block drafters, measure wall-clock throughput or nothing.

代码生成那一项上,串行的 2B 草稿器接受率有 88%,跑出 15.5 tok/s;DFlash2 接受率只有 75%,却跑出 28.1。

也就是说,如果你按接受率来排名——那个几乎每篇论文都在报的指标——你会挑中慢 1.8 倍的那个。

道理其实很简单:串行草稿被拒一次,就是一整次前向传播白费了;而块草稿里某个位置被拒,几乎不花钱。所以比较串行和块草稿器的时候,要么老老实实测墙钟吞吐,要么就别比。

Lossless in distribution, not reproducible in practice分布无损,但实践中不可复现

At temperature 0, DSpark and DFlash2 matched a no-speculation control byte-for-byte on only 6 of 8 tasks (the control itself was fully deterministic). The divergence I traced landed exactly on the sequence’s second-narrowest top-2 logprob margin, 0.022, where a different verification batch shape changed floating-point reduction order and flipped the argmax. The distribution is preserved, but bit-for-bit reproduction on a GPU is not happening. If a regression suite pins exact model output, speculative decoding will break it, and that is not a bug you can fix.

温度 0 下,DSpark 和 DFlash2 跟”不开投机”的对照组逐字节一致的,八项里只有六项——而对照组自己是完全确定的。

我把那次分歧追到了具体位置:它恰好落在整个序列中 top-2 logprob 间距第二窄的地方,0.022。验证时的批次形状变了,浮点归约的顺序跟着变,argmax 就翻了个面。

所以严格说,分布是保住了,但 GPU 上想要逐位复现,别想了。如果你的回归测试锁死了模型的精确输出,投机解码一定会打破它——而且这个 bug 你修不了。

The speedup belongs to the workload, not the setup加速属于负载,不属于配置

Deleting four words from one prompt (an instruction not to answer with an outline) changed what the model wrote, moved draft acceptance from 68% to 55%, and moved throughput 21%. Any single-prompt benchmark of this technique is one sample from a wide distribution. That’s why everything above uses eight tasks, and eight is still too few.

我从一个 prompt 里删掉了四个词——就是一句”别用提纲回答”——模型写出来的东西变了,草稿接受率从 68% 掉到 55%,吞吐跟着变了 21%。

四个词。

所以任何拿单个 prompt 测这项技术的跑分,都只是从一个很宽的分布里随手抓的一个点。这也是为什么上面所有数据我都用了八项任务——而说实话,八项还是太少。

05 — The other path另一条路

What about just running a MoE model?直接跑 MoE 模型不行吗?

The standard advice for bandwidth-limited hardware. A 90 GB DeepSeek-V4-Flash quant generates at 19.5–20.0 tok/s on the same tasks, 1.7× the dense baseline, but it loses to DFlash2 on every task, its prompt processing runs at roughly half the dense model’s rate (90–370 vs 206–753 tok/s), and it leaves 25 GB of headroom on the machine. The 2 GB drafter beats the 90 GB MoE on generation, prefill, and memory simultaneously.

这是带宽受限硬件上最常见的建议,所以我也试了。

一个 90 GB 的 DeepSeek-V4-Flash 量化版,同样的任务跑 19.5 到 20.0 tok/s,是稠密基线的 1.7 倍——听起来不错,但它每一项都输给 DFlash2。而且它处理 prompt 的速度只有稠密模型的一半(90–370 对 206–753 tok/s),跑起来还只剩 25 GB 的余量。

一个 2 GB 的草稿器,在生成速度、预填充速度和内存占用三件事上同时赢了一个 90 GB 的 MoE。

06 — Limits局限

Where this is weak这项工作的薄弱之处

  • DSpark ran from a community GGUF conversion (erlidev/Qwen3.8-27B-DSpark-Q8_0), unverified and untuned, so read “DFlash2 was faster here,” not “DFlash2 beats DSpark.”
  • Single-stream only. Speculative decoding requires --parallel 1 in llama.cpp; concurrent batching was not measured and may win for multi-user serving.
  • One model, one quantization. Drafter quality is model-specific; none of these ratios should be assumed to transfer.
  • No quality evaluation beyond the eight greedy-output comparisons, and the 273 GB/s bandwidth figure is the spec sheet, not a measurement.
  • DFlash2 is not merged. It lives in llama.cpp PR #27342 and needs a local build; released containers ship the incompatible DFlash 1 (wrong number of tensors; expected 81, got 58).
  • DSpark 用的是社区 GGUF 转换版(erlidev/Qwen3.8-27B-DSpark-Q8_0),未经验证也未调优,所以请理解为”DFlash2 在这里更快”,而不是”DFlash2 胜过 DSpark”。
  • 仅限单流。 llama.cpp 中投机解码需要 --parallel 1;并发批处理未做测量,在多用户服务场景下可能反超。
  • 单一模型、单一量化。 草稿器质量与具体模型强相关;这里的任何比例都不应假定可迁移。
  • 除八次贪心输出对比外没有质量评估,而且 273 GB/s 的带宽数字来自规格表,不是实测。
  • DFlash2 尚未合并。 它在 llama.cpp PR #27342 里,需要本地构建;已发布的容器带的是不兼容的 DFlash 1(wrong number of tensors; expected 81, got 58)。
07 — Reproduce复现

Running it yourself自己动手跑一遍

Machine
NVIDIA GB10
Memory
119 GiB unified
Driver
580.95.05
CUDA
13.0 / sm_121
llama.cpp
PR 27342 @ 5ecbe1a
Drafter
DFlash2 Q8_0 · 2 GB

Build & serve构建与启动

git clone --depth 1 https://github.com/ggml-org/llama.cpp.git
cd llama.cpp
git fetch --depth 1 origin pull/27342/head:pr-dflash2 && git switch pr-dflash2
cmake -B build -DCMAKE_BUILD_TYPE=Release -DGGML_CUDA=ON \
      -DCMAKE_CUDA_ARCHITECTURES=121 -DLLAMA_CURL=OFF
cmake --build build -j 20 --target llama-server

./build/bin/llama-server \
  -m Qwen3.8-27B-UD-Q4_K_XL.gguf \
  -md Qwen3.8-27B-DFlash2-Q8_0.gguf \
  --spec-type draft-dflash \
  -ngl 999 -ngld 999 -c 32768 --parallel 1

The drafter is incoai/Qwen3.8-27B-DFlash2-GGUF on Hugging Face; its block size is baked in, so the --spec-draft-* flags have no effect on it (with a conventional draft model, --spec-draft-p-min 0.75 mattered most). If you benchmark any of this: restart the server between measurements, use more than one prompt, and check whether your fastest number is physically possible before you believe it.

草稿器是 Hugging Face 上的 incoai/Qwen3.8-27B-DFlash2-GGUF。它的块大小是写死在文件里的,所以 --spec-draft-* 那一串参数对它统统没用(换成常规草稿模型的话,影响最大的是 --spec-draft-p-min 0.75)。

最后,如果你要自己测这里的任何东西,记住三件事:每次测量之间把服务重启一遍;别只用一个 prompt;还有——在你相信自己那个最快的数字之前,先算一下它在物理上到底可不可能。