Notes·Inference, measurement, hardware

Measuring things that run locally.

Benchmarks of local LLM inference on unified-memory hardware, with the methodology written down, the raw numbers attached, and the mistakes left in.

The DGX Spark isn't slow. It's bandwidth-starved — and speculative decoding helps.

The DGX Spark isn't slow. It's bandwidth-starved, and that's fixable. Five speculative decoding strategies, eight tasks, one binary.

Baseline
11.5 tok/s
Best
29.5 tok/s
Speedup
1.8–2.6×
Runs
n=5 / cell