Notes
← All posts
2026-09-04 ·Deployment notes·NVIDIA GB10 / DGX Spark·MiniMax H3 / ComfyUI

The video model didn't fit. The text encoder was the problem. 装不下的不是视频模型,是文本编码器

A 33B video model shipped with a 32B text encoder, on a box with 119 GB of unified memory and no OOM killer to catch you. One device flag moved 27 GB off the peak for free. MiniMax H3 是个 33B 的视频模型,还配了个 32B 的文本编码器。在只有 119 GB 统一内存、且没有 OOM killer 兜底的机器上,一个 device 参数把峰值削掉了 27 GB,而且不要钱。

1.8GiB

What docker stats reported. The container was holding roughly 54 GB of the machine at that moment. On unified memory, CUDA and UVM allocations bypass cgroup accounting entirely. Every memory number here is measured on the host, because the container numbers are fiction. docker stats 给出的读数。 而那个容器当时实际占着约 54 GB。在统一内存机器上,CUDA/UVM 的分配完全绕过容器的 cgroup 记账。本文所有内存数字都在宿主机上测——容器里的数字是假的。

01 — The wall内存墙

Two big models that have to coexist两个大模型必须共存

MiniMax H3 is an open-weight video generation model: text, images and reference clips in, video with native stereo audio out. Architecturally it is two large models that both have to be resident, or at least resident in turn — a 33B packed-DiT that does the denoising (66 GB at bf16), and a Qwen3-VL-32B text encoder that turns the prompt into conditioning (51 GB at bf16), plus a video VAE (5.2 GB) and a separate audio VAE (0.6 GB).

That is 123 GB of weights before a single activation is allocated, on a machine with 119 GB total — and “total” means everything, because the GPU and the CPU share one pool. bf16 is not a near miss. It is not on the table.

So you quantize. The repackaged ComfyUI weights offer a ladder: a pruned checkpoint at int8 (21 GB), the full checkpoint at int8 (34 GB), an NVFP4 text encoder (15.7 GB). The interesting question is not which of these loads — several do. It is which ones leave room for the activations, and what you give up for the ones that do.

MiniMax H3 是个开源的视频生成模型:进去是文本、图片、参考片段,出来是带原生立体声的视频。从结构上说它是两个大模型,且必须同时(或至少轮流)驻留内存——一个 33B 的 packed-DiT 负责去噪(bf16 下 66 GB),一个 Qwen3-VL-32B 的文本编码器把 prompt 变成 conditioning(bf16 下 51 GB),外加一个视频 VAE(5.2 GB)和一个独立的音频 VAE(0.6 GB)。

也就是说,一个激活值都还没分配,光权重就 123 GB,而机器总共 119 GB——这里的「总共」是字面意义上的全部,因为 GPU 和 CPU 共用同一块内存。bf16 不是差一点,是根本不在选项里。

那就量化。ComfyUI 格式的重打包提供了一条阶梯:剪枝版 int8(21 GB)、完整版 int8(34 GB)、NVFP4 文本编码器(15.7 GB)。真正的问题不是哪个能加载——好几个都能。而是哪个加载完还剩得下激活值,以及剩得下的那些,代价是什么。

Everything below is one machine, one model, mostly single runs. This was a deployment, not a benchmark, and I flag every number that is thin.下面所有数据都来自一台机器、一个模型,大多是单次运行。这是一次部署记录,不是基准测试;凡是证据薄的地方我都会标出来。
02 — The accounting lie看不见的 54 GB

Memory that nothing reports没有工具报告的内存

Before any model work, a simpler problem: the machine claimed 56 GB was in use and nothing could account for it. No process had an RSS above 0.5 GB. docker stats showed the running containers totalling under 2 GB. Free memory, page cache, buffers and slab added up left roughly 54 GB unexplained.

It was a single container holding GPU memory. On a discrete-GPU box that allocation shows up in nvidia-smi and not in free. Here it is the same physical RAM, allocated through CUDA — so it appears in the host’s “used” total, does not appear in the process’s RSS, and does not appear in the container’s cgroup accounting. Stopping that container took available memory from 62 GB to 114 GB.

在碰模型之前,先有个更朴素的问题:机器显示 56 GB 已用,但没有任何东西能解释这 56 GB 去哪了。没有进程的 RSS 超过 0.5 GB,docker stats 显示所有运行中的容器加起来不到 2 GB。把空闲内存、page cache、buffer、slab 全加上,仍有约 54 GB 对不上账。

是一个容器占着 GPU 内存。在独显机器上,这笔分配会出现在 nvidia-smi 里、不会出现在 free 里。而这里它就是同一块物理内存,只是通过 CUDA 分配的——于是它计入宿主机的 used,不计入进程的 RSS,也不计入容器的 cgroup。把那个容器停掉,可用内存从 62 GB 涨到 114 GB。

On unified-memory hardware, docker stats cannot tell you whether a container is using the machine. Compare host free -g before and after stopping it. This is also why the usual safety net is missing: these pages are unswappable, so the kernel livelocks in reclaim rather than firing the OOM killer, and the desktop freezes with no error anywhere.在统一内存机器上,docker stats 无法告诉你一个容器是不是在吃内存。停掉它,对比前后宿主机的 free -g。这也是为什么平时那张安全网是缺失的:这些页面不可换出,内核会在回收里活锁,而不是触发 OOM killer——桌面直接冻住,任何地方都不留错误信息。

That is why I added --reserve-vram 10 to the inference server before doing anything else. It asks the allocator to leave 10 GB for the OS, converting “the machine stops responding” into “the job raises an out-of-memory error.” Measured cost on the standard configuration: peak went from 91 GB to 88 GB, runtime unchanged.

所以我在做别的事之前,先给推理服务加了 --reserve-vram 10:让分配器给系统留 10 GB,把「整机失去响应」变成「任务报一个 OOM 错误」。实测代价:标准配置下峰值从 91 GB 降到 88 GB,耗时不变。

03 — The 2× rule2 倍规律

13 GB of weights, 24 GB of peak13 GB 的权重,24 GB 的峰值

640×384 · 56 frames · 4 steps peak host memory · single run each
TransformerOn diskPeakTime
NVFP4, pruned12.5 GB72 GB65 s
int8, pruned21.0 GB86 GB80 s
int8, full34.0 GB110 GB100 s
A 13 GB increase in weights cost 24 GB of peak. The full checkpoint left 9 GB of headroom at draft resolution — at 720p it would not have survived.

The rule of thumb that fell out of this — every extra gigabyte of weights costs about two gigabytes of peak — held up well enough to plan with. I used it to predict the NVFP4 figure before measuring it: predicted ~70 GB, measured 72.

I do not have a clean mechanism for it. The loader reports “detected mixed precision quantization” and a bf16 manual-cast path, so some working copy is plausibly involved, but I did not instrument it and will not pretend to know. The practical reading was simple, and — as it turned out — wrong: the full checkpoint is unusable on this machine.

由此得到的经验规律是:权重每多 1 GB,运行峰值多约 2 GB。这条规律拿来做规划够用了——我用它在实测前预测了 NVFP4 那一档,预测约 70 GB,实测 72 GB。

至于为什么是 2 倍,我没有干净的机制解释。加载器打印了「检测到混合精度量化」和一条 bf16 手动 cast 的路径,所以中间大概存在某种工作副本,但我没有去插桩,也不打算假装知道。当时得出的实用结论很简单,而且——后来证明——是错的:完整版权重在这台机器上不可用。

04 — The fix解法

Put the text encoder on the CPU把文本编码器放到 CPU 上

The text encoder runs once. It converts the prompt into conditioning tensors and is then dead weight for the entire sampling loop — which, at 720p, is four minutes of a five-minute job. On a discrete GPU you would want it in VRAM anyway, because moving 16 GB across PCIe is expensive. On unified memory there is no bus to cross. The weights sit in the same RAM either way.

ComfyUI’s CLIPLoader has a device input with a cpu option. Switching it:

文本编码器只跑一次。它把 prompt 变成 conditioning 张量,然后在整个采样循环里就是死重——而在 720p 下,采样占了五分钟里的四分钟。在独显机器上你当然希望它待在显存里,因为把 16 GB 搬过 PCIe 很贵。但统一内存没有总线要过,无论放哪,权重都在同一块 RAM 里。

ComfyUI 的 CLIPLoader 有个 device 输入,可以选 cpu。切过去之后:

1280×704 · 124 frames (5.17 s) · 4 steps peak host memory · wall clock · single run each
TransformerEncoder on GPUEncoder on CPU
int8, pruned (21 GB)88 GB · 245 s61 GB · 238 s
int8, full (34 GB)would not fit85 GB · 261 s
The pruned configuration lost 27 GB of peak and, within run-to-run noise, no time at all.

Two things in that table are worth separating. The first is that it unlocks the full checkpoint: 85 GB is comfortable, and the configuration I had written off as impossible now runs at 720p with 34 GB of headroom.

The second is more useful. The default configuration — the pruned checkpoint everyone will actually use — drops from 88 GB to 61 GB. That is the difference between a machine one bad setting away from freezing and one you can leave alone.

It saves more than the encoder weighs: 15.7 GB of weights moving off the GPU removed 25–27 GB of peak. Same ~2× relationship as the previous section, seen from the other direction.

这张表里有两件事值得分开说。第一件是它解锁了完整版权重:85 GB 很宽裕,那个我本来判了死刑的配置,现在能在 720p 下跑,还剩 34 GB 余量。

第二件更有用。默认那档——也就是大家真正会用的剪枝版——从 88 GB 降到 61 GB。这是「离冻机只差一个设置」和「可以放着不管」之间的差别。

它省下的比编码器本身还多:15.7 GB 权重挪出 GPU,削掉了 25–27 GB 峰值。还是上一节那个 2 倍关系,只不过是反着看。

At 640×384 the CPU encoder cost about 20 s. At 1280×704 it cost nothing measurable — the encode takes the same wall time either way, it is just a smaller fraction of a longer job. Both are single runs, so read "cost nothing" as "cost less than the noise," not as a claim about zero.在 640×384 下,CPU 编码器大约多花 20 秒;在 1280×704 下测不出差别——编码本身耗时是一样的,只是在更长的任务里占比更小。两边都是单次运行,所以「不要钱」应理解为「代价小于噪声」,而不是真的等于零。
05 — Methodology方法

Why you cannot diff two quantizations为什么不能直接 diff 两个量化

Having three checkpoints, the obvious question is what the smaller ones cost. My first attempt was to generate the same prompt at the same seed under two quantizations and compute PSNR between the outputs. PSNR came back at 8.9 dB, SSIM at 0.379 — readings that say “the quantization destroyed the output.”

Both readings are meaningless. Both videos are perfectly coherent. They are simply different videos.

Quantization perturbs the numerics, the perturbation moves the sampling trajectory, and at four steps a small early divergence compounds into a different composition. Same seed, same prompt, completely different shot. PSNR between them measures how far the trajectories diverged, which is not a quality signal at all.

手上有三个 checkpoint,自然要问小的那些代价是什么。我第一次尝试是:同一个 prompt、同一个 seed,在两种量化下各生成一遍,然后算 PSNR。结果 PSNR 是 8.9 dB,SSIM 是 0.379——这两个读数看上去在说「量化把输出毁了」。

两个读数都没有意义。两条视频都完全连贯,它们只是两条不同的视频。

量化扰动了数值,扰动改变了采样轨迹,而在只有四步的情况下,早期一点点偏离会滚成完全不同的构图。同 seed、同 prompt、完全不同的镜头。它们之间的 PSNR 衡量的是轨迹分叉了多远,根本不是画质信号。

For diffusion models, per-pixel metrics can only compare things that share a trajectory. They are the right tool for a decoder or encoder change, where the latent is fixed — I use them for exactly that in the next section. They are the wrong tool for anything that touches sampling.对扩散模型来说,逐像素指标只能比较共享同一条轨迹的东西。换解码器、换编码器这类 latent 固定的改动,它们是对的工具——下一节我正是这么用的。但凡动到采样,它们就是错的工具。

So I looked at frames instead: three seeds per configuration, at 1:1 with no rescaling. That is a weak instrument — n=3, one prompt, my eyes — and it is the weakest evidence in this post. But the two comparisons came out very differently, which is at least a sign it was not returning noise:

  • Pruned vs full int8: hard to tell apart. Both hold background detail, both resolve individual guard hairs, contrast matches. The full checkpoint looked marginally cleaner around the eyes on one seed of three. If the memory cost were real, I would not chase it.
  • int8 vs NVFP4: consistent and visible across all three seeds. Backgrounds flatten — a treeline present in the int8 frame is smeared into a gradient. Fur loses micro-detail. Contrast drops.

NVFP4 is 14 GB cheaper and about 25% faster. On this hardware it is a draft setting, not a quality setting.

于是改成看帧:每种配置三个 seed,1:1 不缩放地对比。这是个很弱的仪器——n=3、一个 prompt、我的眼睛——也是本文中证据最薄的一环。但两组对比的结果差别很大,这至少说明它不是在返回噪声:

  • **剪枝版 vs 完整版 int8:**很难分辨。背景层次都在,毛发都能分出根,对比度一致。三个 seed 里有一个,完整版在眼部略干净一点。如果内存代价是真的,我不会去追这个差别。
  • **int8 vs NVFP4:**三个 seed 上一致且明显。背景被压平——int8 帧里有树林的地方,NVFP4 里糊成了一片渐变。毛发丢失微观细节,对比度下降。

NVFP4 便宜 14 GB,快约 25%。在这台机器上它是草稿档,不是画质档。

06 — The complaint那句抱怨

“It looks a bit soft”「片子有点糊」

With the memory problem solved, the remaining complaint was about the output itself: finished clips looked softer than expected. This is the kind of report that invites guessing, so I tried to eliminate suspects with measurements that could actually exonerate them.

Suspect 1 — the H.264 encode. Output bitrate was 1.3–3.4 Mbps at 720p, which looks low. But the encoder runs at CRF 23: constant quality, not constant bitrate. A low bitrate at fixed CRF means the source had little high-frequency content to preserve, which points away from the encoder rather than at it. Confirmed by writing lossless PNG frames from the same graph and comparing crops at 1:1 — the MP4 frame and the PNG frame are near-identical, fur detail intact. Not the encoder.

Suspect 2 — tiled VAE decoding. Decoding 720p in 512-pixel tiles is exactly the kind of thing that softens seams. I ran one sampling pass and fed the identical latent to both a tiled decode and a whole-frame decode in the same graph. Same latent, so per-pixel comparison is valid here. PSNR came back inf; the files were byte-for-byte identical. At this resolution the tiling does nothing at all. Not the decoder.

Suspect 3 — the low-precision speed flags. The server was running --fast fp16_accumulation fp8_matrix_mult cublas_ops and --bf16-vae, inherited from an unrelated model’s configuration. ComfyUI’s own help text calls these “potentially quality deteriorating,” so this was my leading theory. Restarting without them and regenerating at the same seed gave PSNR 49.6 dB against the original, and average lossless PNG sizes of 1248 KB versus 1241 KB. Above roughly 45 dB the difference is not visible. Not the flags — and worth knowing, because they are free.

内存问题解决之后,剩下的抱怨是关于成片本身的:看起来比预期糊。这种反馈很容易引人乱猜,所以我尽量用能真正洗清嫌疑的测量去逐个排除。

嫌疑一:H.264 编码。 720p 下码率只有 1.3–3.4 Mbps,看着确实低。但编码器跑的是 CRF 23,是恒定质量,不是恒定码率。固定 CRF 下码率低,说明源本身就没多少高频要保,这指向的是源而不是编码器。验证方式:在同一个工作流里另存无损 PNG,1:1 裁切对比——mp4 帧和 PNG 帧几乎一样,毛发细节完好。不是编码器。

嫌疑二:VAE 分块解码。 720p 用 512 的分块解码,正是那种会让接缝发软的操作。我跑一次采样,把同一个 latent 同时喂给分块解码和整幅解码。latent 相同,所以这里逐像素比较是成立的。PSNR 返回 inf,两个文件二进制完全相同。在这个分辨率下,分块什么也没做。不是解码器。

嫌疑三:低精度加速开关。 服务上带着 --fast fp16_accumulation fp8_matrix_mult cublas_ops 和 --bf16-vae,是从另一个模型的配置里继承过来的。ComfyUI 自己的帮助文本说这些「可能劣化画质」,所以这曾是我的首要怀疑对象。关掉之后同 seed 重跑,与原图的 PSNR 是 49.6 dB,无损 PNG 平均大小 1248 KB 对 1241 KB。45 dB 以上肉眼分辨不出。不是这些开关——而且这个结论有用,因为它们是免费的。

The answer was in the model card. H3's open weights generate at a 768-pixel short edge; 1280×704 is 0.9 megapixels. Displayed full-screen on any modern panel, that is being upscaled before it reaches your eye. The product's 2K output comes from a regeneration module available only through the hosted API — it is not in the released weights. Nothing was broken. I was looking at a 768p video on a much larger screen.答案写在模型卡里。H3 的开源权重原生就是短边 768,1280×704 只有 0.9 兆像素。在任何现代屏幕上全屏播放,它在到达你眼睛之前就已经被放大过了。产品端的 2K 来自一个只在托管 API 上提供的重生成模块,开源权重里没有。什么都没坏。我只是在一块大得多的屏幕上看一个 768p 的视频。

The local mitigation is an ESRGAN pass after decode. Comparing three ways from 1280×704 to 2560×1408 on the same frame: Lanczos is visibly soft, RealESRGAN ×2 resolves the rim-lit hairs on the ear, and 4×-UltraSharp downsampled to 2× is sharper still with a hint of over-sharpening. Folding the ×2 model into the pipeline cost 2 GB of peak and 65 s on a 56-frame clip.

That does not add detail the model never generated. It stops the display from doing a worse job of the upscale than a purpose-built network would.

本地的缓解办法是在解码后接一道 ESRGAN 超分。同一帧从 1280×704 放到 2560×1408,三种方式对比:Lanczos 明显发软,RealESRGAN ×2 能把耳缘逆光的针毛分出来,4×-UltraSharp 再降采样到 2× 更锐,但略有过冲。把 ×2 那个模型接进管线,代价是峰值 +2 GB、56 帧的片子多花 65 秒。

这并不能补出模型压根没生成的细节。它只是不让显示器用一个比专用网络更差的方法去做这次放大。

07 — Scaling长片的代价

Where the time goes时间花在哪

Sampling cost vs clip length 1280×704 · 4 steps · pruned int8
FramesDurationPer stepTotal job
1245.17 s28 s245 s
36215.08 s148 s859 s
2.9× the tokens, 5.3× the cost per step. Attention is quadratic and you can watch it happen. Peak memory at 362 frames: 93 GB.

Roughly half of a short job is not sampling at all — it is text encoding, VAE decode and muxing. That changes the intuition about step counts: going from the 4-step distilled path to a 20-step run without the acceleration LoRA is not five times slower. Measured at 640×384 it was 121 s against 80 s.

Whether those 20 steps look better is a question I did not answer, for the reason in section 05 — changing the step count changes the trajectory, so I have no way to compare them except by eye, and I did not trust my eye enough to publish a verdict.

短片子里大约一半的时间根本不在采样上,而在文本编码、VAE 解码和封装。这会改变你对步数的直觉:从 4 步蒸馏路径换到不挂加速 LoRA 的 20 步,并不是慢五倍。在 640×384 下实测是 121 秒 对 80 秒。

那 20 步是不是更好看,这个问题我没有回答,原因见第 05 节——改步数就改了轨迹,除了用眼睛看没有别的比较办法,而我不够信任自己的眼睛到可以下结论。

08 — Summary小结

What generalises可以带走的部分

  • Move one-shot models off the GPU. On unified memory, a text encoder that runs once at the start of a job has no business occupying the peak. One device flag was worth more than every quantization decision in this post combined.
  • Container memory tools do not see CUDA allocations. Compare host free before and after, or you will chase ghosts.
  • Reserve headroom explicitly. Without an OOM killer to catch you, the failure mode is a frozen desktop rather than a stack trace. Ten gigabytes of reserve cost nothing measurable.
  • Do not diff samples from different trajectories. Per-pixel metrics are for fixed-latent comparisons. For anything that touches sampling you are stuck looking at frames — so say so, and treat it as the weak evidence it is.
  • Check the model card before optimizing. Three careful eliminations, and the answer to “why is it soft” was the native resolution, documented, in the first paragraph.
  • **把只跑一次的模型挪出 GPU。**在统一内存上,一个只在任务开头跑一次的文本编码器,没有理由占着峰值。一个 device 参数的价值,超过本文里所有量化决策加起来。
  • **容器的内存工具看不见 CUDA 分配。**停掉前后对比宿主机的 free,否则你会一直在追鬼。
  • **显式预留余量。**没有 OOM killer 兜底时,失败形态是桌面冻住而不是一条栈回溯。预留十个 G,测不出代价。
  • **不要 diff 来自不同轨迹的采样结果。**逐像素指标只适用于 latent 固定的比较。凡是动到采样的,你只能靠看帧——那就明说,并且承认这是弱证据。
  • **优化之前先读模型卡。**三次仔细的排除,而「为什么糊」的答案,就写在模型卡第一段的原生分辨率里。
09 — Reproduce复现

Configuration配置

Machine
NVIDIA GB10
Memory
119 GiB unified
Driver
580.95.05
CUDA
13.0 / sm_121
Runtime
ComfyUI 0.34.0
Weights
Comfy-Org/MiniMax-H3

Server flags服务启动参数

--use-sage-attention --bf16-vae --bf16-text-enc \
  --reserve-vram 10 \
  --fast fp16_accumulation fp8_matrix_mult cublas_ops

The --fast group is measurably free here (49.6 dB against a run without it). --reserve-vram is the one that matters; pick a value that leaves your desktop alive.

--fast 这组在这里实测是免费的(与不带它的一次运行相比 49.6 dB)。真正重要的是 --reserve-vram,取一个能让桌面活着的值。

The flag this post is about本文的主角

CLIPLoader:
  clip_name: qwen3vl_32b_minimax_h3_nvfp4_awq.safetensors
  type:      minimax
  device:    cpu          # default -> cpu: -27 GB peak at 720p

Graph shape工作流形状

UNETLoader --+-- LoraLoaderModelOnly -- MiniMaxH3SigmaShift --+
             |                                                |
CLIPLoader --+-- MiniMaxH3ImageToVideo --+-- positive --------+
                                         +-- AV latent -------+
                                                           KSampler
                                           steps 4, cfg 1.0, euler/simple
                                                              |
                             VAEDecodeTiled (video) ----------+
                             LTXVAudioVAEDecode (audio) ------+
                                         |
                             [ImageUpscaleWithModel x2]   <- optional
                                         |
                                CreateVideo -> SaveVideo

Frame counts must land on the model’s 17k+5 grid (5, 22, 39, … 124, … 362) at 24 fps; the nodes round up silently otherwise. cfg stays at 1.0 because the released checkpoints are CFG-distilled — that is not a speed compromise, it is the intended setting.

帧数必须落在模型的 17k+5 网格上(5、22、39、…124、…362),24 fps;给别的值节点会静默向上取整。cfg 保持 1.0,因为放出来的权重本身就是 CFG 蒸馏过的——这不是为速度做的妥协,就是它该用的设置。