Home / File reference / Pruned NVFP4

MiniMax H3 Pruned NVFP4: Compatible GPUs, VRAM and Download

Main model (diffusion) · NVFP4 · 11.66GB · pruned · updated 2026-09-18

ItemValue
File nameMiniMax-H3_FL2VA-NVFP4.safetensors(DmitryDB)
ComponentMain model (diffusion)
Quant / formatNVFP4
Size11.66 GB
ComfyUI folderdiffusion_models (main model folder)
SourceDmitryDB/MiniMax-H3-ComfyUI-Quants

Which GPUs can run it

736 × 416 · 5s: 40 GPUs workable

GTX 1060 6GB GTX 1070 8GB GTX 1080 8GB GTX 1080 Ti 11GB RTX 2060 6GB RTX 2060 12GB RTX 2060 Super 8GB RTX 2070 Super 8GB RTX 2080 Ti 11GB RTX 3050 8GB RTX 3060 8GB RTX 3060 12GB RTX 3060 Ti 8GB RTX 3070 8GB RTX 3070 Ti 8GB RTX 3080 10GB RTX 3080 12GB RTX 3080 Ti 12GB +22 more

1344 × 768 · 5s: 40 GPUs workable

GTX 1060 6GB GTX 1070 8GB GTX 1080 8GB GTX 1080 Ti 11GB RTX 2060 6GB RTX 2060 12GB RTX 2060 Super 8GB RTX 2070 Super 8GB RTX 2080 Ti 11GB RTX 3050 8GB RTX 3060 8GB RTX 3060 12GB RTX 3060 Ti 8GB RTX 3070 8GB RTX 3070 Ti 8GB RTX 3080 10GB RTX 3080 12GB RTX 3080 Ti 12GB +22 more

1344 × 768 · 15s: 31 GPUs workable

RTX 2060 12GB RTX 2080 Ti 11GB RTX 3060 12GB RTX 3060 Ti 8GB RTX 3070 8GB RTX 3070 Ti 8GB RTX 3080 10GB RTX 3080 12GB RTX 3080 Ti 12GB RTX 3090 24GB RTX 4060 8GB RTX 4060 Ti 8GB RTX 4060 Ti 16GB RTX 4070 12GB RTX 4070 Super 12GB RTX 4070 Ti 12GB RTX 4070 Ti Super 16GB RTX 4080 16GB +13 more

“Workable” = estimated status ok or warn with 32GB RAM / fast-disk (warn means weight offloading and noticeably slower).

Download · HuggingFace →

→ Run the numbers for your RAM/budget in the calculator (this file pre-filled)

FAQ

How much VRAM does MiniMax-H3_FL2VA-NVFP4.safetensors(DmitryDB) need?

The file is 11.66GB (Main model (diffusion)). Runtime VRAM also depends on resolution and duration; a 720p-class short clip typically needs roughly 8GB+ of free VRAM. Run the calculator for your exact card.

Where does this file go in ComfyUI?

diffusion_models (main model folder). Keep the filename as-is; refresh ComfyUI and select it in the matching loader.

How does it compare with other quantizations?

See sibling entries in the model library and the INT8/BF16 comparison article. Rule of thumb: 8GB → GGUF Q3/Q2; 12-16GB → pruned INT8 or W4A8; 24GB+ → full INT8 or NVFP4 (RTX 50).

Similar files: BF16 · INT8 · P- BF16 · P- INT8 · P- FP8 · BF16 · INT8 · P- BF16 · P- INT8 · P- FP8

Related: Model library · INT8/BF16 comparison · 8GB VRAM setups