Minimax H3 INT8 / INT6 / INT4 ConvRot
Sampler settings I currently use are: er_sde/res_multistep simple
Quantized weights of the Minimax H3 video generation model (FL2VA and REF2VA variants), packaged for ComfyUI. Four precision tiers are available so you can pick the right point on the speed / quality / VRAM/RAM curve for your hardware.
About these variants
Four precision levels are available for both FL2VA and REF2VA (all sizes are pruned checkpoints), trading VRAM and speed against generation quality:
BF16 — full-precision reference (~40 GB). Highest quality, largest footprint. Use if you have the VRAM and want the ceiling to compare against.
INT8 ConvRot (W8A8) — 8-bit weights and 8-bit activations (~21 GB). Near-lossless quality, runs on INT8 tensor cores (RTX 30-series and up). The default recommendation.
INT6 (W6A8) — 6-bit weights with 8-bit activations (~16 GB). Middle-ground between INT8 and INT4: noticeably smaller than INT8 with a small quality trade-off. Good pick when INT8 is close to fitting but not quite. Uses the W6A8 kernel path from comfy-kitchen PR #191.
INT4 (W4A8 mixed) — 4-bit weights with 8-bit activations (~12.5 GB), quantized by Kijai. Smallest footprint of the group. Uses a ConvRot codebook layout (comfy-kitchen #90) that decodes to INT8 at compute time, so it runs on the same INT8 hardware as W8A8 with the same throughput — the win is memory, not speed. Larger quality trade-off than INT6.
All quantized variants keep the sensitive layers (I/O, modulation, embeddings, output head) in BF16.
If you're not sure: start with INT8. Move to INT6 if you need a bit more room. Move to INT4 if you're actively hitting VRAM limits — the quality gap to INT8 (and INT8 to INT6) is small enough that it's a good tradeoff on tight budgets.
Model variants
FL2VA — first-last frame to video / audio
REF2VA — reference-conditioned: up to 9 reference images, 3 reference videos (each may carry its own paired soundtrack), and 3 standalone reference audio clips
Both models support t2v (text to video), i2v (image to video), v2v (video to video), a2v (audio to video), and multiple mixed references (image / video / audio). Both were further fine-tuned for higher-quality outputs in their intended use cases — FL2VA for first-and-last-frame interpolation, REF2VA for reference-driven generation.
Requirements
ComfyUI at a version that supports Minimax H3 natively.
For the INT4 (W4A8) model: ComfyUI v0.31.0 or newer (PR #15308).
An NVIDIA GPU with useful INT8 throughput (RTX 30-series and up recommended).