# MiniMax-H3 Ref2VA (4-Bit INT4 Safetensors) - Experimental Model quant, may have zombie sway and artifacts in 10 second runs. Pure int4 version, meant for lowest possible vram usage for experimentation.
This repository provides an optimized 4-bit INT4 Safetensors quantization of MiniMax-H3-Ref2VA, engineered for consumer GPUs (16 GB VRAM, RTX 4080 / RTX 4090 / Laptop GPUs).
## 🚀 Model Details
- Base Architecture: MiniMax-H3 Ref2VA Diffusion Transformer (DiT)
- Format: Zero-Copy .safetensors (Single 9.80 GiB file)
- Quantization Scheme: Symmetric 4-Bit Linear FastInt4Linear) with FP16/BF16 Scale & Bias preservation
- AdaLN Conditioning: Pruned & Restored Continuous AdaLN Curve (1025 grid points)
- VRAM Footprint: 9.80 GiB (100% GPU Resident, Zero PCIe Swapping)
- Loading Time: ~0.03 seconds
## âš¡ Performance & Compatibility
- TeaCache Compatible: Full native support for Block-Level TeaCache (~50% compute bypass).
- LoRA / Adapter Support: Fully compatible with LightX2V 4-step Turbo LoRA minimax_h3_ref2v_turbo_4step_v0.1_bf16.safetensors).
- Attention Engines: Supports SageAttention 2.0 Patched, FlashAttention-2, and native PyTorch SDPA.
## 📜 License & Attribution
- Base model licensed under the MiniMax H3 Community License Agreement, Copyright © 2026 MiniMax.
- Powered by MiniMax H3.
