# Qwen Image 2.1 [MXFP8] 🚀
This is the optimized MXFP8 quantized version of the powerful Qwen Image 2.1 model. The goal of this upload is to bring the excellent visual generation and prompt comprehension capabilities of Qwen 2.1 to setups with limited resources, drastically reducing VRAM consumption with no noticeable loss in quality.
## ✨ Main Highlights
* MXFP8 Efficiency: Utilizes the Microscaling FP8 format to compress the model. This means it takes up almost half the disk space and VRAM compared to FP16/BF16 versions, while maintaining detail precision and structural fidelity.
* Accelerated Inference: Modern GPUs (RTX 5000 series) benefit greatly from FP8 compute, resulting in significantly faster generation and processing times.
* Accessibility: Perfect for running locally on graphics cards with lower VRAM, allowing for heavier workflows or higher resolutions that would normally cause an Out of Memory (OOM) error on the base model.
Description
Comments (2)
the real advantage is only for rtx 5000 and rx 9000 line cards, they have mxfp8 hardware acceleration, the rest will be decompressed programmatically into fp16/bf16 on the fly and work in it, the gain will only be in the size of the model, but int8 convrot gives the same or even better quality, the same size, and will accelerate starting from rtx 2000 and rx 7000. The maximum performance even for rtx 5000/rx 9000 in int8 is equal to mxfp8, for the rest int8 is faster. int8 convrot is the new "gold standard"
Yes, hardware acceleration is the big difference
In my tests on my RTX 5000 series, the MXFP8 actually runs up to 15% faster than the INT8. It only takes up a few extra megabytes of space/VRAM, so at least for the 5000 series, that small size difference is totally worth the extra speed boost
