Qwen Image 2.1 7B INT8 and INT4 (W4A8) ConvRot Quantizations, for use in ComfyUI
Qwen-Image-2.1
Qwen-Image-2.1, A unified text-to-image generation and image editing model in the Qwen family. With just 7B parameters in its visual generation component (32 Single-Stream DiT layers), Qwen-Image-2.1 balances generation quality, inference efficiency, and versatility.
Highlights
Efficient Image Generation: Qwen-Image-2.1 combines strong visual performance with fast inference and a compact design, making high-quality image creation accessible across a wide range of creative workflows.
Flexible Creative Control: With support for diverse inputs, outputs, and localized edits, Qwen-Image-2.1 gives creators the flexibility to explore ideas and refine details within a unified workflow.
Four key improvements define this release:
Compact and Efficient — A lightweight architecture with mixed-granularity attention and prefix KV cache reuse delivers strong image quality at low computational cost.
Native Transparency, Unified Creation and Editing — Generate regular or transparent (RGBA) images from text, edit transparent layers, and extract subjects from photographs—all in one model.
Versatile Editing — Support up to 10 reference images, specify local edits via circles, painted annotations, or separate masks, and preserve identity for people and products.
Realistic Textures and Refined Aesthetics — Improved typography, portrait lighting, and fine details for more visually compelling results.
Description
FAQ
Comments (17)
damn Qwen has this lifeless ai vibe aesthetic in most images
Yeah, it seems like it, LoRAs are definitely needed for T2I. The edit capabilities are a bit more interesting.
@tsolful yeah loras can fix this without problem however artifacts are very difficult to fix. This image has that
As long as it's open sourced, trainable, and the base knowledge and quality is ok,
all you need is to trust the community.
@ikekph5 I trust the community 100%, I'll probably train a new fantasy realism refiner when AItoolkit supports. Non-commercial licensing tho :/ https://huggingface.co/Qwen/Qwen-Image-2.1/blob/main/LICENSE
@tsolful the license is ... unfortunate ... Why just follows other qwen model like qwen3.8 with a revenue cap. Qwen image 2.1 is an old model to them anyways.
edit: GPT tells me they probably don't care about the old model. But the model doesn't seem to have safety filter and is too good at nsfw so they probably want the "research or evaluation purposes only" shield for lawsuits.
@ikekph5 This was posted to the QwenDev twitter account, so maybe it is not as bleak as it seems.
"We’ve received so much love for Qwen-Image-2.1 over the past 24 hours, thank you!! Also gotten a lot of questions about the license, especially around model outputs. So here’s the answer: Outputs are not part of the licensed Materials. Users retain the rights to images and other content they generate using the model."
It's really poor in comparison of Krea 2. I hope the editing features are better. Thanks for the INT8 anyway.
the license kills this model
It is far and away the best editing model I've messed with, robust and surprisingly fast.
It's an omni edit model so it's more like a replacement to the Flux 2 Klein.
@VeerGeer Yes, the editing is very nice. I keep Krea 2 for image generation and replace Flux Klein 9B by Qwen 2.1 for editing !
cant wait for the Loras, its actually good
still seeing the old halftone thing going on in image edit, so they didn't fix that in their new vae, lol
very light NSFW test showing that fine‑tuning is certainly needed, but even the bare knowledge is quite good. At the very least, there’s no censorship.
The camera angle changes are weak, but it's the best image enhancer (open source) I've seen so far. The consistency of character is very good. The screen-shot improvement is quite good.
It worked well when I used a prompt like this.
---
General Order: Rotate the camera angle 90 degrees to face the front of the character in image1.
Order1: The character is now directly facing the camera in a full front-facing view.
Order2: Both eyes and the front of the face/body should be fully visible, looking straight at the viewer.
Note: Original image1 is left side view of the subject.
---