A small FFT of the Qwen_Image model. This is the beginning of work on this model, so don't judge it too harshly. I tried to improve the realism, and I think it worked. I am waiting for constructive feedback and suggestions.
Description
A small FFT of the Qwen_Image model. This is the beginning of work on this model, so don't judge it too harshly. I tried to improve the realism, and I think it worked. I am waiting for constructive feedback and suggestions.
FAQ
Comments (21)
IMO aesthetically its far superior to the base QWEN model. Although I dont know what is causing the scan line effect in my images, its 100% my fault though, something wrong in my workflow. I love it, it has now replaced the base QWEN model for me.
PJ0 follows the prompt even more coherently than the base QWEN model. In my sample image below P0J Real followed the facial structure prompt perfectly, while the base model failed it.
EDIT: Wow amazing, It even gets text right at 1920x1088 resolution. I could rarely get the base QWEN model to get text legible/readable at that resolution.
Edit: After futher testing I cant seem to consistently get a woman with pale skin via prompting, they're almost always super tanned.
The demo images have the same issue,
Thank you for your feedback. Unfortunately, I do not know what can cause vertical lines, I do not see such an effect. And I'm glad you like it.
@becausereasons the grid on some images is a problem of the architecture of the base model not this checkpoint, it is a problem of the architecture of the DIT and MMDIT (Qwn-Images)
I am very interested to get a Q3_KS quantized version of this so I can run it in my Mac Mini. Do you have this on HuggingFace too? If so then this GGUF my Repo can be used to make quantized version - https://huggingface.co/spaces/ggml-org/gguf-my-repo
Hi! I have an HF model, but at the moment I'm not sure if there will be a quantized version. https://huggingface.co/speach1sdef178/PJ0_QwenImageArt_stage1
Loving this model, and so early in Qwen's release too. Fixes the plastic Qwen look.
I truly believe the base QWEN model is very much like how SDXL base and pretty much all of the Stable Diffusion bases came to us, in need of some serious fine tunes. QWEN has so much potential though, an incredible amount.
@fox23vang226 I totally agree, and that's why I started working on this model
@fox23vang226
Yeah, same sentiment here. Also worth mentioning how its creators are eagerly engaging with the community. It really feels like an open-source project, rather than a corporate product.
Although I'm worried that it's going to be increasingly difficult to gather community support considering the current turmoil with CivitAI and its payment providers. The website definitely peaked during late-SDXL early-Flux days, until The Great Purge started. Many creators will be flying over to other websites, often with their content paywalled which saddens me. This is effectively the first full fine-tune of Qwen-Image on this website, and it's receiving so little attention...
Well, let's stay positive for now and hope for the best. Rome wasn't made in a day, and SDXL was already out for a while when Pony and later Illustrious came out.
Thanks @speach1sdef178 for working on this, keep up the good work!
For best results, I'm running lightning lora 4 step with strength under 0.5, and using 5 to 6 steps in sampler to make up for the low lora strength. This helps to make it even less plastic or over contrasted.
@shinonomeiro Thanks for the feedback. I'm surprised myself that I don't see any other improvements from other authors. I really liked the model. Yes, it definitely has something to fix, but I'm sure it's worth it and I hope that more versions of this very interesting model will appear soon.
What about mixing with other char LORAS trained on QWEN? Working good?
@LDWorksDavid I just posted a pic here, combined with Polaroid lora
@speach1sdef178 I tested your Stage 2 version, and it seems to have less weight on Asian woman from the outputs? Not getting the same broad results as v1 in terms of asian faces. Much cleaner details though in v2, but I'm also running it at 0.7 cfg to avoid the v2's over-contrast and over-saturation compared to v1. Polaroid lora also adds to the non-asian weights, so when combined together it really westernizes faces.
Please make a quantized version I have a Horrible Laptop with only 8GB of VRAM lol
Let's see what we can do about it, but I can't promise 100%
Would be great to hear how you are training this amazing model! Learning rate, number of steps, bs, dataset etc.
And how the hell i will run it on my 4070 16 VRAM?😠
well it's always RAM + VRAM. You could find a script to turn the full weights into FP8, or just wait the extra time on generation
Download distorch2 nodes (Comfyui-multiGPU). I'm running Qwen FP8 on my 16 GB VRAM card and it's faster than Q5_1 gguf


















