Z-Image-Turbo is live for Generation! We're monitoring performance, and we'll be tweaking functionality behind the scenes over the next few days!
Z-Image is a powerful and highly efficient image generation model with 6B parameters. It is currently has three variants:
🚀 Z-Image-Turbo (this model) – A distilled version of Z-Image that matches or exceeds leading competitors with only 8 NFEs (Number of Function Evaluations). It offers ⚡️sub-second inference latency⚡️ on enterprise-grade H800 GPUs and fits comfortably within 16G VRAM consumer devices. It excels in photorealistic image generation, bilingual text rendering (English & Chinese), and robust instruction adherence.
🧱 Z-Image-Base – The non-distilled foundation model. By releasing this checkpoint, we aim to unlock the full potential for community-driven fine-tuning and custom development.
✍️ Z-Image-Edit – A variant fine-tuned on Z-Image specifically for image editing tasks. It supports creative image-to-image generation with impressive instruction-following capabilities, allowing for precise edits based on natural language prompts.
Original ComfyUI Models: https://huggingface.co/Comfy-Org/z_image_turbo/tree/main/split_files
Original HF Repo: https://huggingface.co/Tongyi-MAI/Z-Image-Turbo
Description
FAQ
Comments (829)
Showing latest 429 of 829.
Decent model, images looks better than flux - there is no that weird light effect on a face, but it is not much cheaper, seems like similar i have 5 mins per 1Mpx image (on T4).
Good it has nudes out of the box, but things like nipples, and other details from under the clothes, are messed. Hands good, even at distance - no need for hand detailer. Also decent toes - something IL lacks a lot.
If it was like x5 times faster it would be good candidate for general purpose model - which can generate everything from regular text description.
It shouldn't take 5 mins on Tesla 4, the GPU isn't that slow, unless you run the model on CPU instead
@qek 14s/it + about min for loading model files. That's for 832x1216. And not run cpu for sure.
@Sunkarie You said: i have 5 mins per 1Mpx image
@Sunkarie If you use TAEF instead of AE for drafts, you won't waste like 3 seconds for decoding
@qek 14s * 15 iterations(maybe overkill) + 1min(loading, because 16gb ram not enough to keep all files loaded in memory - it load some from disk each time it start running ksampler) - thats where 5 mins totals.
Adding lora makes it even slower - more stuff to load from disk.
Just to clarify: 1Mpx = 832x1216 = ~1 megapixel.
I'll look into settings/setup one day, maybe there is a way to make iterations itself faster, i did tried change sampler, form res_multistep to euler - it did no impact on time and almost no impact on image itself.
牛逼王,这他妈出图也太快,质量也太好了,就这显存占用这他妈也太牛逼了,这他妈也太支持NSFW,无话可说。目前最完美的模型我草了
Hey CivitAi community,
I have a system with low VRAM. I’m able to run a full size FP16 model using a Qwen3 4B GGUF version as the text encoder, since running the full-size model with its original encoder is too heavy for my system. I was wondering aside from the general quality loss that comes with quantized (GGUF) models is there any difference in prompt understanding if I don’t use the official encoder?
"general quality loss that comes with quantized (GGUF) models" May be slower and slightly lower quality, unsloth's Dynamic 2 quants appear to be superb
"...if I don’t use the official encoder?" Quantization doesn't make Qwen3 4B unofficial, do you use Qwen3 4B or a fine-tune of it?
Personally I don't recommend you use GGUFs as TE. They make generation slower (as I can understand they need a time to ungguf). If you use a ggufed checkpoint probably it'll be impossible to use it due two unggufing processes.
Here is a FP8 of the encoder:
https://huggingface.co/jiangchengchengNLP/qwen3-4b-fp8-scaled
@qek
I’m using Qwen3-4B-Q4_K_S.gguf from unsloth/Qwen3-4B-GGUF.
I was wondering if the text encoder included with the Z-image model might be better optimized for prompt understanding in text-to-image tasks, or specifically for Z-image. I’m not very familiar with the technical differences, so I might be mistaken.
@mphobbit Thank you!
@ALFARANKO I experimented with FP8 of vanilla LLM as a text encoder first time and it gave me mess, not a picture. GGUFs gave me only "size mismatch error". I'm not Comfy user so I can mistake in case of Comfy.
@mphobbit I have to use a 4-bit quant of it because I got OOM when tried ControlNet
@qek full-size can provide OOM, due its LLM origin and as I can understand it consumes memory in a different way than previous-gen TE.
@ALFARANKO Thank you! This variant 30% faster than standart model!
Qwen3 4B official + zimage turbo can i run on 8gb vram?
@Wendy_Earth It works but LLM can freeze for a while your PC. Use FP8 TE (the link is in my comment above)
@mphobbit okay
@Wendy_Earth i'm using the Forge Neo UI, and honestly it runs pretty well. I can handle an FP16 model with FP8 CLIP on just 4 GB of VRAM with almost no errors. So if you’ve got 8 GB, it should be smooth overall there might be a bit of offloading, but you really shouldn’t hit any errors.
Can anyone give me a workflow for img2img?
Just use VAE Encode instead of Empty Latent Image
can you share this maybe? I cant find any working workflows
Z Image with no Turbo https://huggingface.co/ostris/Z-Image-De-Turbo
Is this any good? What's the difference?
@Juanamalaa Not distilled
@qek it is distilled, it was then de-distilled. so it's still worse than base.
@TienX k
@TienX It seems Astralite will fine-tune it https://huggingface.co/purplesmartai/zony-v8-256px-exp-de-distilled
@Juanamalaa In theory you can use cfg scale and more steps, that doesn't mean it's good/better - I believe it's mostly convenience for training and checkpoint merging. Ostris' adapter worked just fine though for training. I think most people would simply want to use the turbo version for general use.
Woah?
Why is nobody mentioning the water issue?
Any time you gen some animal with fur, it looks wet. Even genning humans their hair is very often somewhat wet/oily, but fur is a much bigger problem. It's almost impossible to gen a completely dry looking animal.
Genning a bunny for some reason made it gen something with normal dry fur. But bear and cheetah are wet and cats and dogs a little bit but not as much it seems.
You are right, I noticed that with my grinch image too. You can try using 2 samplers, for example 0-4 and 4-9 steps. The "euler_ancestral" sampler for the 2nd half gives smoother results.
It generates cars WITH driver. And that beautifully. For that alone I'm happy. Thank you.
I'm not sure. For me, z-image is a bit tricky to control. There's a tendency for the results to be too narrow. However, it's fast to create, lightweight, and follows prompts accurately. If FLUX is a cake, Z-image is like a brownie. Flux is rich but dull, and Z-image is small but intense.
Flux2 is like an elephant I can't run without extreme quantization and no fast generations. Z Image can't edit images like Flux2 does, but there will be Z Edit
@qek a caged elephant too, censored so much that it will never understand basic anatomy
Очень чувствительна к подсказкам, это новая эпоха генерации, сейчас надо более точно описывать то, что вы хотите получить.
@qek Yes, Flux2 is not attractive at all.
@gausssidorov928 Да, я всегда предпочитал абстрактные концепции, поэтому Z-изображения кажутся мне немного сложными. Например, мне приходится описывать настроение визуально, а не эмоционально. Сложнее использовать воображение, но если удаётся описать его точно, можно сократить количество итераций. Z-изображения определённо выглядят лучше предыдущих моделей.
@halfBurn Sadly, Z's prompt adherence is not excellent, but better than Chroma.1. I still can't get some images I want even when I use clear instructions because it may not get what I need
Stop dislikebombing it
Funny trick:
if add in promt with NFSW girls this: "her finger fully covering her well-depicted soft genitals with hidden inside moist labia."
Woman's pussy image getting not best, but much-much better, without guts spilling out.
:)
There's a problem with character training. The random characters are generated by a model that's not at all stable.
Post your config
Asian face too strong. even when prompting for (CAUCASIAN,WHITE,EUROPEAN,SLAVIC,RUSSIAN:1.3) i still get 100% korean looking girls. i really hope base doesn't have this problem
Try my LoRA, some of the images in my training set did have Asian faces so it didn't completely remove them, fighting with the model, but it also had other images to help with this. Base model does not have a problem at all, I've found Z-Image to be amazing so far. https://civitai.com/models/2176378/afterdark-z-image-turbo
The prompt must be your problem. Model may default to Asian character, but will respond if you provide detailed description. It seems to expect more elaborate prompts. SDXL style bag of tags doesn't seem to work that well here. Try using Qwen to expand your prompts.
I think u are using some incorrect words or sth, this model can easily generate many type of women.
I never get Asian women unless I specifically prompt for them. You are clearly doing something wrong.
https://civitai.com/articles/23083/humans-of-z-image-races-cultures-and-geographical-descriptors-as-understood-by-z-image
There are lots of examples of how the model handles different racial/cultural/geographical descriptors. Z-Image may default to Asian, but I've found it very willing to give me other types of people.
@Stardeaf Забудьте про деревянный стиль SDXL, здесь нужны более подробные подсказки точные.
https://civitai.com/posts/25011970 if you mention anything like black hair, punk aesthetic, goth aesthetic, even something like "film grain" it makes the girl asian. the model is overtrained on asian girls, or they just didnt properly tag the training data. it doesn't take a long time using this model to realize this.
try to specify the race, ethnicity, type of appearance and so on, here is my implementation: https://civitai.com/posts/25006046
@thighenjoyer You may be on to something here. I think it's about the eyes shape. But forget the girl and look at the bass... This is definitely made in China.
@Snowden_Effect Try his prompt. He's on to something. It may depend on broader context.
My lora helps a lot with this issue: https://civitai.com/models/2194714/jib-mix-realistic-z-image-lora
Just a great checkpoint. I love the ability create different artistic or filtered effects. (Watercolor vs Oil vs Ink wash vs pencil vs charcoal vs film)! There are not enough "thank yous" I can offer to accurately convey my appreciation for your efforts here. Kudos!
Kudos!!!
"fits comfortably within 16G VRAM consumer devices."
Not on my 5060 Ti 16 GB, apparently. Tried to use this thing on ComfyUI; system commit increased to almost 50(!) GBs, system froze completely for over 10 minutes, had to reset.
The model itself is 12 GBS, but you also have to install that 8GB Qwen thing on the side, and that is murder for my PC.
5060 Ti 16GB,不仅可以通过 ComfyUI 使用 Z_Image Turbo 生成图片,速度比 Flux1 更快,而且还可以用于训练 Z_Image Turbo LoRA
Что за бред несешь, модель спокойно работает на 8 гигах видеопамяти
How much system RAM do you have?
I use it in Comfy on a 4070 Ti with 12GB of VRAM, and it produces very good images very quickly. You can use the workflow from this image; it works very well.
I love it when users talk about VRAM but not about RAM, vice versa
it works on rtx 3060 12GB VRAM + 64GB System RAM. I am offloading the text_encoder to the CPU (RAM). I have not tested it without offloading yet, but it should work as well, because the text-encoding stage and the sampling stage do not run simultaneously, so Comfy should offload unused models to system RAM.
I also have a 5060TI w/ 16GB, and the bf16 model & default clip works with part of the model being offloaded to system RAM.
For no noticeable loss in quality and a big boost in speed switch to an fp8-e4m3fn version of the model and a GGUF q8_0 clip; those are half the size so the clip and model stay in VRAM with plenty of extra space.
been running most models that are around 12-16 gb in size on my 3080 10gb model just fine offloading to my 32gb of ram. and not that slow generation times. not sure if its because 3080 has really high performing VRAM similar to 3090 ti
fien for me
I'm use z-image-turbo-fp8-e5m2.safetensors (5.73Gb) -- on my 4060Ti 16 Gb -- works well.
Probably you should use a simple workflow , declutter your custom nodes , maybe use comfyui copilot to debug the workflow , or just start with a fresh comfyui install... i have 3060 12gb with 32Gb ram and use this version with vae and josiefied-qwen , 2048X2048 ~120ss , 1024X1024 ~25s , no lora
I got a 5070ti and it fits fully in vram, i use the fp16 model and the fp8 text encoder, 10 sec per small img
try: --disable-pinned-memory on comfy launch.
@Amatiramisu --dont-upcast-attention --reserve-vram 150
@qek afaik upcast is only enabled on the attention layers and that is also only if you are not using xformers/pytorch/sage etc. - but I might be talking crap. - disabling pinned memory though has pretty much stopped some memory leaks / over-use, especially on WAN workflows for me, on linux as well as on windows. they added that function only recently, few weeks ago, but made it the default option pretty quickly as well, which I don't understand, but oh well.
@Amatiramisu I'll try again
There's one concept it seems to be incapable of: low angle photo shot. Now matter how much I poke it toward low angle perspective it comes out with regular front view. Anyone had more luck with it? Assuming without a lora.
Yep I tried few options but non of them worked for me :/
Include details about the sky/ceiling in your prompt
@FourByteBurger So far I've got front shot of the subject and an 'upright ceiling' behind her head... But thanks for the suggestion I'll explore this lead more.
I've heard that if you translate your prompt to chinese it works better.
@Chimaeraic_Generation Sometimes it does, but not this time. Also I observed that Qwen LLM doesn't recognize the concept when you task it with captioning. So maybe that is the underlying reason.
I poked it around a lot for that and when i used: photo taken from below, camera pointing 90 degrees upward, face visible at the top of the frame looking down, view from below it worked.
I think it's because it's the turbo model it's a little more rigid but it can do it if you poke it right
@blank_name Thanks for the tip. So far I'm getting subject front view, but upside down... ROTFL. I'll investigate.
This is as close as I get. We need to base model to train a controlnet depth map to truly get the depth of it. The current controlnet is not good because it was trained on the turbo model which causes it to massively degrade image quality. Either that or a Lora trained on low angle worms eye views(which is very easy to do imo as LoRa training for turbo takes like 1 hour max, just takes someone to do it).
https://civitai.com/images/112965674
@fox23vang226 Yes, I can get this close too. I aimed at something like this: https://civitai.com/images/99787091 I guess we'll see when/if base model arrives. It may be the most anticipated premiere of the year.
@fox23vang226 404 The page you are looking for doesn't exist
@qek it was just a picture of a woman squatting while taking a selfie with one hand/arm from her feet.
@fox23vang226 I also meant I tried, it made some kind of front view anyway
this is the best I got https://civitai.com/images/113017090
@NNAI DAMN nice one!
@NNAI Good thinking, thank you. I started from your prompt. Gradually removing and adding things. It looks like it's easy concept to break, but works if you're careful. And it needs different elements to work together. Here's what I've got:
@Stardeaf Very nice!
Something like these two images? I posted these on a throwaway since I'm not sure that I want them on my profile. I included the prompts used as well: https://civitai.com/posts/24997736
does this work on forge?
ForgeNeo updated to run Z image. https://github.com/Haoming02/sd-webui-forge-classic/tree/neo
Forge Neo yes, i have no luck though. Basic model, basic encoder oom with my 8gb Vram. GGuf encoder, gguf model i get error(some multiplication blabla)
Get SwarmUI, the basics are super simple + it has everything you need (plugins you have in Forge and more) + constant updates + Comfy background once you become an advanced user. Any other UI is a waste of time. Just a little advice. :)
I have the model, qwen text encoder, and ae. vae and it will not work for me in forge
@Tozi_White No, get ComfyUI
https://civitai.com/articles/23234/z-image-and-generation-guide-on-neo-upscale
@qek But did you know that SwarmUI also includes the entire ComfyUI? Having Comfy alone gives you fewer options than having SwarmUI, which is even better for beginners.
@Tozi_White Who am I to command, right? At least it isn't SD WEBUI
thanks all for the comments @Tozi_White thank you , so swarm is like comfy but for lazy people like me? I haven't tried swarm before, does forge neo or swarm requires to relocate all the models and loras to it's new folders ?
@_Tigerman_ Thank you! can I just update my current forge or I need to install the neo and move all models and loras to it's new directory ?
@Tozi_White ah ok , thanks for the advice! so swarm it is, any links to installation?
@qek Never. ComfyUI is for programming, not creativity.
@bigdog11 ForgeNeo is a new branch from Forge. I would install it separate 1st, test it out then move your models/lora if you want. I used the install instructions on the ForgeNeo github page. Some UI's have a home setting, so you can point to all your models and lora, because you don't want to duplicate GB's of stuff between say comfy and forge I use one directory structure shared between comfy and forge. I think in SwarmUI click on server->server configuration near the top of the page. This section allows you to set where all your models are lora are stored. If you are comfortable with more advanced windows command prompt stuff, google how to use the "mklink" command and the "subst" commands. Using these you can make new drive letters and/or directories that point to data in another directory. So for instance you can make a central models directory like D:\Ai\Models\ and make the models directory in your favourite UI point to the shared models directory on another drive. In the forge neo webui-user.bat I have this as part of the start command line
set COMMANDLINE_ARGS=--lora-dir C:\AI\Lora that way I have a central store for lora. I can also as a work around use the mklink command to make a directory say for instance in comfyui model\lora, and make that lora directory point to the contents in C:\AI\Lora without duplicating it.
So if i post a image with overswollen breasts and big clit of a middle age woman it gets blocked to get moderated by mods but if others posts lolis or things that make you puke your guts out they say nothing...get your AI $Hit to do a better job please...this is why i stopped paying ( i wanna pay but my things get deleted and other posts not)
Hmm, the question is whether that makes you better than people who post lolis...
@VisionaryAI_Studio well ...i think not , but i wanted to post that big tit girl ... it was funny and creepy at the same time... it reminded me of SD :D
Real, I reported some cubcon to see how the mods work, it didn't get removed
@qek i have no idea what that is but i think the mods have their hands full (i don't know with what) and they are trying really hard...i think :) i deleted the images even the ones that were posted (the one used with a lora but the one without lora was too "real" ) oh well , shitposting goes on ...
I honestly don't know what you're posting here that's causing your renders to be taken down. I post a lot of NSFW stuff and there's zero waiting or removal.
@Tozi_White "there's zero waiting", no, my images appear after big delays, up to 90 minutes
@Tozi_White none of your images get the yellow tringle with a ! in the upper corner of the image ? lucky you ...i tried to post yesterday and all of them were waiting for moderation...they had prompt , checkpoint and workflow embedded ... a few hours ago they were all still waiting for moderation and i deleted them ( i don't want to get banned for a curvy woman with big tits in thongs)
does this explode 8 gb vream cards?
not a bit! I'm getting 2 image generations at a time in about 70-80 seconds. The nature of the model means those two images are very similar, but sometimes there's a clearly better output between them. Big differences require changing prompt or clever noise injection techniques.
@eldritchadam thanks
@Only_Fuuka_OF also worth noting, I'm generating fairly large - 1024 by 1360. 9 Steps. Basically the recommended default workflow published with the model, added the SeedVarianceEnhancer custom node which helps a bit with variation. But all-in-all, this model is loads of fun to play with.
@eldritchadam I'm using FORGE so i have little more than steps, cfg, resolution, vae and enconder
@Only_Fuuka_OF Forge does not work, need Forge Neo
@qek ah, won't use it then
i loaded the zimage template from the latest comfyui ran in lowram mode and switched model to a fp8 version, it works, 25secs x pic. edit, the e5m2 version
@4070444890 all good but i don't wana give up normal forge
@Only_Fuuka_OF give up!
@qek NEVAAAH
it's exploding here on clip node with 12 vram
@instahentaigames2111 🔥👨🚒🚒
@Only_Fuuka_OF Neo is 100% same, like Forge. Plus has support for new models. Memory managament is the same, what is good for sdxl, but not for models with huge encoders.
@kholsaksham533 idk about it being the same, went from a111 to forge and suddenly the exact same prompts and settings was giving out diferent stuff so I'm pretty on the fence
@instahentaigames2111 I'm running it perfectly well on 12GB.
do yourselves a favor and install stability matrix and use it to install comfui, if your not using comfyui you are causing yourself all kinds of headaches for no reason trust me , comfyui is what newer models are designed to work on, its what the model designers themselves use and its what gets the most up to date support. using forge or a111 means your stuck using the same internal workflow for every project that was created by someone else and your at their mercy as to when and if you get updates to it and that simply doesnt work for every project, sometimes you need to use a different method to achieve the results you want and thats what comfui allows you to do, comfyui seems complicated at first but play around with it a bit and it becomes super intuitive, it has templates for every major model built into the menu that work out of the box and then you have the freedom to combine and use parts of other workflows at will, learning to use comfyui gives you so much freedom to rearrange and build workflows that do what you want them to do, i switched to comfui so i could run a model that wasnt available on other platforms at the time and since i have learned to use it everything else seems generic, i have built workflows for specific use cases and saved a huge library that is ever evolving using bleeding edge technology as it gets released. when people say "until its available for forge/a111 im not using it" dont realize that its not that the models arent compatible its forge/a111 that aren't compatible and specific support for new models has to be written for them leaving you always behind the curve , by using comfyui you get day1 support for most if not all models. another exciting part about it is coming here to civitai and finding premade workflows for all kinds of new features and plugins built and ready to go. comfui is the standard and most everything else is a pretty gui slapped on top of what is essentially comfyui workflows running in the background. in basic form all of these models start with the same basic components
a model loader
a text(clip) loader and encoder
a vae loader and encoder/decoder
some random noise in the form of an image
a sampler that takes all of this and makes stuff out of it
load up a simple template and youll see its really not very hard to customize your own workflows
1 load a model connect it to your ksampler
2 load a text encoder and connect it to a prompt input box and then to your ksampler
3 load a vae , connect it to a vae encoder and connect a blank image to that vae encoder then connect that vae encoder to your sampler
boom thats a basic workflow
most models and workflows follow this same principal, its really not that hard trust me
your model is the brain, your text encoder is a translator for human language to ai language and vae is the translator for pictures to ai language , all of this gets sent to the sampler in ai language (latent data) and it draws an image
in essence its converting words and image data into a format that can run on your gpu then using the sampler as a calculator to crunch the numbers on your gpu
if i can do it, anyone can do it trust me
for anyone who has taken the time to read this long message i challenge you to install comfyui using stability matrix as your installer on windows, then load a basic sdxl workflow from the templates in the menu on the top left and go to each of the boxes it creates and give it the model that it is asking for, once you have done that you will be well on your way to using comfui with ease
@loudhippo840689 bro ain't nobody reading all that
@loudhippo840689 I am not agree with everything there
How I use this on A1111
A1111 has not been updated in over 1 year, it is dead.
My advice would be to move to Comfyui, there is an easy one click installer here: https://github.com/Tavris1/ComfyUI-Easy-Install
Use it on imagination~
You can't. ForgeNeo has support an it has a very similar interface to A1111
SD.Next just made the update to FLUX2 , Z-image , you can try that |||||| yeah , later edit ... it's slower than comfyUI and harder to get it started ...
I'm using SD Webui Forge Neo.
GitHub - Haoming02/sd-webui-forge-classic at neo https://share.google/8K8WNofZPfrPgieyF
@J1B any notebook comfyui for colab ? like TheLastBen for A1111 ?
@hakfull95539 I have used this one Click ComfyUI Runpod template with Qwen-image as the default a few times: https://console.runpod.io/hub/template/t1t5m22wdy?id=t1t5m22wdyit
is not free if thats what you are hoping, but you might be able to get a $5 referral link from somebody.
@hakfull95539 Create it yourself or fork another one
I confirm as you have already been told that with Forge-neo it works perfectly.
@J1B thank you bro
@J1B install forgeneo, it looks like A1111 but supports almost everything. Definitely dont use comfy you will spend weeks troubleshooting nodes instead of generating.
@bitzupa false
@bitzupa skill issue tbh, comfy is easy
Like others have mentioned here, Forge Neo works well, i started with comfy but moved to forge neo for ease of use.
@huggar why would one prefer the inferior one (Neo)
@J1B comfy is cancer and will make you want to redact yourself. worst trash ever invented. nobody wants to spend 3 days troubleshooting except linux clowns.
I HAVE COME HERE TO CHEW ASS, AND KICK BUBBLEGUM.... euhh wait..
Checkpoint is amazing! It's incredibly fast on the RTX 3090 with just 8 steps. The realism is very good, and I was impressed with the quality and precision of the text generation!
It also works pretty well with Forge. Do you know if there's a way to use Lora for Flux with Z-Image?
Would you mind sharing your settings? I've got a 3090 but it's been taking 15 seconds per image and not loading everything into VRAM. So I'll sit at about 14GB used VRAM with the other 10 wasted
Steps: 10, Sampler: Res Multistep, Schedule type: Simple, CFG scale: 1, Shift: 3.5, Seed: -1, Size: 896x1152, Model hash: 2407613050, Model: Z-Image_Turbo bf16, Denoising strength: 0.25, Detail Daemon: "D1:1,0,both,0.12,0.2,0.8,0.4,1.5,0.07,-0.04,0.15,1", Hires Module 1: Use same choices, Hires sampler: Euler, Hires CFG Scale: 1, Hires upscale: 1.5, Hires steps: 5, Hires upscaler: 4x-ClearRealityV1, Version: neo, Diffusion in Low Bits: Automatic (fp16 LoRA), Module 1: ae, Module 2: ZImage_qwen_3_4bText_Encoder_2168935
VERY stupid in undestanding promps, why uses 6Gb LLM clip?
Flux.2 uses a 24B LLM, its prompt adherence may suck anyway
The Qwen3 model used for CLIP is not configured as a general purpose reasoning/thinking AI. Don't think of it as an LLM, that's not what it does.
Compared to most models I've found it very good at following prompts; short tag style prompts, short natural language sentences and long word-jumbles enhanced by an external LLM all work.
If you have problems with prompting, u can use this guide by sweetmax797 to improve them. He created an app but also describes a prompt template for any LLM you prefer.
In z-image LLM is Qwen3-4B-Instruct-2507
you can even use Qwen3-4B-Thinking-2507 model for fun (without any advantages and miracles)
Pointless, and it uses Qwen3-4B by default. I downloaded an abliterated version just in case anyway
It seems it's a adopted version for image generation. Usage FP8 versions of original LLM gave mess or size mismatch errors instead pictures in my case.
They do not give extremely different results btw, but may be just variations
@qek I noticed that it follows the prompts more closely (thinking model)
I created pictures with a negative prompt a week ago. But I can not see where I can write my negative prompt today. What happens with online generation?
Newb question - Can you use turbo loras with the full Z Image version, or vice/versa - can you use non turbo loras for the Z image Turbo checkpoint?
I'm using comfyui. I'm pretty good at image to video now but surprisingly photos are harder than video so far to me.
Z-Image-Turbo has a "turbo" built in; hence the name and use of 8 or 9 steps. We don't have the base model yet; we should be getting a Z-Image-Omni-Base soon.
So any existing loras are intended for use with Z-Image-Turbo, even if people shorten that to "Z-Image" in descriptions.
@Nabonidus non turbo is Z Image De-Turbo
@qek which still isnt the full model.
@qek ...which is a modified Z-Image-Turbo intended for use training loras for Z-Image-Turbo.
I'm doing something wrong, 8 steps of this model takes the same wall clock time or more than 25 steps of illustrious
Check that you're not losing time swapping models in/out of VRAM (use GGUF model/clip if needed) and make certain cfg is set to 1.0.
If you're confident with python environment you can get some big speedups with sage attention and torch compile.
It's still going to be slower than an optimized illustrious, but it's far more capable... and it's an order of magnitude faster than FLUX.2 on most consumer hardware.
What's the program?
LLM text encoder consumes most time and memory. Try FP8 one. Unet has nothing special in fact.
Or try Lumina. This model is tuned towards anime and in fact it's similar to Illustrious (2B model); Z has the same ideas as Lumina (usage LLM as a text encoder) .
@mphobbit What about Cosmos Predict2 2b t2i?
@qek Hadn't tried, but outputs from its civit gallery look interesting. It seems that community doesn't not support it comparing with Lumina.
@mphobbit I still have Cosmos Predict2 2b t2i, got some images looking like real photos, but I should dedicate some time to try again to see more
@Nabonidus I'm using Q6_k gguf. Comfy is running with --highvram to not unload models, also I'm using --use-pytorch-cross-attention because I gave up installing sage attention. I am able to use torch compile to cudagraphs, (I also gave up installing triton/inductor), will see if I this improves time when I ask for multiple images.
From what I looked it is expected to be slower, but with better prompt understanding?
@qek swarmui/comfy
Did some tinkering and got faster results, only some seconds faster than Illustrious, but maybe can help others.
What I found out is that the gguf text encoder and model versions execute slower, maybe because of how SwamUI/Comfy handles them, or (most likely) because of my processor architecture, which is AMD Strix Halo.
With the gguf version I was getting 35 seconds for 1024x1024 image. With safetensors I got 25s, and after adding --bf16-unet --bf16-text-enc to comfy initialization I am getting 22s. I'm using torch cross attention, and no torch compile.
Illustrious here executes 25 steps in about 25seconds, so z-image is now 10% faster.
For others with strix halo, this is my current Comfy startup parameter string (I'm using windows without WSL):
```
--windows-standalone-build --fast --highvram --disable-mmap --use-pytorch-cross-attention --disable-auto-launch --force-non-blocking --supports-fp8-compute --bf16-unet --bf16-text-enc
```
@WittyBitty if ComfyUI's backend is 0.4+, it should be faster, refer to https://github.com/comfyanonymous/ComfyUI/pull/11057
Stop, "8 steps of this model takes the same wall clock time or more than 25 steps of illustrious" is almost valid
@qek yes, I just realized this, the gain appear to be easier to write natural language prompts and have texts in the image.
@WittyBitty Depends on other things like sampler...
@qek I'm using the dpm++ sde heun with SGM uniform
This is normal
@qek 5 sec (z-turbo) vs 45 sec (cosmos) look strange.
@210881175 Cosmos models aren't distilled
is this model uncensored? why ni**les and vag*na is weird?
its not fine tuned like the models youre using where they dont look weird. the fine tunes should come after the non turbo zimage model is released
It hasn't been specifically trained on adult bits. That will improve over time, with community finetunes which improve the NSFW capabilities.
@whateverr already available: LoRAs
Hey man i can say "nipples" and "vagina"
And nothing will happen to me :)
@210881175 p*nis
It's uncensored, which means it doesn't actively fight you. That's not the same as actually being trained on NSFW content, but compared to other base models it does nipples and vaginas better than any I can think of. Just don't ask it to make a penis.
Because it's not fighting NSFW generations it's easy to make loras to fill in the bits it doesn't know, as opposed to a model like SD 3.5 which was so badly censored it couldn't generate a woman laying on grass or many other SFW images.
The censorship comes from the text encoder, and there isn't one without censorship yet.
@asukaylink663 good to know. thanks
Great model, but it delivers slightly better results when using an 8B LLM as the text encoder.
I get an error when I switch my text encode from Qwen3-4B-UD-Q6_K_XL.gguf to Qwen3-8B-Q8_0.gguf (4B-8B)
RuntimeError: Given normalized_shape=[2560], expected input with shape [*, 2560], but got input of size[1, 128, 4096]
They are not compatible.
I don’t use the Comfy GGUF loader; there’s another one. Just use the Comfy Manager search function and look up “GGUF.” I ran into the same error.
@Mayer2003 What specific gguf downloader do you use?
model: https://huggingface.co/DavidAU/Qwen3-8B-Hivemind-Instruct-Heretic-Abliterated-Uncensored-NEO-Imatrix-GGUF?not-for-all-audiences=true
comfy node: https://github.com/calcuis/gguf
@Mayer2003 thk
@byvalyi People say something strange with this "calcius gguf" node - https://github.com/city96/ComfyUI-GGUF/issues/230#issuecomment-2704748235
https://github.com/city96/ComfyUI-GGUF/issues/256#issuecomment-2849342007 And suggest, may be better than using this "Calcius Heretic text ecoder" to use Qwen3 8B from here - https://huggingface.co/ggml-org/Qwen3-8B-GGUF
@xNzX Sounds good, I’ll give it a try, thanks.
I cant get it to work, i use all the huggingface files correctly but it says my gpu sucks, but my gpu is decent and has enough vram.
What's the repo?
try to use zImageTurboAIO
https://civitai.com/models/2173571/z-image-turbo-aio
(take workflow from sample images)
I get a long size mismatch error for a buncha things.
size mismatch for noise_refiner.0.attention.qkv.weight: copying a param with shape torch.Size([11520, 3840]) from checkpoint, the shape in current model is torch.Size([3840, 2304]). size mismatch for noise_refiner.0.attention.out.weight: copying a param with shape torch.Size([3840, 3840]) from checkpoint, the shape in current model is torch.Size([2304, 2304]). size mismatch for noise_refiner.0.attention.q_norm.weight: copying a param with shape torch.Size([128]) from checkpoint, the shape in current model is torch.Size([96]). size mismatch for noise_refiner.0.attention.k_norm.weight: copying a param with shape torch.Size([128]) from checkpoint, the shape in current model is torch.Size([96]).
etc
@marcyyyy updated comfyUI ? if yes , best you use a fresh copy of portable , it's faster than trying to find out where all the bugs are , just save checkpoints , lora and workflows before
Which GPU, and what is the actual error?
@marcyyyy yes, get latest ComfyUI
I have a problem: ComfyUI won’t let me use this model as a checkpoint, only together with Stable Diffusion. I tried using https://civitai.com/models/372465/pony-realism and connected it as a LoRA, and then it works.
Maybe I did something wrong? Can I use only this model as a checkpoint? And is it normal that as soon as I move away from a portrait-style image, it immediately starts generating anomalies and two or more characters?
"I tried using https://civitai.com/models/372465/pony-realism and connected it as a LoRA, and then it works."
?? as a LoRA?
In ComfyUI, Z-Image safetensor model file is located in: ...\ComfyUI\models\diffusion_models
and Z-Image GGUF model in:
...\ComfyUI\models\UNET
Yes, you should use this model as only one checkpoint. Here is basic workflow - https://www.imagebam.com/view/ME19694D You can upload this PNG image in ComfyUI and get a workflow.
@xNzX thanks for help
That's not a LoRA model, if you don't already know you should understand the difference between checkpoint and LoRa.
Enjoying this one quite a lot, sharper than other models and it really got that realistic pro photo feel to it (like studio photos).
Didn't download it before because it is 'older' than other ZIT models but apparently not worse :)
impressive
I have a question. :
Who has ever been able to run z-image on GTX cards like 1070 or 1080 ti ???
confyUI versions newer than 0.3.45, I can't install - either a DLL error, or writes update nvidia driver, but the driver is fresh.
A fresh forge-neo will not be installed either...
comfyui installs the newest torch by default. I'm not familiar with your cards but I found using python 3.12 and torch 2.7.1 works the best for me.
You could try this for a fresh comfy install:
go to a folder where you want to install comfyui and open a command line
git clone https://github.com/comfyanonymous/ComfyUI.git
cd ComfyUi
python -m venv .venv
.venv\Scripts\activate
pip install "torch<2.8" "torchvision<0.23" "torchaudio<2.8" --extra-index-url https://download.pytorch.org/whl/cu128
pip install "xformers<0.0.32" --index-url https://download.pytorch.org/whl/cu128 --no-deps
pip install -r requirements.txt
pip install "triton-windows<3.4"
Everything breaks down at this stage:
pip install "torch<2.8" "torchvision<0.23" "torchaudio<2.8" --extra-index-url https://download.pytorch.org/whl/cu128
I can't download some of the files, so the process stops.
I was trying to use ChatGPT to help me out, but after a few days I realized it wasn't up to that task at all. Then I asked Claude Sonnet 4.5 and had the solution in no time. I think I've already deleted that chat window, but you might want to try yourself. BTW I run Z-Image on a 1070 card with only 8GB of system RAM.
@zooltool how much time does it takes for one generation?
So far not bad, relatively fast, but
- it has serious fetish for ears, earrings and piercings, I can't make any hair style where ears are hidden reliably.
- if I mention 3d, everyone become a child and age control becomes very hard.
- darker images are usually very noisy.
- tails on humans only work from front view, or it grows from weird places.
- animal fur is usually messy.
- with more detailed prompt the seed variations are low, outputs are almost identical.
For people on window with 4000 series cards I recommend using KJNodes for sageattention + torch compile and staying with torch 2.7.1
"I can't make any hair style where ears are hidden reliably" I know it does not have an excellent prompt adherence, but it is not unsolvable!
"if I mention 3d, everyone become a child" I couldn't reproduce it, I made some fake 3D images with Z and I didn't even mention ages, there were adults. Proof: https://civitai.com/images/112964657
"darker images are usually very noisy" Try other samplers+schedulers and shift values. Not noisy, proof: https://civitai.com/images/115828687
"tails on humans only work from front view" Why? I couldn't reproduce it either
"animal fur is usually messy" Why? I don't have the issue
"with more detailed prompt the seed variations are low" I am aware of the issue, but I don't have the problem, my outputs with different seeds are very different, no (noise) modifications, no fine-tuned Z
"For people on window" Windows*? "To be on a window" sounds awkward :c
I absolutely love the model, I do have a challenge with the nipples, they all feel a bit postpartum. Not just by my own generations, but also based on what I see in the comments/posts below. Anyone found a solution for that?
A lora
@qek Yeah... thing is though: I work a lot with wildcards, and I do not always want to have nipples. But when they're there, they should be good. Such Lora's may exist, but it's kind of a challenge.
@qek if interested: https://civitai.com/models/2291993/zimage-better-boobs
I'm facing 1 issue: the moment I interact with comfyui while this model runs, my iteration speed drops from 5sec/it to 70+sec/it. Even just tweaking a prompt or navigating is enough to do the harm.
Anybody else experienced that?
Of course, it uses 100% of a CPU and a GPU, do you have another GPU? Have enough reserved VRAM?
You can try disabling hardware acceleration in your browser while using CumfyUI, or try starting ComfyUi with --reserve-vram 0.8 option to allow more VRAM reserve for the system. default is 0.6 on windows
ComfyUI doesn’t recompute nodes whose inputs haven’t changed since the previous run. For example, on the first generation it has to load the text encoder, upscaler, VAE, etc., but if you press Generate again without modifying the prompt or changing any node connections, it reuses the cached results and should run much faster.
A jump from 5 sec/it to 70+ sec/it is extremely unusual. This usually happens when some part of the pipeline gets offloaded from GPU to CPU/RAM — for instance, if your CLIP Loader or another model component doesn’t fit in VRAM and silently falls back to system memory. In that case even a small UI interaction can trigger a reload or re-allocation, which stalls the entire iteration.
On my setup the difference is minimal (around 2.13–2.30 sec/it), even when interacting with the UI, so such a huge slowdown strongly suggests VRAM pressure or offloading issues rather than ComfyUI's recomputation logic.
You might want to check:
- whether Z-Image or other nodes force additional models into memory,
- VRAM usage right before / after interacting with the UI,
- whether your CLIP/Text Encoder is being offloaded to CPU,
- if any background software is stealing VRAM mid-run.
哦,我也有同样的问题
so far not bad but will need to experiment more
Idk why but it seems like no matter how i tried to adjust the prompts or other parameters, the face cannot be as stable and good as i want (especially if i do latent upscale)... Maybe its because this is just a turbo model?
Because algorithmic latent upscale sucks, it makes your latent noisy and broken and it needs a high denoise value to denoise
@qek i do feel latent upscale sucks in face in realistic models but it does help on improving the details and fixing some glitches. Like if the character has very complicated nail polish, a basic 720p image tends to not perform well, but after latent upscaling it to 1080p, its gonna be much better, the only bad thing is like latent upscale will destroy the face in 90% cases. But I dont wanna spend much time on fixing hands or some details using detailer nodes, nor do i wanna use nodes like ultimate sd upscale to upscale and meanwhile fix some detail errors cuz it costs too much time.. I also tried seedvr2, the result is good but it really relies on a well-generated base image cuz it wont help fixing any detail errors. And also I tried using super-resolution model to upscale the image and then use i2i with a very small denoise to polish the image still the result is not as good as I want... Is there any good way to prevent latent upscale from destroying the face? (or in other words to have things except face polished like how latent upscale does while the face remains good? I know i can do this with mask stuff but i just wanna know if there's some quick way to do so)
@SakanakoChan Why not just using a GAN upscaler and then tiled img2img to refine?
Latent upscale usually need denoise at least 0.5 which will alter the imege. Instead use image upscale and lower denoise 0.15-0.25
Vae Decode -> Upscale image by -> Vae Encode -> Ksampler with lower denoise -> Vae Decode
It is a bit slower but allows lower denoise.
From my experience you'll need to add about 5% contrast back before Encoding.
Also this is subjective but depending on the prompt for some images I prefered the original image over the upscaled one.
If you load any of my newer images, it has an upscaler subgraph on the upper-right part, and ignore all my custom nodes.
No matter how hard i try, i can't get this model to generate a woman with a fkat chest...
It is easy
dafkuq?
Lol I can generate breasts of all sizes, just need a good prompt and a LoRA for bigger breasts
@qek you misread. It's VERY difficult to get flat chested females. Especially flat chested anime illustrated females. TRY IT
I agree, I tried "very small breasts" and "flat chest", but it always gives medium breasts results.
@wktra No, I just generated flat chested women, lol. Of course I misread "fkat chest"
@qek Fkantastic!
It skews heavily toward 40 year old roastie hags. Try this LORA Tits Size Slider - Z Image Turbo - v1.0 | ZImageTurbo LoRA | Civitai
Skews heavily towards 40 year old roastie hags. There is a lora called tits size slider. that might help.
Usefull, replace my current main model for picture generation
Censorship must be removed from the model!!!
this is the original SFW model,the creator also uploaded a uncensored version
@kuska128 That's... not correct. There is only one official Z-Image-Turbo model.
@kuska128 Realy? Could you please share a link to it?
@theally Dude, I can't even get the characters to kiss properly! And what's wrong with nipples under clothing? Who declared a holy war on them?
@zalupamirok696 Limitations in the model training data, not censorship. Wait for the Z-Image-Base release, then the community will be able to go ham finetuning it into something more capable of NSFW generation.
This is not due to censorship, but rather poor content quality resulting from an insufficient proportion of training data; the model has not undergone specialized training on NSFW-related content.
You can refer to FLUX.2 for this. True censorship means the model will ignore your prompts entirely. For example, if you prompt for "nude from head to toe", the model will simply disregard it and generate a person wearing clothes instead.
@zalupamirok696 You can try this model for NSFW https://civitai.com/models/958009/redcraft-or-redzimage-or-updated-dec03-or-latest-red-z-v15?modelVersionId=2462789
@6vidit9 The creator made a new, upgraded version: https://civitai.com/models/2242173
@zalupamirok696 "I can't even get the characters to kiss properly! And what's wrong with nipples under clothing?" I have neither of these issues even with the original Z Image Turbo
@qek It's a pity that there isn't even a single piece of evidence to support your words in your profile)
@6vidit9 The examples didn't impress me, unfortunately...)
@zalupamirok696 That's fine ig, I have generated many from that same and have loved so far.
@zalupamirok696 You need to work on your prompt ig, if you'd see people here, they have generated plenty of NSFW images.
@zalupamirok696 "It's a pity that there isn't even a single piece of evidence to support your words in your profile" Ok, I will post 100 images of kissing pairs to comply with your trolling then? I already smashed NNAI's comment, saw it? We have nothing to talk about then, I only see smiling AI generated woman on your profile
@qek It's a shame I'll never see them because you banned me, hahahaha!
@6vidit9 Controlnet doesn't count; the model can't even draw female genitalia properly, let alone male ones. And what's the reason for that? Lack of time again? Seriously, there wasn't enough time to draw a woman's pussy and nipples? How can such a powerful model draw NSFW text2img? Forgive me, but all NSFW finetunes look like the work of children with special needs. No offense.
@zalupamirok696 It is hard to expect a large company to blatantly develop a model that includes NSFW content — neither FLUX nor SD can do this, as it would arouse the vigilance of society and the government.
@theally yeah my bad, confused you with somone else.
I keep getting a Qwen3 error on Forge Neo. No idea what that means.
I just went through this you have to download the encoder and vae and put it in neo the encoder and vae is listed on the checkpoint make sure to use those
@ian538 Thanks. Cant believe I didnt see that.
@Corgblam I have the same error, I already download the encoder and vae but I still have this problem. Are you using Windows ? Linux ?
@caramelrudy801 Windows 10. And really all I did was drop them in the Text Encoder folder and VAE folder respectively, loaded them up at the same time in the VAE dropdown menu along with the Z-Image model, and it worked.
@Corgblam thanks for your answer, I read on a forum that you need to have version 2.6 of NEO and install Triton for Windows, I'll try it tonight.
Does this just not work for A1111?
Nope
get sd-webui-forge-neo if you dont like comfy
A1111 is dead
Why do you suspect updates of the GUI when it clearly says that the last commit was made 2 years ago!
Hello guys, I install Neo today with python 3.11 (before I was using A1111 with python 3.9). I would like to test ZIT checkpoints, but I have a quen3 error ? Somebody can help me ? (I am beginner, my OS is win10). Tks !
ZIT in Neo uses the UI preset LUMINA in the top left corner. You will need a z image turbo(ZIT) checkpoint. Use a 6GB model to start then upgrade if your vram allows a bigger model. Put in neo\models\Stable-diffusion. You will need the flux vae (319mb) often named ae.safetensors in your neo\models\VAE. Lastly you will need the text encoder that is called qwen_3_4b.safetensors (7.5GB) in neo\models\text_encoder. Then select these files from the top of the page drop down boxes called Checkpoint for the zit checkpoint file and VAE/ Text encoder dropdown box for the ae and qwen_3_4b files. Lastly set the slider called CFG scale to 1. ZIT doesn't use the negative prompt 1 = disable neg prompt. Write a prompt then hit generate.
@_Tigerman_ Thanks a lot for you answer.The only way that I found to use ZIT on NEO was to erase my folder, and restart a clean install using the Stable Diffusion tutorial (https://www.stablediffusiontutorials.com/2025/11/forge-neo-installation.html) now it is working fine ! BTW, what is the "switch" parameter ? and do somebody know if ther is a way to use it in the X/Y/Z plot ?
I really like ZIT. Prompt adherence is great. But with the size limit of the prompt window of CivitAI online generation is very hard. Can we make a petition for it to be enlarged?
Prompt adherence is terrible. What are you talking about? Try making a woman without hips that had 5 kids. it's nearly impossible.
@bougyakumahou your fear mongering again
@qek Ok reddit
We'll raise Z-Image-Turbo and Base to 20,000 characters shortly!
@theally Very kind, but I'd hate to affect site functionality. Really, I doubt we need more than 18k. 19k at most.
Got this thing working, and it mostly ignores prompts (especially for clothing), if they are longer than 5 words, sucks at lighting prompting, completely ignores negatives, can't get different faces, heavily tends towards unattractiveness, and gives women 40 year old mom hips after 5 kids.
Z Image Omni coming very soon!
https://github.com/Comfy-Org/ComfyUI/pull/11979
https://github.com/Comfy-Org/ComfyUI/pull/11985
not today :(
@210881175 https://github.com/huggingface/diffusers/pull/12857 from past month
As far as my research goes, this is currently the most efficient and powerful model available for local generation for anyone with less than 20GB of VRAM. It's a very good combination of efficiency and prompt accuracy. I think this is a good replacement for Flux.
edit: after testing, yes it's definitely because people like their 'fast' generation. Overall a decent model, not as in depth when it comes to finer detail and long prompts but it's good for use cases.
Hi, does anyone know which sampler is used for ZIT in the civitai generator? You can't choose one in the settings and the generation data just says "undefined".
All the Loras stopped working. Now they do nothing to the image suddenly. Using Civit onsite.
The Dev team is aware, fix soon! If you notice issues, please submit them to [email protected]
@theally thank you!
Working again.
Any chance for remacri upscale of Z-image turbo on civit?
We have upscaling in the Generator, from the Magic Wand menu on generated images.
@theally I know but it's not creative upscaling. With Pony I use to do remacri upscale which in essence increases details to an incredible fidelity, where I can choose the level of Denoise of the upscale, and then I do a final 2x on that to reach around 4k.
Its nuts that this state of the art model is free :D its on par whit aurora in image department 😎
Thanks for this amazing model 😍, and special thanks for your support and generosity 🍻💯
I tried using this model in ForgeUI in Pinokio but it keeps saying "ValueError: Failed to recognize model type!" I tried the ComfyUI layout you have for one of the images and it errors out almost immediately. HELP!?
I've been testing for a few days and it's unbelievable how fast it is in low memory (8GB). A bit dumb sometimes about understanding prompt in comparison to Qwen, but so fast.
Hi! I also have 8GB of VRAM. But it doesn't work for me at all. :'(
Great model, good prompt understanding, but, i always get blurred images, no surface details, washed out colors, skin is one solid piece, no variations, no details.....
I tried all variants of samplers/schedulers, i tried various cfg, various steps... always the same.
Skin enhancements loras simply don't do anything, so...what i do wrong?
what platform are you using?
I reckon you need to upscale
hey people,
i need some help with forge ui. i am really confused about which sampler and scheduler combo to use. i’ve tried a bunch of them but i keep jumping from one combination to another every time i change models and it’s getting frustrating.
does anyone have a go-to combo that works well for the original models and the tuned ones from the community? just looking for something consistent so i can stop guessing.?????????????
This page is useful for showing which text encoders and vae you should use. It also has links for where to download them. https://github.com/Haoming02/sd-webui-forge-classic/wiki/Download-Models
The SD.Next documentation states that it supports ZIT, but there are no installation instructions. Could you tell me which folder I should put the model in so that SD.Next will see it? It's not visible in the model drop-down list, but the text "encoder" and "vae" are visible. I've tried folders Diffuserors, Unet. and also created a "Zimage" folder. When automatically downloading from the interface itself (default downloadable models list), an error occurs and the interface crashes.
SD.Next + Zluda Radeon 6800 16GB VRAM + 32GB RAM
Try ~\models\diffusion_models
Is there a unet folder, if so try it.
I used to have an amd gpu and had to switch to nvidia, sd.next was the only thing I found that worked and I was stuck to sd 1.5
Does anyone have a prompt for a realistic penis and scrotum they would care to share?Circumcised, uncut, erect or flaccid. Realistic size and appearance. Granted I don't have a lot of experience with this, but no matter what I try I end up with porn star size absurdities. I am using this model in Draw Things, but I don't think that makes a difference. Thanks in advance.
THis model censured against human's henitalia. Only ugly "pieces of something" generated - that is censoreship settings.
And so, @Civitai gives me a new reason yesterday to consider leaving them again: apparently their new politics either giving moderator right to amateur haters that will drink your blood if they could OR they now want to blackmail me by blocking on TOS w/o an actual reason (btw from same set its images that remain unblocked) and demand yellow currency to unblock others. where lies the moral line, eh? CIVITAI? you are allowing TONNS of pure disgusting porn to be hosted on your server, your images section full of this filth , But if someone posts lightly erotic theme, you play a moral 'angel' card? SHAME ON YOU!
What?
Send an email to [email protected] and I'll look into it.
Lightly erotic theme!🤡
Let me guess... banned image was from "semester exams" set? Civitai banned all "school" NSFW works as "minors in porn". All other images from this set will be banned too.
@xNzX there no minors in prompt, do homeworks
@theally i know you would review ,@theally, but what for, do you think it would make any difference? Its just easier not to post anymore, to make the haters happy.
@BzzzDarklord Who talk about prompt? Images are banned for their appearance. And you don't answer - was I'm right in my guess? Earlier, I've got banned set with enough adult female - but is school settings. So, as I understand from that, school board on wall = cause ban.
"btw from same set its images that remain unblocked" — Yes, that confused me too. But here is practice of Civitai's moderation: moderator get report for one image -- quickly ban them - and quickly goes away. HE DO NOT look on other images from this set. LATER, he look on other images - and ban them all, AND SET "warning point" to your account. And if you get 3 warning points - they simply ban your account. So, be careful. And, if ONE image from set is banned - YOU should manually remove all other similar images. THAT is a Civitai's moderation logic.
@xNzX Not planning removing ANYTHING... And don't give a flip about "harmed" feelings of some "new European" perverts practicing Sharia. Its like the idea to cover up statue of liberty because of its breasts. Nude body will remain and will celebrate its freedom. if you under 18, go get permission from your parent.
@BzzzDarklord Of course, you're free to have any opinion you want. But censorship is a matter of policy. The site will make its decision, and if there are a lot of prohibited images - they'll give you a strike point. Three strikes - and they'll delete your account. I've lost my first account with thinking exactly like you.
ok bye, you won’t be missed 🤡
Which Workflow can you recommend,? (beginner friendly), is there one with an IMG2IMG option? thanks for your time.
Here is very BASIC img2img workflow for you (with switcher to text2img) - simple, with minimum custom nodes. You can load this PNG image in your ComfyUI and get workflow (use denoise 0.5-0.6 for img2img) - https://civitai.com/posts/26510791/
@xNzX Thanks for the reference, Ill check it ASAP.
What Clip should I use?
I suppose I should also ask how to, as I cant figure out how to download the qwen_3_4 thing from the workflow I saw.
download from here the version you need, you can try all just make sure difussion model would be fp8 if you got fp8 text encoder, if missing vae, or other stuff ,you can find under same parental folder on HF text encoders
p.s. personally strongly recommend bf16 version
how can i avoid asian forever?
it always causes me problems
use 'caucasian'
shocking i know.
@TheP3NGU1N thanks
@theally are custom ZIT models ever coming to the onsite Generator? or would the demand be too much for the GPU's?
Why aren't my images uploaded? All images are marked "Not published." Is there a way to fix this?
some of mine get deleted if they have the yellow triangle on top
usually image uploading is not equal to image posting on civitai, images you posted should appear under your profile, it takes some time to civitai to account those for 'images to display' or 'images to limit their visibility', sometimes process requires a moderator review to approve, if automatic bot detected key words /key expressions it did not like, once became 'images to display' you will see those under your images section
A wonderful model, thank you very much!
I'm not sure if I should recommend this model to anyone. It's addictive.
I love you so much! Thank you <3 Hi everyone! I am so happy for Z-Image!
I'm using forge neo, and ZIT works just fine, but when I try to use ZIB, I get the error message: "RuntimeError: The size of tensor a (1280) must match the size of tensor b (160) at non-singleton dimension 1"
Ask ChatGPT - he knows. As example, here - https://gpt-chatbot.ru/gpt-5 Ask in your native language.
bad. needs too much resources to do what other models can do better
I was just checking which is the best-ranked model, and yours came up first. After installing it, I realized people were right—the realism and image quality are amazing. I’ve tried different checkpoints, but none feel like this one. It’s so addictive; every time I see a photo of an Instagram model, I think I can recreate it the same way. There are only minor skin realism issues, but the LORAs are doing a fantastic job. Please keep up the great work—much appreciated.
Creator: Tongyi-MAI (also referred to as Tongyi-MAI Lab) is an AI research lab under Alibaba Group / Alibaba Cloud, part of their broader Tongyi (通义) AI ecosystem
@Addictedd are you telling me to download that one?
@prabas029548 No, just telling who is behind the model :) what you downloaded here on (civitai) is correct one, no worries.
@AddicteddGot it, thanks for clarifying! It’s great to know Tongyi-MAI Lab is the team behind it. I’m glad the version I found on Civitai is the right one—really appreciate the reassurance. Amazing work by them! Hopefully in the future we’ll be able to create 1080p AI videos on 8GB VRAM in just 40 seconds.
just checked the videos section how can i make the same using which tool / workflow are they using for video genration
if you have comfyui usually you can click and drag the image from civitai to the comfyui workplace and it will show the workflow used to make the image. I have a few
Is there any specific reason i cannot generate nsfw images? No matter the prompt the model always presents clothed charaters.
Of course. This model doesn't have any naked images how do you expect it to generate naked images?
This checkpoint does have nude images in its training data, and it is perfectly capable of generating semi-nudes and nudes. Sex scenes is another matter though.
You can try some very basic prompts (e.g.: "photo of a nude woman") on the civitai generator, and you'll see ZiT works fine as is.
Other online generators often have nsfw filters in place which I guess might be your issue.
@oneeyed2 Thank you, i was going for nudity, not action. Yet the prompt did not yield ever partial nudity, despite several attempts.
@Iceisnice88 Were you trying on civitai or locally? What was the exact prompt? Any lora used? Are you using the regular civitai (.com) or the .green one?
Here is a test image generated on civitai: https://civitai.com/images/126099962
As you can see, it does work. Try and remix it.
A wonderful model
what sampler and schedule type do I use for the best results
Sampler: res_multistep
scheduler: simple
(these setup is fast and outputs nice quality, also dpmpp_sde/beta setup is good as well)
Deis/deis_2m_ode/beta is great for realism and skin texture.
Can it run on 8gb vram
Yeah google or bing "z-image turbo gguf"
RuntimeError: ERROR: clip input is invalid: None
help?
get a clip model like qwen 3.4b
https://huggingface.co/Comfy-Org/z_image_turbo/resolve/main/split_files/text_encoders/qwen_3_4b.safetensors
I use this model with Forge Neo.
I like the quality and, of course, the speed.
Unfortunately, the output variations aren't very good...see my article on this topic: https://civitai.red/articles/28925/z-image-turbo-seed-variations
This is an extremely high-quality model. It's a pity that it doesn't support NSFW content; otherwise, I would play it constantly. If possible, I hope to see an upgraded version in the future that incorporates new training datasets—specifically, high-quality nude photography and images featuring spread genitalia.
Try Moodyporn mix v10. Its based on ZIT and IMO does the best NSFW (without needing loras)
nice
Doesn't seem to work for me.
E: Figured it out. I dun a dumb.
Compared to KLEIN or ZIB, ZIT has lower prompt compliance, resulting in more randomness. It produces images quickly, uses relatively little VRAM.
I know this may a stupid question but how can I use vae, after I choose this vae it went back to Automatic and render only grey picture. Also where to put file text encoder? Thank u
X:\ComfyUI-aki-v2\ComfyUI\models\vae for vae;X:\ComfyUI-aki-v2\ComfyUI\models\text_encoders for txt
@fastdances Thanks ^^
😥😥😥
RuntimeError: ERROR: clip input is invalid: None
If the clip is from a checkpoint loader node your checkpoint does not contain a valid clip or text encoder model.
File "E:\ComfyUI\ComfyUI_windows_portable\ComfyUI\execution.py", line 534, in execute
output_data, output_ui, has_subgraph, has_pending_tasks = await get_output_data(prompt_id, unique_id, obj, input_data_all, execution_block_cb=execution_block_cb, pre_execute_cb=pre_execute_cb, v3_data=v3_data)
^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^
File "E:\ComfyUI\ComfyUI_windows_portable\ComfyUI\execution.py", line 334, in get_output_data
return_values = await asyncmap_node_over_list(prompt_id, unique_id, obj, input_data_all, obj.FUNCTION, allow_interrupt=True, execution_block_cb=execution_block_cb, pre_execute_cb=pre_execute_cb, v3_data=v3_data)
^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^
File "E:\ComfyUI\ComfyUI_windows_portable\ComfyUI\execution.py", line 308, in asyncmap_node_over_list
await process_inputs(input_dict, i)
File "E:\ComfyUI\ComfyUI_windows_portable\ComfyUI\execution.py", line 296, in process_inputs
result = f(**inputs)
File "E:\ComfyUI\ComfyUI_windows_portable\ComfyUI\nodes.py", line 78, in encode
raise RuntimeError("ERROR: clip input is invalid: None\n\nIf the clip is from a checkpoint loader node your checkpoint does not contain a valid clip or text encoder model.")
me either.. how can i use it?
@vahn94831 It's rubbish! It doesn't work!
Update your comfyui
@fastdances It's not working, there's no workflow. I hate it!
The images posted here are usually in NPG format and come with their own workflow. You download the image to your computer and drag it into your COMFYUI, and you will discover a miracle.
look like clip error. just try check your "Load CLIP" -node type- is 'lumina2'
@PixelTree I installed FP8 and it's working now, but I don't have Loras or Negative Prompt.
I use it fine in ComfyUI daily. What issues are you having?
skill issue
You wrong, it's works well. Here is my very BASIC workflow Z-image-turbo with img2img option (use denoise 0.5-0.6 for img2img) - https://civitai.com/posts/26510791/
@esstx Fk you !
Do you know how to use this model???
bruh, two totally free things cant be a scam lol. but yeah, it works completely fine.
Hello! Except for male genitalia, ZIT is the PERFECT model! I would kindly ask that future revisions include accurate male genitalia instead of just mangled lumps with hair. None of the loras out there for this purpose give me what I want. Thank you for your consideration!
I had to download a workflow then recreate it piece by piece to understand how to make this work - but it works!
It isn't Flux but its a step above Illustrious.
why it gives chinese girls, when i prompt for gipsy babes?
just describe what you want more detailed
it's from a chinese company. heavily trained on chinese datasets
chinese girls better
how do you fix the "ValueError: Failed to recognize model type!" on webui forge?
you better use comfyui
@2672497124 does it now work on webui?
@biuwqbiuwqdibu l donnot know but it works on comfyui, l find it hard to use on webui because of a lots of error, maybe you can fix it
Forge Neo, works fine.
Maybe a dumb question: I’m trying to set this up on Draw Things for Mac, but it says the VAE is incompatible and I’m unsure what to do with the text encoder. Can anyone explain?
wild how the default pic for this upload from the official site account is an asian woman with the negative prompt "blurry ugly bad asian, asian, japanese" attached, even though Z-IT doesn't use negative prompts lool.
Any ideal upscale settings?
Thanos meme "Fine. I'll do it myself"
4x_UniversalUpscalerV2-Sharp_101000_G
Hires steps > 3-4
Denoising strength > 0.15-2
Upscale > 1.5-2
Shift > 3/default/did not tamper with
Hires CFG scale > 4
***These are the preliminary settings I've been messing with and have so far generated some consistent, upscaled quality images
***Feel free to add to
Update 1: I've found lora Midjourney_Luneva_Cinematic increases image quality substantially chef's kiss
IMO best way to upscale locally is using SEEDVR2. Just make sure to downscale your image around 0.7 before processing it inside seedvr2.
Hi! Its worked on 3060ti, but on 5070ti i have this errors:
NotImplementedError: No operator found for memory_efficient_attention_forward with inputs: query : shape=(1, 29, 32, 128) (torch.float32) key : shape=(1, 29, 32, 128) (torch.float32) value : shape=(1, 29, 32, 128) (torch.float32) attn_bias : <class 'torch.Tensor'> p : 0.0 [email protected] is not supported because: requires device with capability < (8, 0) but your GPU has capability (12, 0) (too new) dtype=torch.float32 (supported: {torch.bfloat16, torch.float16}) attn_bias type is <class 'torch.Tensor'> requires device with capability == (8, 0) but your GPU has capability (12, 0) (too new) [email protected] is not supported because: dtype=torch.float32 (supported: {torch.bfloat16, torch.float16}) attn_bias type is <class 'torch.Tensor'> operator wasn't built - see python -m xformers.info for more info cutlassF-pt is not supported because: requires device with capability < (5, 0) but your GPU has capability (12, 0) (too new)
cool
Seems we're getting a new Z-Image Turbo competitor very soon from Alibaba in their Qwen Series line ;)
Qwen-Image-Flash: Beyond Objective Design
https://huggingface.co/papers/2606.03746
https://civitai.red/articles/31530/an-interesting-research-paper-qwen-image-flash (Wrote a quick summary for those interested). Hard to complain about a new image checkpoint seemingly coming soon :)
It used to work on SwarmUI like a month ago, but now trying to load it or a merge of it crashes the entire backend. It gives me a bunch of warnings complaining about "HybridCache" being missing and then produces this fatal error and dies:
16:45:24.192 [Error] Self-Start ComfyUI-0 on port 7821 failed. Restarting per configuration AutoRestart=true...
16:45:24.244 [Error] Error loading model 'G:/stable-diffusion-webui-1.1.0/models/Stable-diffusion/moodyProMix_zitV13.safetensors' (arch=z-image) on backend 0 (ComfyUI Self-Starting): System.Net.WebSockets.WebSocketException (0x80004005): The remote party closed the WebSocket connection without completing the close handshake.
---> System.IO.IOException: Unable to read data from the transport connection: An existing connection was forcibly closed by the remote host..
---> System.Net.Sockets.SocketException (10054): An existing connection was forcibly closed by the remote host.
--- End of inner exception stack trace ---
at System.Net.Sockets.Socket.AwaitableSocketAsyncEventArgs.ThrowException(SocketError error, CancellationToken cancellationToken)
at System.Net.Sockets.Socket.AwaitableSocketAsyncEventArgs.System.Threading.Tasks.Sources.IValueTaskSource<System.Int32>.GetResult(Int16 token)
at System.Net.Http.HttpConnection.ReadBufferedAsyncCore(Memory`1 destination)
at System.Runtime.CompilerServices.PoolingAsyncValueTaskMethodBuilde1.StateMachineBox1.System.Threading.Tasks.Sources.IValueTaskSource<TResult>.GetResult(Int16 token)
at System.Net.Http.HttpConnection.RawConnectionStream.ReadAsync(Memory`1 buffer, CancellationToken cancellationToken)
at System.Runtime.CompilerServices.PoolingAsyncValueTaskMethodBuilde1.StateMachineBox1.System.Threading.Tasks.Sources.IValueTaskSource<TResult>.GetResult(Int16 token)
at System.IO.Stream.ReadAtLeastAsyncCore(Memory`1 buffer, Int32 minimumBytes, Boolean throwOnEndOfStream, CancellationToken cancellationToken)
at System.Runtime.CompilerServices.PoolingAsyncValueTaskMethodBuilde1.StateMachineBox1.System.Threading.Tasks.Sources.IValueTaskSource<TResult>.GetResult(Int16 token)
at System.Net.WebSockets.ManagedWebSocket.EnsureBufferContainsAsync(Int32 minimumRequiredBytes, CancellationToken cancellationToken)
at System.Runtime.CompilerServices.PoolingAsyncValueTaskMethodBuilde1.StateMachineBox1.System.Threading.Tasks.Sources.IValueTaskSource.GetResult(Int16 token)
at System.Net.WebSockets.ManagedWebSocket.ReceiveAsyncPrivate[TResult](Memory`1 payloadBuffer, CancellationToken cancellationToken)
at System.Net.WebSockets.ManagedWebSocket.ReceiveAsyncPrivate[TResult](Memory`1 payloadBuffer, CancellationToken cancellationToken)
at System.Runtime.CompilerServices.PoolingAsyncValueTaskMethodBuilde1.StateMachineBox1.System.Threading.Tasks.Sources.IValueTaskSource<TResult>.GetResult(Int16 token)
at System.Threading.Tasks.ValueTask`1.ValueTaskSourceAsTask.<>c.<.cctor>b__4_0(Object state)
--- End of stack trace from previous location ---
at SwarmUI.Utils.Utilities.ReceiveData(WebSocket socket, Int64 maxBytes, CancellationToken limit) in G:\SwarmUI\src\Utils\Utilities.cs:line 312
at SwarmUI.Builtin_ComfyUIBackend.ComfyUIAPIAbstractBackend.AwaitJobLive(String workflow, String batchId, Action`1 takeOutput, T2IParamInput user_input, CancellationToken interrupt) in G:\SwarmUI\src\BuiltinExtensions\ComfyUIBackend\ComfyUIAPIAbstractBackend.cs:line 387
at SwarmUI.Builtin_ComfyUIBackend.ComfyUIAPIAbstractBackend.LoadModel(T2IModel model, T2IParamInput upstreamInput) in G:\SwarmUI\src\BuiltinExtensions\ComfyUIBackend\ComfyUIAPIAbstractBackend.cs:line 1087
at SwarmUI.Backends.BackendHandler.LoadModelOnAll(T2IModel model, Func`2 filter) in G:\SwarmUI\src\Backends\BackendHandler.cs:line 790
How well it works with faces
Can it not do NSFW?
Looks like we don't have an active mirror for this file right now.
CivArchive is a community-maintained index — we catalog mirrors that volunteers upload to HuggingFace, torrents, and other public hosts. Looks like no one has uploaded a copy of this file yet.
Some files do get recovered over time through contributions. If you're looking for this one, feel free to ask in Discord, or help preserve it if you have a copy.