Qwen 3 VL 4B quantized as w4a8.
Only tested as text encoder / clip for Krea2.
Needs minimum ComfyUI v.0.31.0
Description
FAQ
Comments (12)
[Generate Text] int8 and int4 take 16 seconds, while this int8 asym 78 seconds, does it have to be so slow? upd: a ComfyUI bug
it's not int8 actually, its w4a8 , so quantized 4 bit running at 8, also are you con comfyui v0.31.0?
@TiwazM I have already updated comfy kitchen and the back end, and I already had the newest Torch for CUDA 13
@who_is_civet I don't know w4a8 support is very new and like a day old in ComfyIi. I am using it as a text encoder / clip with no issues.
@who_is_civet are you trying to run it as an LLM ?
@TiwazM Generate Text only, so the ComfyUI devs should fix it, there is nothing to do yet. Works as intended to encode text
@who_is_civet is it fixed by now? thanks
as of now didn't find any issue or any quality loss or bad outputs. works really well and the lower file size is a extreme plus. thank you
Not noticing any quality differences. Without extreme testing, it's only SLIGHTLY slower in my 8gb VRAM / 32gb RAM setup. May just keep for sheer size savings. :D
Benched aganist a FP8_Scaled version
@xFennec777 thanks for the feedback, I did it mostly for fun to see if it works ;)
Same as I did with one of my zimage turbo which is about SDXL size in total Unet and TE ;)
@TiwazM Oh, totally! That's what I am doing - I spent a lot of time making GGUF tooling for Flux Klein 2. I personally enjoy feedback on my work, happy to not give so much if I become too noisy on your work.
