My primary intent of this quant is to make it run within 32 GB RAM while still having enough memory to run some of my other tasks without performance degradation. If quality is your priority and hardware is not an issue, I would recommend seeking INT8 Convrot (~14 GB) or better
Experimental quantization of Kroma v0.2 by Lodestones, primarily aimed at reducing VRAM usage and improving usability on older consumer GPUs.
This build uses a mixed quantization approach with INT8 tensor-wise quantization and W4A4 layers.
v0.1 is kind of that sweet spot I found, not too destructive in terms of quality
v0.2 and v0.3 are further smaller is size and of course relatively lower quality
v0.2 is now up and is 0.66 GB smaller compared to v0.1
Used a simple T2I workflow, no second stage or upscaling done (I'm too lazy)
Tested on:
NVIDIA GTX 1660 Super 6 GB
32 GB System RAM
Test settings:
8 steps
CFG 1
Euler / Simple
696 × 1048 resolution
Approximately 8 s/it on a GTX 1660 Super
This is an experimental release, so performance and compatibility may vary depending on hardware and workflow.
Description
FAQ
Comments (4)
So I thought this was int8convrot since I saw the word and didn't read the full text. Then I compared it to an actual int8convrot (cicaloo from HF) and the quality is nearly identical. In the scheme of Krea 2 models whatever this Kroma 0,2 is really will not be anywhere near the top 10 models right now, so I have no idea what the point of this was, or what's the point of publishing obscure quantizations of it.
I guess this Kroma 0,2 is good for specific purposes, I just have no idea what those purposes are.
I'm kinda surprised it would be nearly identical, there should be some significant quality loss given it's about 4GB less in size? Especially the pixelated effect around the hair (found in a good number of generations by one peep in the gallery), I figured it wouldn't be there in full INT8 Convrot but didn't try that so will take your word for it.
I went with INT8 over INT8 Convrot cause I found it to be faster, at least for my GTX 1660 Super and my minimal testing. I actually tried my hand at quants with Krea 2 Turbo directly but I was curious about Kroma since I used Chroma before and I was surprised at how diverse it is, Kroma being from the same creator made me try it. Thankfully Kroma is way easier to prompt than Chroma.
As for the point of it, it's just 0.2 and very early so I would give it some time and see where it goes. I personally use it since it doesn't necessarily need loras, at least for the random image gens I like to do. Adding loras on Krea 2 Turbo increases my gen times drastically so Kroma is good in that aspect.
Very fine! v0.2 and v0.3 Interested too
Will work on getting it uploaded tomorrow!
Just keep in mind that they are just smaller in size and relatively worse in quality

