CivArchive
    Krea 2 Turbo — SVDQuant 2.42x Faster - Krea 2 Turbo — SVDQuant
    NSFW
    Preview 139424847
    Preview 139424901
    Preview 139424917

    Krea 2 Turbo — Activation-Aware SVDQuant W4A4 for ComfyUI

    Activation-aware INT4 quantization + SVDQuant low-rank correction for Krea 2 Turbo — ~2.4x faster than BF16, about a third the file size of the usual FP8 setup, no calibration dataset needed. Built for GPUs without FP8 tensor cores (Turing/Ampere, RTX 20/30-series), where the standard FP8 advice actually runs slower than plain BF16.

    Speed (RTX 3090, 1024x1024, 8 steps, warm):

    • BF16 (reference): 18.80 s — 1.00x

    • W4A4, no low-rank branch: 6.49 s — 2.90x

    • SVDQuant rank 64: 7.16 s — 2.63x

    • SVDQuant rank 256 / rank 256 actaware: 7.77 s — 2.42x

    The low-rank branch (the accuracy correction) costs ~9-10% of step time and barely varies with rank — going from rank 16 to rank 256 is only ~4% slower.

    Fidelity, measured, not guessed: LPIPS against a BF16 reference, 16 prompts x 2 seeds, paired stats. The activation-aware objective (branch fit against real per-channel activation energy instead of assuming it's uniform) is the biggest single accuracy win in the whole project and costs nothing at inference — same shapes, same kernels, only different numbers inside the branch.

    Links: 🔧 Custom nodes (GitHub): https://github.com/alperktt/Krea-2-SVDQuant-ComfyUI 📦 Checkpoints + full benchmarks + example images (Hugging Face): https://huggingface.co/AlperKTS/Krea-2-SVDQuant-ComfyUI

    Setup:

    1. Clone the custom nodes: git clone https://github.com/alperktt/Krea-2-SVDQuant-ComfyUI custom_nodes/krea2-svdquant Restart ComfyUI.

    2. Download a checkpoint from Hugging Face into ComfyUI/models/diffusion_models/. Start with rank256-actaware — closest to BF16 of anything in the sweep, same speed as plain rank 256. rank64 if you want the smaller file and never load a LoRA. W4A4-noLowRank for max speed/min size with no branch at all.

    3. Grab the text encoder (qwen3vl_4b_fp8_scaled.safetensors) and VAE (qwen_image_vae.safetensors) — standard for any Krea 2 Turbo workflow.

    4. Load this workflow, select your checkpoint in the loader node, generate.

    Loader node by checkpoint:

    • W4A4-noLowRank → stock UNETLoader

    • SVDQuant-W4A4-rank* (16/64/128/256/actaware) → Krea2 SVDQuant W4A4 Loader (this repo)

    Using LoRAs? Use Krea2 SVDQuant LoRA Loader. On a quantized checkpoint the stock loader would have to dequantize the weight, add the LoRA, and requantize it — losing the 4-bit format and requantizing the LoRA delta along with it. This repo's node applies the LoRA as a parallel branch instead, so the quantized weight is never touched. If you load LoRAs, use rank 256 — rank 64 loses most of its accuracy advantage under one.

    Tested on: RTX 3090 (Ampere) — works on any GPU with INT8/W4A4 tensor cores (Turing and up). Experimental, from-scratch project on ComfyUI's native comfy_kitchen backend. Benchmark reports from other GPUs welcome — open a discussion on the repo.

    Description

    Native W4A4 (activation-aware INT4) quantization + SVDQuant low-rank correction for Krea 2 Turbo, built from scratch on ComfyUI's own comfy_kitchen backend (not Nunchaku — no Krea 2 support there).

    Why activation-aware INT4:
    Weight-only 4-bit quantization (W4A16) doesn't actually speed anything up — if activations stay 16-bit, the matmul still runs on bf16 tensor cores at bf16 speed. The real gain comes from quantizing activations too, down to 4-bit, onto hardware that has native INT8/W4A4 tensor cores (Turing/RTX 20-series and up — much wider support than FP8, which only helps on Ada/Hopper/Blackwell). The convrot (Hadamard rotation) step handles the outliers that 4-bit activations create, and the SVDQuant low-rank branch claws back additional accuracy on top of that — no calibration dataset needed for any of it.

    What's included:

    • 4 checkpoints: W4A4-noLowRank (7.50 GB, fastest) and 3 activation-aware SVDQuant-W4A4 variants at rank 16/64/128 (7.60–8.30 GB) — higher rank = larger correction branch = generally closer to BF16, though not strictly monotonic at a single seed.

    • 2 custom nodes: Krea2 SVDQuant W4A4 Loader (reads the extra low-rank tensors the stock UNETLoader can't) and Krea2 SVDQuant LoRA Loader (stock LoRA loader silently skips all 224 quantized transformer layers on these models — matches only ~32/256 with no warning; this one patches correctly via a parallel low-rank branch).

    • quantize_krea2.py: reproducible script, regenerates any of these from a BF16 source in under two minutes, plus lets you produce ranks not included here (32, 256) or the INT8 (non-activation-aware-4-bit) variant if you want the higher-accuracy/lower-compression tradeoff instead.

    Speed (RTX 3090, 1024x1024, 8-step Euler, warm run):

    formatwarm runvs. BF16BF1621.3s1.0xFP8 (emulated on Ampere)19.2s1.1xActivation-aware W4A4 (no low-rank)10.1s2.1xActivation-aware SVDQuant W4A4 (any rank)10.1–10.2s2.1x

    Accuracy notes: Rank 128 was the closest SVDQuant variant to BF16 in my qualitative text-rendering tests, which is why it's included alongside rank 16 (smallest) and rank 64 (middle ground). If your use case is dense small text (menus, labels, documents), the low-rank branch narrows the gap but doesn't fully close it — large text and general composition hold up fine across every checkpoint here, including the no-low-rank file.

    FAQ

    Comments (3)

    xFennec777Aug 11, 2026
    CivitAI

    How does smaller VRAM benchmark? Like 8gb cards - is this set up in such a way to speed up instances where memory swapping kicks in?

    woaiyitiaofengAug 12, 2026
    CivitAI

    I'm getting some kind of video memory error message, but I think the current support and stability are poor.

    velantegAug 12, 2026
    CivitAI

    So this strictly requires CUDA 13?
    While workflows with KJ stil require Python 2.9 because KJ nodes incompatible with new versions of some dependencies.
    CUDA 13 requires python above 3.0.
    In simple this incompatible with KJ nodes.

    Checkpoint
    Krea 2

    Details

    Downloads
    158
    Platform
    CivitAI
    Platform Status
    Available
    Created
    8/11/2026
    Updated
    8/12/2026
    Deleted
    -