CivArchive
    Preview 128452827
    Preview 128452852
    Preview 128452864
    Preview 128452869
    Preview 128452901
    Preview 128452900
    Preview 128452904
    Preview 128452906
    Preview 128452907
    Preview 128452908
    Preview 128452912
    Preview 128452913
    Preview 128452914
    Preview 128452915
    Preview 128452916
    Preview 128452919
    Preview 128452921
    Preview 128452923
    Preview 128452925
    Preview 128452927

    ๐ŸŽŒ Z-Anime | Full Anime Fine-Tune on Z-Image Base

    Full Fine-Tune โ€ข Rich Aesthetics โ€ข Strong Diversity โ€ข Full Negative Prompt Support

    BF16 & FP8 & GGUF & AIO โ€ข Natural Language Prompts โ€ข 8GB VRAM

    ๐Ÿค— Now also on Hugging Face: huggingface.co/SeeSee21/Z-Anime โ€” including the full Diffusers folder for ZImagePipeline.from_pretrained() use.


    โœจ What is Z-Anime?

    Z-Anime is a full fine-tune of Alibaba's Z-Image (Base) architecture โ€” not a LoRA merge, but a completely retrained model optimized for anime aesthetics from the ground up.

    Built on the S3-DiT (Single-Stream Diffusion Transformer) with 6 billion parameters, Z-Anime inherits everything that makes Z-Image Base special: rich diversity, strong controllability, full negative prompt support and a high ceiling for fine-tuning โ€” now fully tuned for anime.

    This page contains the complete Z-Anime family:

    • ๐ŸŽŒ Z-Anime Base โ€” Full quality, full control, full creativity

    • โšก Z-Anime Distill-8-Step โ€” Great results in 8 steps

    • ๐Ÿš€ Z-Anime Distill-4-Step โ€” Maximum speed, 4 steps

    • ๐Ÿ“ฆ GGUF Variants โ€” Q8_0 + Q4_K_S for low VRAM / CPU / AMD

    • ๐Ÿ“ฆ AIO Variants โ€” All-in-one checkpoints (Base + 4-Step + 8-Step)

    Each main variant is available in BF16 (~12 GB) and FP8 (~6 GB).


    ๐ŸŽฏ Key Features

    • โœ… Full fine-tune on Z-Image Base โ€” not a LoRA merge

    • โœ… Rich anime aesthetics with strong style diversity

    • โœ… Natural language prompts โ€” detailed descriptions, not tag lists

    • โœ… High diversity across characters, poses, compositions and layouts

    • โœ… LoRA training ready โ€” perfect base for further fine-tuning

    • โœ… Partially NSFW capable

    • โœ… 8 GB VRAM compatible

    • โœ… All variants supported by the official Z-Anime ComfyUI Workflow


    ๐Ÿ—บ๏ธ Z-Anime Roadmap

    โœ… Released

    ๐ŸŽŒ Z-Anime Base โ€” Full fine-tune on Z-Image Base, BF16 & FP8

    โšก Z-Anime Distill-8-Step โ€” fast anime generation in 8 steps, CFG 1.0, BF16 & FP8

    ๐Ÿš€ Z-Anime Distill-4-Step โ€” ultra-fast anime generation in 4 steps, CFG 1.0, BF16 & FP8

    ๐Ÿ“ฆ GGUF Variants โ€” for low VRAM and AMD GPUs. Since CivitAI currently has no dedicated GGUF category, here is what the files represent:

    • Z-Anime-Base-Q8_0 = Pruned Model FP8 (6.73 GB)

    • Z-Anime-Base-Q4_K_S = Pruned Model NF4 (4.2 GB)

    ๐Ÿ“ฆ AIO Versions โ€” All variants with VAE + Text Encoder integrated in a single file:

    • z-anime-base-aio (BF16 + FP8)

    • z-anime-distill-8step-aio (BF16 + FP8)

    • z-anime-distill-4step-aio (BF16 + FP8)

    ๐Ÿ”ง Z-Anime ComfyUI Workflow โ€” Official workflow, supports all variants (auto-detects Diffusion / GGUF / AIO loaders, optional LoRA, optional 1.5ร— upscale)

    ๐Ÿค— Hugging Face Repo โ€” full mirror including the Diffusers folder for Python users: huggingface.co/SeeSee21/Z-Anime

    More updates coming โ€” follow to stay notified! ๐ŸŽŒ


    ๐Ÿ“ฆ Versions Overview

    ๐ŸŸข BF16 (~12 GB)

    Maximum precision. BFloat16 format, no quality compromise. Best for professional or commercial work and LoRA training. Still runs on 8 GB VRAM.

    ๐ŸŸก FP8 (~6 GB)

    Recommended for most users. Half the file size, much faster downloads. Excellent quality, barely distinguishable from BF16. Perfect for everyday use and testing.

    ๐Ÿ”ต GGUF

    Optimized for lightweight inference setups, especially useful for low VRAM, CPU inference, or alternative backends.

    ๐ŸŸฃ AIO

    All-in-one checkpoints with image model + Text Encoder + VAE integrated into a single file. Single-file convenience, no extra loaders needed.


    ๐ŸŽŒ Z-Anime Base

    The foundation of the Z-Anime family. A full fine-tune with the highest quality ceiling, the widest creative range and full negative prompt support.

    Recommended Settings:

    Steps:      28โ€“50
    CFG:        3.0โ€“5.0  (up to 9.0 possible)
    Sampler:    euler_ancestral
    Scheduler:  beta
    Negative:   strongly recommended โ€” very responsive!
    

    CFG Guide: 3.0โ€“5.0 is the sweet spot for balanced quality and creativity. 5.0โ€“7.0 gives tighter prompt adherence. 7.0โ€“9.0 is for maximum control โ€” watch for over-saturation. Above 9.0 is not recommended.

    Negative prompts have full effect on Z-Anime Base. The official workflow ships with an optimized negative prompt ready to use.


    โšก Z-Anime Distill-8-Step

    The sweet spot of the family. Distilled from Z-Anime Base, delivering strong anime results in just 8 steps. Much faster than Base while keeping most of the quality intact.

    Recommended Settings:

    Steps:      8
    CFG:        1.0  (max ~1.5)
    Sampler:    euler_ancestral
    Scheduler:  beta
    Negative:   limited effect
    

    CFG Guide: Runs best at CFG 1.0 by design. Small nudges up to 1.3โ€“1.5 are possible for slightly tighter prompt adherence. Do not go above 1.5 โ€” artifacts may appear.

    Negative prompts have limited effect at this distillation level. Use ConditioningZeroOut (included in the workflow) instead of writing a full negative prompt.


    ๐Ÿš€ Z-Anime Distill-4-Step

    The fastest Z-Anime variant. Built for maximum throughput โ€” rapid prototyping, batch generation and situations where speed matters most.

    Recommended Settings:

    Steps:      4
    CFG:        1.0  (max ~1.5)
    Sampler:    euler_ancestral
    Scheduler:  beta
    Negative:   limited effect
    

    CFG Guide: At 4 steps the model has very little correction room. Stay at CFG 1.0 for the most stable results. Nudging up to 1.3โ€“1.5 is possible but increases instability. Do not go above 1.5.

    Tips for 4-Step: Be specific and front-load the most important details early in your prompt. The optional upscaler (hires fix or SeedVR2) in the workflow is especially useful here to recover fine detail.


    ๐Ÿ“ Resolution Guide

    | Use Case | Resolution | |---|---| | โญ Portrait / Character art | 832 ร— 1216 | | Landscape / Scenes / Backgrounds | 1216 ร— 832 | | Square / General purpose | 1024 ร— 1024 | | Tall / Full body / Phone wallpaper | 768 ร— 1344 | | Cinematic / Wide scenes | 1920 ร— 1088 | | High quality / Detailed portraits | 1024 ร— 1536 |

    Supported range: 512 ร— 512 to 2048 ร— 2048, any aspect ratio. All resolutions run on 8 GB VRAM.


    ๐Ÿ’ก Prompting Guide

    Natural language โ€” not tag lists!

    โœ… Good

    A young anime girl with long silver hair and golden eyes, wearing a
    traditional shrine maiden outfit with white haori and red hakama.
    She stands in a sunlit bamboo forest, cherry blossoms falling softly
    around her. Warm afternoon light filtering through the trees,
    detailed fabric shading, expressive face, calm serene expression.
    High quality anime illustration with fine line work.
    

    โŒ Avoid

    anime girl, silver hair, shrine maiden, bamboo, cherry blossom, warm light
    

    Character portraits

    Detailed anime portrait of [character], soft rim lighting,
    expressive eyes with detailed reflections, fine hair strands,
    clean linework, professional anime illustration quality.
    

    Action scenes

    Dynamic anime [scene], dramatic angle, motion energy, speed lines,
    particle effects, cinematic composition, detailed shading,
    high quality anime art.
    

    Backgrounds & landscapes

    Anime [location] at [time of day], [lighting], [atmosphere],
    Studio Ghibli inspired detail level, beautiful background art,
    wallpaper quality.
    

    ๐Ÿ”ง Installation

    Step 1 โ€” Download your version (BF16, FP8, GGUF or AIO) for the variant you want.

    Step 2 โ€” Place the files:

    Standard BF16 / FP8 models:

    ComfyUI/models/diffusion_models/
    โ”œโ”€โ”€ z-anime-base-bf16.safetensors
    โ”œโ”€โ”€ z-anime-base-fp8.safetensors
    โ”œโ”€โ”€ z-anime-distill-8step-bf16.safetensors
    โ”œโ”€โ”€ z-anime-distill-8step-fp8.safetensors
    โ”œโ”€โ”€ z-anime-distill-4step-bf16.safetensors
    โ””โ”€โ”€ z-anime-distill-4step-fp8.safetensors
    

    GGUF variants:

    ComfyUI/models/unet/
    โ”œโ”€โ”€ z-anime-base-q8_0.gguf
    โ””โ”€โ”€ z-anime-base-q4_k_s.gguf
    

    Text Encoder & VAE (for the non-AIO variants):

    ComfyUI/models/clip/
    โ””โ”€โ”€ qwen_3_4b.safetensors
    
    ComfyUI/models/vae/
    โ””โ”€โ”€ ae.safetensors
    

    AIO variants โ€” single file, no extras needed:

    ComfyUI/models/checkpoints/
    โ”œโ”€โ”€ z-anime-base-aio-bf16.safetensors
    โ”œโ”€โ”€ z-anime-base-aio-fp8.safetensors
    โ”œโ”€โ”€ z-anime-distill-8step-aio-bf16.safetensors
    โ”œโ”€โ”€ z-anime-distill-8step-aio-fp8.safetensors
    โ”œโ”€โ”€ z-anime-distill-4step-aio-bf16.safetensors
    โ””โ”€โ”€ z-anime-distill-4step-aio-fp8.safetensors
    

    Step 3 โ€” Load in ComfyUI:

    • Use the Load Diffusion Model node for the model file, a CLIPLoader for the text encoder and a VAELoader for the VAE.

    • For the GGUF versions: load the GGUF model from the models/unet/ folder, use the same CLIP and VAE files as above.

    • For the AIO versions: just use a standard Checkpoint Loader โ€” no extra CLIP or VAE loading required.

    • Or use the official Z-Anime ComfyUI Workflow โ€” it handles all variants and precisions with a built-in model switch.


    ๐Ÿ“ฆ Custom Nodes (for the official workflow)

    • rgthree-comfy

    • ComfyUI-Lora-Manager

    • ComfyUI-GGUF (only for the GGUF variants)

    • ComfyUI-SeedVR2_VideoUpscaler (optional, only for SeedVR2 upscale)


    ๐Ÿค— Hugging Face Repo

    The complete model family is also mirrored on Hugging Face:

    ๐Ÿ”— huggingface.co/SeeSee21/Z-Anime

    The HF repo additionally contains:

    • The full Diffusers-format folder (diffusers/) โ€” drop-in compatible with ZImagePipeline.from_pretrained() for Python users

    • An alternative Text Encoder by BennyDaBall โ€” Engineer V4 (full fine-tune of the Z-Image text encoder with SMART training, drop-in compatible โ€” often produces more varied outputs from the same seed)


    ๐Ÿ“ˆ Version History

    v1.0 โ€” Initial Release

    • Z-Anime Base in BF16 & FP8

    • Z-Anime Distill-8-Step in BF16 & FP8

    • Z-Anime Distill-4-Step in BF16 & FP8

    • GGUF Variants added:

      • Z-Anime-Base-Q8_0 = pruned FP8 model (6.73 GB)

      • Z-Anime-Base-Q4_K_S = pruned Q4_K_S / NF4-style model (4.2 GB)

    • AIO Variants added (all 6):

      • z-anime-base-aio-bf16 / -fp8

      • z-anime-distill-8step-aio-bf16 / -fp8

      • z-anime-distill-4step-aio-bf16 / -fp8

    • Official ComfyUI Workflow included โ€” supports all variants

    • Hugging Face mirror with full Diffusers folder for Python users

    • Optimized for euler_ancestral + beta, simple practical use across the family


    ๐Ÿ™ Credits


    Z-Anime โ€” Anime at its finest, powered by Z-Image Base. ๐ŸŽŒ

    Description

    null

    FAQ

    Comments (25)

    Seii1Apr 24, 2026ยท 1 reaction
    CivitAI

    compared to anima base model, how good this finetune ? , the fenerating time also so much longer on zib

    SeeSeeLP
    Author
    Apr 24, 2026

    It was trained for around 160,000 steps. I did not use a ready-made dataset โ€” I built my own over the course of several months by creating and curating my own images and training material.

    The training itself took about 3 weeks on two NVIDIA Tesla cards with CPU offloading XD, so I would rather not even think about the total runtime or electricity bill.

    I used OneTrainer as the base, with some custom adjustments on my side.

    Seii1Apr 24, 2026ยท 1 reaction

    thats a lot of work,u really did a good job, i was asking about multiple concept character in 1 frame, or even nsfw really in Z image isnt it censored heavly ?

    SeeSeeLP
    Author
    Apr 24, 2026

    @Seii1 Thanks a lot, I really appreciate it.

    To answer your question more directly: in its current state, the model was mainly trained on a subset of my dataset, around 15,000 images, to test what it learns well and whether the direction is right. A lot of the training material included solo characters, including some NSFW content, so it can already handle that area to a certain extent.

    Where it is still weaker right now is group scenes, very explicit content, and some specific poses. That is mainly because I did not use the full dataset yet, but only a selected part of it for this first version.

    The dataset included around 90 different characters, although I do not have a final exact list yet. Also, since the dataset is based on images I created myself, some characters may not always be reproduced 100% perfectly.

    I am already working toward a V2. After generating over 1,000 images with this model myself, and also getting feedback through PMs, I now have a much clearer picture of what already works well and what is still missing. At the moment I am reworking parts of the captions, so I am not training the next version yet, but that is the current focus.

    The plan for V2 is to push it further so users are not limited to just saying โ€œanime,โ€ but can also describe a more specific look or style they want, for example something closer to a One Piece-inspired look. Of course, we will have to see how well that translates once it is actually trained, but that is the direction I am aiming for. ๐Ÿ™‚

    SeeSeeLP
    Author
    Apr 24, 2026ยท 1 reaction

    One thing I would add here: the main issue with NSFW content is not really the text encoder itself. The text encoder can still pass the appropriate embeddings to the model. The bigger limitation is the Z-Image model itself, since the original creators did not train it on that kind of data.

    That is exactly why it was interesting to me in the first place โ€” to see whether those capabilities can actually be taught through fine-tuning. We already know that this can work to some extent with LoRAs, but that is still not clear proof that the same behavior transfers equally well through a full fine-tune.

    The same question applies to multiple-character scenes. So for me, this project is also partly about testing how far the base model can be pushed in those areas.

    I am honestly very curious myself to see how V2 turns out. Letโ€™s see. ๐Ÿ™‚

    jimzlfApr 26, 2026ยท 1 reaction
    CivitAI

    really wonder how many characters does it know๐Ÿ˜since the illustrations have show many

    jimzlfApr 26, 2026
    CivitAI

    really wonder how many characters do it know๐Ÿ˜since the illustrations have show many

    heatwoodzachary763May 1, 2026ยท 1 reaction
    CivitAI

    i am looking for were to put the negative prompts but can not find it ,i kind of hoping it just under my nose, but it eludes me.

    SeeSeeLP
    Author
    May 1, 2026

    If you're using my workflow, the negative prompt node is actually hidden behind the positive prompt node and collapsed. Just drag the positive node to the side to reveal it. Once you expand it, you'll see one of my standard negative prompts inside! :-)

    heatwoodzachary763May 1, 2026ยท 1 reaction

    @SeeSeeLPย many thanks

    DocueiMay 1, 2026ยท 1 reaction
    CivitAI

    I've always wanted something similar Illustrious/WAI in Z-Image so I can finally get proper scenes going with actual prompt adherence vs those dumb models that cannot distinguish actor positions.

    yokinarudesu351May 4, 2026

    if you want that use anima preview 3, this one is very limited, the characters are always posing the same very neutral, anima is way more dynamic and varied in its generations

    DocueiMay 4, 2026

    @yokinarudesu351ย doesn't handle multiple characters and multiple positions well.

    z1145May 2, 2026ยท 1 reaction
    CivitAI

    Z-Anime Distill-8-Step=Z-Anime Base+Z-Image-Fun-Lora-Distill-8-Steps-2603๏ผŸ

    z1145May 2, 2026ยท 1 reaction

    Can the anime character Lora I trained on ZImage base be used on this model?

    SeeSeeLP
    Author
    May 2, 2026

    Yes and yes ๐Ÿ˜Š

    heatwoodzachary763May 3, 2026ยท 2 reactions
    CivitAI

    i got it working i had to turn of sage attention and use quad cross attention instead.

    SeeSeeLP
    Author
    May 3, 2026

    Awesome, Iโ€™m really glad you got it working!
    And thank you so much for all the feedback and the extra info โ€” Iโ€™m sure this will also help other users who run into the same issue. ๐Ÿ˜Š๐Ÿ‘Œ

    anon101May 4, 2026ยท 1 reaction
    CivitAI

    does this understand physics and anatomy as good as og zit? m away from gpu atm. cannot test myself for now. asking author & other testers. thnx in advance

    SeeSeeLP
    Author
    May 4, 2026ยท 1 reaction

    I wouldnโ€™t say it is worse than the original Z-Image in terms of physics or anatomy. From my own testing, it keeps the general understanding of the base model quite well, but the output is more tuned toward anime / illustration aesthetics.

    So anatomy should still be solid, especially for anime-style characters, poses, and compositions. Of course, as with most image models, very complex hands, extreme poses, or unusual physics can still need a few retries.

    feedback from different prompts and workflows is always very helpful. Thanks for asking! ๐Ÿ˜Š

    DocueiMay 4, 2026ยท 1 reaction
    CivitAI

    When you recover from your huge power bill,

    I am hereby humbly requesting for extra blowjob/paizuri imagery in the training data for us plebs out there. The doggystyle outputs are pretty good so thank you lol.

    tigerlizardboyMay 5, 2026
    CivitAI

    Hey! I've been fine-tuning Z-Anime with OneTrainer for furry/NSFW and I'm getting beautiful aesthetics but broken structure CFG above 1 makes it worse instead of better. Did you use min_SNR_gamma during training? And what timestep sampling worked for you? I'm on logit-normal right now and wondering if that's part of the problem. Any tips on the training recipe would be hugely appreciated ๐Ÿ™ My traning runs were garbage.

    rivdemon1221554May 11, 2026
    CivitAI

    Is this just a low-step or distilled version thing, but you can clearly see the z-image artifacts real bad in these results. Like, people say the qwen grid is ugly, because they use the fp8 version without knowing how to do better. This on the other hand does not go away with z-image.

    AetherlynMay 14, 2026
    CivitAI

    ๅœจไธ่ฎค่ฏ†ๅ‡ ไธชๅŠจๆผซ่ง’่‰ฒไธŽ็”ปๅธˆ็š„ๅ‰ๆไธ‹ๆˆ‘็š„่ฏ„ไปทๆ˜ฏไธๅฆ‚ Anima๏ผŒ่ฏ•็”จไน‹ๅŽๆœ€็›ด่ง‚็š„ๆ„Ÿๅ—้™คไบ†่‚ขไฝ“ๆ•ˆๆžœ่ฟ˜ๅฏไปฅไปฅๅค–ๅ‡ ไนŽๆฒกไป€ไนˆไผ˜ๅŠฟ๏ผŒๅŒ…ๆ‹ฌๅƒ็ฏ‡ไธ€ๅพ‹็š„้ป˜่ฎค็”ป้ฃŽใ€‚

    yk34Jun 1, 2026
    CivitAI

    I think there is something wrong with the negative prompt words in the official workflow provided for the base model. On the one hand, the negative prompt should directly fill in the things you want to avoid instead of "avoid something"; on the other hand, even if you change the directly provided prompt, there does not seem to be much difference in the overall picture quality.

    Checkpoint
    ZImageBase

    Details

    Downloads
    867
    Platform
    CivitAI
    Platform Status
    Available
    Created
    4/23/2026
    Updated
    7/31/2026
    Deleted
    -

    Files