I. Introduction
AnimaYume is a text-to-image model fine-tuned from Anima, a high-quality anime-style image generation model developed by CircleStone Labs. It builds upon Cosmos 2, a model developed by NVIDIA’s research team.
II. Information
For version 0.1:
This model is a preview version fine-tuned from the Anima base model using a custom dataset. Training was conducted across multiple resolutions ranging from 768 to 1280 pixels, with a primary focus around 1024. The goal of this release is to improve stability and minimize unwanted artifacts when producing high-resolution images.
Notes: All the example images at this version were generated at the resolution 1024x1536 or 1536x1024
For version 0.2:
This model is a continuation of AnimeYume v0.1. In this version, I improved the quality of my dataset and used several techniques to prevent oversaturation and low-quality outputs. Based on my testing phase, I observed that the prompt coherence is better than v0.1, and the model remains very stable when generating images at a resolution of 1536.
Note: I am still waiting for the final version of Anima and testing some methods to make my training process faster. I know the license might make the model less popular, but I only care about whether the model is good or not. I’m aware that many others use better licenses, but I’m too lazy to spend a bunch of money training a model from scratch.
For version 0.25:
This version was trained on Anima Preview 2. Due to several issues with the base model, such as overfitting, black/white borders, quality inconsistencies, and problems with artist tags, I decided to focus primarily on improving the model’s knowledge, reducing these issues, and making it as stable as possible.
Note: In this version, I did not attempt to improve the model’s style. I tried doing so, but it caused the model to forget some of its existing knowledge. The training process is similar to v0.2, but the dataset has been adjusted to better address the issues present in Anima Preview 2.
For version 0.3:
This version was trained using Anima Preview 2. It is an experiment with a new training method for the model. You can consider it as another branch of AnimeYume 0.25, developed in parallel. However, this version uses new techniques and a larger dataset compared to v0.25.
Note: In this version, I experimented with a new training approach, so the model is slightly different from v0.25. Additionally, all example images were generated using prompts shared with users on CivitAI to evaluate whether this new method.
For version 0.4:
This version was trained on Anima Preview 3 using a custom dataset. In this release, I improved prompt understanding and artist style. Based on my testing, some artist styles match my expectations, although I haven’t tested everything in detail since I’m currently quite busy :<. Additionally, I fixed several issues from Anima Preview 3 that also appeared in Preview 2.
Note: I’ve only tested with simple test cases, not comprehensively, so if you encounter any issues, feel free to let me know. I also used a larger AI computing cluster to speed up the training process :D.
All example images were generated using prompts shared by users on CivitAI, as I wanted to evaluate the model’s performance.
For version 0.5:
This version was trained on Anima Base v1.0 using my custom dataset (a mix of a small e621 dataset and Danbooru). In this release, I added many new characters and improved the existing ones. I also enhanced support for various artist styles, allowing the model to generate results that are much closer to the original styles. In addition, the model now understands some concepts and knowledge from e621, although the support is still limited.
Notes: I’ve only tested the model with a few simple test cases so far, so if you encounter any issues, feel free to let me know. This release can be considered a demo version showcasing my new training method, which focuses on preserving existing knowledge while adding new knowledge at the same time. The release also came sooner because I was finally able to use all the resources I had available :D
All example images were generated using prompts shared by users on CivitAI, as I wanted to evaluate the model’s performance using real user prompts.
For version 1.0:
This version was fine-tuned on Anima Base v1.0 using a variety of datasets, including Danbooru, e621, Gelbooru, and Konachan. The training approach differs from the original model in several ways. Most notably, I did not include quality score tags in the training data.
I also experimented with multiple captioning styles, ranging from traditional tag-based annotations to different forms of natural language descriptions, similar to the approach I used for Netayume Lumina.
Note:
This release has two versions, each trained using different methods. For the v1.0 Demo, I experimented with a mixed training approach, but it was difficult to control. The v1.0 Final is different from the v1.0 Demo, so please don't compare them directly, even though they were trained on the same dataset. I created both versions to test different ideas for training diffusion models. Since Anima is small enough, it gives me the flexibility to experiment with various training methods and see what works best.
Unlike Netayume, I did not use my full dataset for training. (The complete dataset contains around 25 million images, and if I were to use the entire dataset, I would train a model from scratch rather than fine-tune an existing one.) Additionally, I have been quite busy recently, so I have not been able to test this model as extensively as I would like. Moreover, this model dont have any default style :L
V1.0 final has a watermark in the model, which is not effect the results of generated images. Moreover v1.0 currently support chain of though prompt as you can see in my example images i did it
If you encounter any issues or have feedback, please feel free to share them with me.
All example images were generated using prompts provided by CivitAI users, as I wanted to evaluate the model's performance under real-world prompting conditions.
III. File Information
This file contains only the diffusion model and does not include a VAE or text encoder. To use it properly, you will need to download those components from the link here
IV. Notes & Feedback
This is an experimental fine-tuned release, and I am waiting for the final version release to tune it :D
Your feedback, suggestions, and creative prompt ideas are always welcome, every contribution helps make this model even better!
V. Acknowledgments
Big thanks to narugo1992 for the dataset contributions.
Credit to Circlestone Labs and Nvidia for the fantastic base model architecture.
If you'd like to support my work, you can do so through Ko-fi!
Description
FAQ
Comments (34)
Has any additional character learned compared to the base version?
Hi, version 1.0 was trained on the same dataset as v0.5 in dabooru but expand in e621 konanchan and gelbooru, therefore it has some knowledge of the characters in these new sites
Hi everyone, I want to clarify the situation regarding the watermark in AnimaYume v1.0 Final.
First of all, this watermark will not affect your generated results in any way.
Secondly, it cannot be tracked by any external tools (except by me). It is not a trigger word too, because you can actually bypass or overwrite this when merging with a high LoRA dimensional rank or multiple models. The method i will keep secret :<
So, why did I do this?
I recently heard a rumor that someone was selling my model. To be exact, they merged their LoRA with my model and sold it as their own. Honestly, claiming ownership like that is straight-up fraud. While I’m not 100% sure if the rumor is true, I needed to implement a safeguard (guardrail) just to check.
If you merge my model into yours, I am more than happy! All I ask is that you credit my model. As I’ve mentioned before, you are free to do whatever you want with it - except selling it or claiming you are the original creator. That's just rude :L
If you have any question or concern about this model please comment on this, thanks
is there any reason why I can't sell this model generated images?
@PertaliteMeister You can do it just dont sell my model :L
Is this a veiled dig at WAI? His forcefully merged model only managed to top the Anima leaderboard again by leaning on Yume's foundation and the reputation built up during the Illustrious era. Even as just a passerby, I'm furious at this kind of plagiarism of others' work. If I'm reading too much into it, please allow me to apologize. Selling stolen models is also utterly despicable.
@liu0ying No it is not :L, i am telling someone else i can not tell here
Don't hesitate. f those thieves.
@liu0ying Want to see something more dramatic and disgusting? This f**ker even dared to threaten the original creator, after being found he stole others model and put them behind paywall. He pretended to be innocent, could win an Oscar for Best Actor (see his comments in that article, and here. We had a long conversation)
He doesn't know much about AI, doesn't train the model himself. And this f***** is one of the top "creator" on tensorart.
@duongve13112002 Ah, the permissions confuse me because it says, "❌ Sell images they generate"
This new model is finally a return to form for animayume. It's insanely stable, love it so far.
Just one question about artist styles,just @ is enough or ''by'' is a must?
Hi i think @ is enough but you can try both of them to make the better result
@duongve13112002 thanks for your attention and answer!
I have a question, since you trained the model with e621 images too, how well does it generate furry/anthro characters?
Hi, it is depend on your preferences, currenlt i tested on furry and it can generate very well from my view point, but it might be different to you
@duongve13112002 Yeah, i should probably test it myself too. Can i know approximate amount of images you used from e621 to train the model? Cause most of the arts on e621 are furries
@Myrukora Hi it is about 300k
please make int8 version..
Hi here is int8 version (A finetunned version) you can try it: https://civitai.com/models/2356447/rdbt-or-anima?modelVersionId=3071010
There are a new warnings in comfyui logs with the final 1.0
[WARNING] unet unexpected: ['pos_embedder.dim_spatial_range', 'pos_embedder.dim_temporal_range', 'pos_embedder.seq']Made an int8 AnimaYume here. For ComfyUI latest hardware int8 support.
thank you king
he knows a structure (not natural language) because in the anime models he doesn't take 2 characters well, the male character comes out with a female face and he doesn't take the artist combination well
with the illustrious model I don't have that problem
yes i have this also problem
Because Anima base wasnt trained with bbox coordinates and used castrated TE instead of Qwen 3 VL
@Lynx2025 so it will never turn out well with 2 characters? are they going to improve it? because I see a lot of potential in this model
great job on checkpoints, so far im keeping two version at hand both v1.0 and v0.5.
v1.0 have over all better finish, more crisp and polished results, when when it comes to style accuracy and texture i still find v0.5 superior.
v0.4 is still the best for me. Guess I'd have to use it still. Anima v1.0 Base seems to be completely different.
Do the CLIP strength and LoRA strength need to be kept the same when using a LoRA? Thanks.
Changing the clip shouldn't have any effect; as for the LORA strength, there are usually instructions from the person who posted it. Otherwise, use a range between 0.80 and 1.00.
I understand that there are not a lot things to make but will there be animayume update? or its a final model until new anima 2.0
Hi currently not because i am focusing on experiment my new custom architecture :D. Which i have to train from scratch :v
Idk why, but Anima image generation on ComfyUI looks worse on Ubuntu/Kubuntu than Fedora Workstation and Windows 11.
Maybe because I had to use Pipx and couldn't do that python version 3.11 virtual environment command. Maybe the python version really affects it




