I. Introduction
AnimaYume is a text-to-image model fine-tuned from Anima, a high-quality anime-style image generation model developed by CircleStone Labs. It builds upon Cosmos 2, a model developed by NVIDIA’s research team.
II. Information
For version 0.1:
This model is a preview version fine-tuned from the Anima base model using a custom dataset. Training was conducted across multiple resolutions ranging from 768 to 1280 pixels, with a primary focus around 1024. The goal of this release is to improve stability and minimize unwanted artifacts when producing high-resolution images.
Notes: All the example images at this version were generated at the resolution 1024x1536 or 1536x1024
For version 0.2:
This model is a continuation of AnimeYume v0.1. In this version, I improved the quality of my dataset and used several techniques to prevent oversaturation and low-quality outputs. Based on my testing phase, I observed that the prompt coherence is better than v0.1, and the model remains very stable when generating images at a resolution of 1536.
Note: I am still waiting for the final version of Anima and testing some methods to make my training process faster. I know the license might make the model less popular, but I only care about whether the model is good or not. I’m aware that many others use better licenses, but I’m too lazy to spend a bunch of money training a model from scratch.
For version 0.25:
This version was trained on Anima Preview 2. Due to several issues with the base model, such as overfitting, black/white borders, quality inconsistencies, and problems with artist tags, I decided to focus primarily on improving the model’s knowledge, reducing these issues, and making it as stable as possible.
Note: In this version, I did not attempt to improve the model’s style. I tried doing so, but it caused the model to forget some of its existing knowledge. The training process is similar to v0.2, but the dataset has been adjusted to better address the issues present in Anima Preview 2.
For version 0.3:
This version was trained using Anima Preview 2. It is an experiment with a new training method for the model. You can consider it as another branch of AnimeYume 0.25, developed in parallel. However, this version uses new techniques and a larger dataset compared to v0.25.
Note: In this version, I experimented with a new training approach, so the model is slightly different from v0.25. Additionally, all example images were generated using prompts shared with users on CivitAI to evaluate whether this new method.
For version 0.4:
This version was trained on Anima Preview 3 using a custom dataset. In this release, I improved prompt understanding and artist style. Based on my testing, some artist styles match my expectations, although I haven’t tested everything in detail since I’m currently quite busy :<. Additionally, I fixed several issues from Anima Preview 3 that also appeared in Preview 2.
Note: I’ve only tested with simple test cases, not comprehensively, so if you encounter any issues, feel free to let me know. I also used a larger AI computing cluster to speed up the training process :D.
All example images were generated using prompts shared by users on CivitAI, as I wanted to evaluate the model’s performance.
For version 0.5:
This version was trained on Anima Base v1.0 using my custom dataset (a mix of a small e621 dataset and Danbooru). In this release, I added many new characters and improved the existing ones. I also enhanced support for various artist styles, allowing the model to generate results that are much closer to the original styles. In addition, the model now understands some concepts and knowledge from e621, although the support is still limited.
Notes: I’ve only tested the model with a few simple test cases so far, so if you encounter any issues, feel free to let me know. This release can be considered a demo version showcasing my new training method, which focuses on preserving existing knowledge while adding new knowledge at the same time. The release also came sooner because I was finally able to use all the resources I had available :D
All example images were generated using prompts shared by users on CivitAI, as I wanted to evaluate the model’s performance using real user prompts.
For version 1.0:
This version was fine-tuned on Anima Base v1.0 using a variety of datasets, including Danbooru, e621, Gelbooru, and Konachan. The training approach differs from the original model in several ways. Most notably, I did not include quality score tags in the training data.
I also experimented with multiple captioning styles, ranging from traditional tag-based annotations to different forms of natural language descriptions, similar to the approach I used for Netayume Lumina.
Note:
This release has two versions, each trained using different methods. For the v1.0 Demo, I experimented with a mixed training approach, but it was difficult to control. The v1.0 Final is different from the v1.0 Demo, so please don't compare them directly, even though they were trained on the same dataset. I created both versions to test different ideas for training diffusion models. Since Anima is small enough, it gives me the flexibility to experiment with various training methods and see what works best.
Unlike Netayume, I did not use my full dataset for training. (The complete dataset contains around 25 million images, and if I were to use the entire dataset, I would train a model from scratch rather than fine-tune an existing one.) Additionally, I have been quite busy recently, so I have not been able to test this model as extensively as I would like. Moreover, this model dont have any default style :L
V1.0 final has a watermark in the model, which is not effect the results of generated images. Moreover v1.0 currently support chain of though prompt as you can see in my example images i did it
If you encounter any issues or have feedback, please feel free to share them with me.
All example images were generated using prompts provided by CivitAI users, as I wanted to evaluate the model's performance under real-world prompting conditions.
III. File Information
This file contains only the diffusion model and does not include a VAE or text encoder. To use it properly, you will need to download those components from the link here
IV. Notes & Feedback
This is an experimental fine-tuned release, and I am waiting for the final version release to tune it :D
Your feedback, suggestions, and creative prompt ideas are always welcome, every contribution helps make this model even better!
V. Acknowledgments
Big thanks to narugo1992 for the dataset contributions.
Credit to Circlestone Labs and Nvidia for the fantastic base model architecture.
If you'd like to support my work, you can do so through Ko-fi!
Description
FAQ
Comments (32)
Looks great if you used your massive 25 million images it would definitely be a Super Anima model. 😁
"Most notably, I did not include quality score tags in the training data."
Quick question on this: Does this mean we should remove quality tags from prompting? Is there a difference between adding them vs. not?
quality tags are relics of pony era, using it in positive prompt can create pony-like outputs
how much of % is e621?
For non-danbooru sources, like e621, for the tag based prompting, did you use the tags from the sources, or re-tag with danbooru specific tags?
Hi i used the tag of e621 from original source
Tested some of my older prompts and I think to characterize this model the best I'd say it has a flatter bias, more like cel shaded stuff and this makes it difficult to prompt for styles that requires alot of details. Details are swallowed by the model. Most surprising is that dpm 2m sde looks like euler now, lots of VAE noise that would be surpressed by the sampler and would only come out in euler (for example in the original hikaru ice skating prompt which I wrote, the ice dust with euler was very grainy with big grains that looked unnatural but typical for qwen VAE and euler, but looked sharp and sparkly with dpm 2m sde gpu. With Yume 1.0 the dpm has now just the same lack of detail in the dust as before with euler).
This is like pony 7
Unfortunately, to be honest, I feel like something is off with v1.0. Many artist tags are less effective at reproducing the artist's style.
Yes, V10 might have some issues; training LoRa on it yielded somewhat poor results. With the same parameters and dataset, V05 basically reproduced the dataset's style, while V10... not even close.
The advantage of the V10 over the V05 is that it's better on the fingers, but I'd still rather put up with this minor issue and go back to the V05.
The new version represents a massive change from start to finish. Let’s start with some of the most obvious changes: firstly, almost all the images have reverted to that ‘greyish’ look from the SD 1.5 era, before VAE was introduced; secondly, the art styles of a range of artists I’ve tested have undergone significant changes; and finally, the model’s adherence to prompts has deteriorated somewhat. This version is actually rather unstable—one might even call it something of a failure. I believe you need to redesign your training approach and content. BUT To be honest, thank you for your hard work and bold exploration; you are one of the pioneers who were among the first to experiment with the Anima model to unlock new possibilities.
What is that pfp son...
Hi all. As noted in the description, version 1.0 introduces major changes, so there might still be some oversight on my end. My current assumption is that the Anima (T5 + LLM adapter) is causing a bottleneck, making it underperform when processing new knowledge.
Interestingly, when I trained on Ideogram V4 and other models that use an LLM as a text encoder with the exact same dataset (not the one used for v1.0), Anima's behavior was very weird. Please don't worry, though! I welcome any feedback to help identify and fix what's happening here. Looks like I have another round of troubleshooting to do! :D
Just build your version of Anima on Cosmos 3 Nano and gguf it 🤣
https://github.com/NVIDIA/cosmos/commit/444d86120a57c42e527e4938eadd476c64c58062
@Lynx2025 nah if i can choose, ideogram come to my mind first :L
When you tried training your model based in Ideogram 4, You've managed to disable its safety filter?
@duongve13112002 with those safety filters 🤣
@duongve13112002 Alibaba might release Qwen Flash soon as a replacement for ZIT https://huggingface.co/papers/2606.03746
I know it might be a massive undertaking but would you ever do a NetaYume fine-tune of anima with your multi million image data set as that could completely fix anima.
I think the problem lies in the architecture and the text adapter because, according to current research, the true knowledge of text-to-image matching resides in the text adapter.
Once training matures, as long as the TE can match the adapter or the text modality, it can continue to work effectively.
However, the models you exemplified or used for comparison are all single-stream or hybrid-stream architectures.
Furthermore, Anima's understanding of text is extraordinarily strange, making it unable to have a massive amount of new descriptions trained into it like Z-image and similar models, which results in very bizarre trigger effects.
It would be highly difficult to handle if one were to train drastically different new things into it.
In reality, the internal DiT still stores the training data of the images, but the triggering mechanism is abnormally peculiar.
In the same scenario where the LLM is not fine-tuned, and practically speaking cannot be fine-tuned, the architecture used by Anima's Cosmos Predict2, combined with its current series of designs, feels insufficiently sensitive to text, and its matching and comprehension capabilities are not very good.
Am I crazy or does this not play nice with Anima Turbo?
You can try to increase Turbo Lora weight from 1.0 to about 1.3 - 1.5
it does not play nice with my loras trained on base 1.0, probably AnimaYume v1.0 is fine tuned too much (ill -> noob situation).
I've only tried genning a few images but it feels definitely worse than animayume 0.5. Something about the styles feels off and definitely different compared to the previous version. Prompt adherence also seems worse. I think I'll stick with 0.5 for now
It's alright at best, definitely not better than AnimaYume v0.5 :/
Hi,
Please note that the old v1.0 (demo) is different from this final release. As I mentioned before, the old version was mainly an experiment where I wanted to try some new ideas. This final version is based on the same concept, but it's much more reliable and polished. 😄
Why i released the dumb version because i dont want to hide anything and i can make mistake :L
Also, this v1 Final version includes a watermark. So if anyone plans to merge it, reuse it, or build upon it, please make sure to give proper credit. That way, I can easily track who is using my model.
Notes: this watermark wont be effect anything to your results
Maybe you can share your recommendation settings like resolutions that work best, CFG values range, most optimal steps, or even how to prompts properly so it can maximize the result. You know, just like a quick tips n trick.
I'm aware that there are showcases with all those settings information that I can try, but I still want to know what do you recommend
Hi! My model works best at a resolution of 1536x1536 (though you can go higher or lower).
The optimal step count is around 30. For the sampler, I prefer Euler a with a CFG scale between 4.5 and 6 (you can use CFG 1.0 for faster generation).
As for prompting, this model was trained on various formats from tags to natural language and chain-of-thought prompts. However, combining tags with natural language will likely yield the best results.
@duongve13112002 Thanks for answering :)
Oh I forgot one thing, do you have recommendation hi res fix settings?
@zetsubou i dont use hi res too much, so i dont know which is the best setting :<
@duongve13112002 Ahh.. unfortunate then. Thanks again for answering
@zetsubou Use realersgan x4 anime 6b for 20 steps with a denoise strength of 0.5, resize by 1.5x for best results



















