🌸💀 DaSiWa LTX 2.3 💀🌸
My new LTX 2.3 model for I2V, T2V, V2V generation.
Version overview: https://civarchive.com/articles/23495/dasiwa-model-versions-and-timeline
A comparative study:
https://civarchive.com/articles/29961/dasiwa-or-major-ltx23-model-comparison-part-1
https://civarchive.com/articles/32224/comparison-between-major-ltx23-models-part-2
Expect that not everything is perfect and mind LTX2.3 is not as stable as WAN 2.2 finetunes.
⚠️ Make sure to open the DOWNLOAD dropdown to see all quants possible.
🔮 Key Features:
🔥 Best With I2V and V2V
🧪 Optimized mixture
🔊 Better Sound
🗣️ Better Voices
🌟 Enhanced Quality and Reasoning
🔞 Unrestricted
🪄 Better Prompt Responsiveness
🥺👉👈Better understanding of anime/manga style composition
🪡 Finetuned mixed precision's
😵💫 Reduced some hallucinations
👘 Strengthened visual consistency/understanding for anime
Different versions may have additional customization (read the version notes)!
❎ Included VAE and not included VAE
🧪Distillation and non-distillation
🍒Workflow
Make sure to checkout my easy to use Workflows!
🍄LoRA's
But: This checkpoint is not meant to replace all LoRAs, it is meant to:
Perform better overall at his own
As easy as possible to use
With LoRAs to be more awesome
⚠️ Read the corresponding announcements.
📢 Make sure to check it out for in-depth information and a complex comparison!
🛠️ Recommended Settings
CFG 1
Euler_CFG_PP/linear_quadratic
8-10 Steps (distilled)
Dependencies
VAE
LTX23_audio_vae_bf16.safetensors
LTX23_video_vae_bf16.safetensors
Dual CLIP (Encoder and Projection)
gemma-3-12b-it-heretic-v2_fp8_e4m3fn.safetensors
ltx-2.3_text_projection_bf16.safetensors
🩻 Known issues
Tell me 🫵🫢
LTX2.3 be LTX2.3 🫣
Hands are sometimes unstable
Shifting of fine details (e.g. eyes) without prompting or really high resolution
Needing way more runs for good results than WAN22
Like all LTX23 checkpoints at the moment this can have 👻 ghosting dependent on the scene, motion and used Loras and workflow's settings. After all LTX23 is still unstable.
LTX23 is very dependent on the used workflow + settings!
🩺 Fixes & Feedback
If you use LoRAs, try to respect the LoRA training triggers and try some versatile descriptions, most LoRAs will work with 0.3-1.2 (start with 0.3)
Do not mass add LoRAs, just add 1 or 2
Negative prompting do not work with cfg 1, thats a limitation of speed-ups with cfg 1
Before posting any questions I suggest reading my guide.
Update your ComfyUI ❗
🖤 Why I Made This
Pushing LTX2.3 to its limits!
This checkpoint is also my personal playground.
Closing words
🤩 I want to thank all the fantastic other creators who made super nice LoRAs and concepts to play with! Support that awesome creators by using their LoRAs and post to their gallery and share the meta-data!
⚠️ I made all this with permissions or open-source resources (the time it is incorporated).
I share as much insights as I can without compromising my work. I'm doing this for fun as my hobby and just do not want my hobby to be destroyed.
More details can be obtained in the corresponding announcements!
If you would like to contribute in my awesome (😉) checkpoint or willing to share resources I'll gladly give credit! Just contact me!
✅ All credits / resources are mentioned inside the announcements! - Since different versions may have different resources.
YOU are responsible for outputs as always! If you make ToS violating content and I get aware I WILL report this.
Disclaimer
This models are shared without warranties and with the condition that it is used in a lawful and responsible way. I do not support or take responsibility for illegal, harmful, or harassing uses. By downloading or using it, you accept that you are solely responsible for how it is used.
LTX-2.3 Custom Addendum: Fine-Tune Integrity & Attribution
Base License: LTX-2.3 Community License Agreement
1. Verification & Integrity Requirement
This model is a fine-tuned or merged derivative of LTX-2.3. To ensure users receive the correct weights, safety metadata, and version updates, the Official Source is maintained at: https://civarchive.com.
Notice of Non-Support: Any versions hosted on third-party platforms (mirrors) are considered "Unverified." The creator provides zero warranty, support, or safety guarantees for unverified files.
2. Trademark & Branding Restriction (Pursuant to LTX-2.3 Section 8)
While the underlying weights are subject to the LTX-2.3 distribution rights, the name "DaSiWa [Model Name]" and any associated logos or promotional imagery are the intellectual property of the creator.
Renaming Rule: Any Entity or individual redistributing or mirroring this model on a third-party platform (including but not limited to Hugging Face, Tensor.art, or SeaArt) MUST remove the original model name and branding unless explicit written permission is granted.
Source Attribution: Redistributors must provide a prominent link back to the Official Source as the primary point of origin.
3. Commercial Platform Restriction (Pursuant to LTX-2.3 Section 2)
Commercial Entities (as defined in the base license) that generate revenue through the provision of "Generation-as-a-Service" or ad-supported hosting are prohibited from using the official branding of this model to market their services without a separate agreement.
If your platform charges "credits" or subscriptions to access this specific fine-tune, you are required to contact the creator to ensure compliance with my project.
Description
🔥 Best With I2V and V2V
💎 True Vision (no distillation)
🧪 Optimized mixture
🔊 Better Sound
🗣️ Better Voices
🌟 Enhanced Quality and Reasoning
🔞 Unrestricted
🪄 Better Prompt Responsiveness
🥺👉👈Better understanding of anime/manga style composition
🪡 Finetuned mixed precision's
😵💫 Reduced some hallucinations
👘 Strengthened visual consistency/understanding for anime
FAQ
Comments (126)
Hi! v2 seems to have slower animations, but the consistency is very strong. I also have a strange issue to report: during the generation process, I saw the preview initially perform a certain action correctly, such as lifting someone up, but in the next preview, the character remained stationary. The final generated video also shows the character remaining stationary. I don't know why.
it's because at a CFG of 1 you prompt doesn't have enough weight to keep certain actions that defy what is "common sense" and someone leaving the ground is against the model's common sense and if you're doing I2V it's also against the reference image, so your fighting against the reference image and the model's "common sense" go do a render with 50 steps and a CFG of 4 without the distillation lora and you will almost certainly get the motion you were looking for, although it's likely that if you are doing I2V that the model will also fairly significantly change the look of the characters, not that I have done a base generation yet.
probably a middle ground of a CFG of 2.5 and ramping down to a CFG of 1 for the final low sigmas stages would fix that though, but you would need to look up or ask a LLM around about's where LTX 2.3's sigmas fall is then match the CFG curve to decay during the low sigmas pass. Again I'm guessing here, the changing the CFG might very well screw up the sigmas being real.
@Etheoma ok,thx.
Any chance for fp16?)
Not for v1 and v2, at least. They are not considered stable.
For some reason, SolsticeCoin causes rather frequent BSODs on my system. No idea why is that but that's what's happening. Previous version was fine.
A checkpoint itself cannot BSOD, for sure.
It might be from moving large models from your SSD and into your System memory all the time. It's not related to this model. A BDOS will occur if some sectors of your SSD are starting to fail and windows loses something critical while doing a task.
Not a checkpoint issue, pretty sure it's related to LTX 2.3 model memory management. Try adding the argument "--disable-pinned-memory" to your run_nvidia_gpu.bat file (assuming you're using comfyui).
I downloaded a new LTX2.3 workflow (using the standard ltx-2.3-22b-dev-fp8 models) a week ago, and it would crash my system repeatedly (4090 24GB VRAM, 64GB RAM). Strange thing was that it didn't crash during the generations. It would crash after I closed the browser or the terminal. Wan workflows with larger models and even other LTX 2.3 workflows worked fine for me though. Claude AI recommended adding the --disable-pinned-memory argument to the bat file which finally fixed the issue for me.
Did you just updated your bios? amd push a new memory feature which may automatically enable but some memory stick won't be stable.
@quadrazz93101 No, I did not do any changes to my system in the past half a year. I also recently ran Memtest86 with flawless result.
@barry9000 I use SwarmUI but it's basically a ComfyUI frontend.
@PseudoGrafx the problem could be your drive, if you offload it will fill your pagefile if your pagefile get 100% and RAM you may crash.
@Darksidewalker I got 48GB RAM and over 100GB of free space on the drive where my pagefile resides. I could try freeing up some more space but not really sure it needs that much...
@PseudoGrafx That depends on the used settings how much it needs. You can monitor this with perfmon.
Update…lip synching didn’t work using Dasiwa workflow initially but once I turned off "First-Frame Anti-Burn” it works great!
it's good and fun, but not very stable for flf2v, looking forward for next version
That's the nature of ltx23
@Darksidewalker hey i just wanna thank you for all the work, everything i've used work really nice.
masterpiece
very curious as to which distill lora you merged in. condsafe? and which strength?
V2 had condsafe, v3 is not distilled
Honestly kinda wish you had left this as not distilled so that you can apply the distillation lora at whatever weight is appropriate to audio etc, like I'm sure for 80% of cases what you came up with works really well, but... yeh... kinda wish you had created a workflow instead with the lora applied the way it's applied in the model right now.
Dev version, maybe, pleas.
Although I'm wondering whether applying the lora with a negative weight might work? Which one did you use? I assume 1.1_fro90_ceil72_condsafe?
Edit: V2 and V3 are already undistilled, kinda wish the creator would change the page as it's kinda missleading.
v3 fixes all of the old bugs from v1 and v2 AND it lets you make infinite length videos on just 500 megabytes of vram
500 MB VRAM? typo?
Sarcasm?
@lolmao500 I'm AI I cannot understand. O.o
do you really get better results using heretic encoders?
I did not extensively test it, maybe not :)
The audio for NSFW is beyond awful, I hope this can be fixed in next version.
Prompting issue, I assume. Sorry to say that.
Base LTX23 is not trained in NSFW audio, and Sulphor-2 also is medium at this, this added some extra audio, so yeah....
You may to have improve your prompting for sounds and audio, it totally depends on what you write.
@Darksidewalker Thanks, will try that.
When I changed the model from version V2 to version V3 in the same workflow, the generated video had a ghosting effect, which did not exist in version V2. How can I solve this problem?
Not possible that you are using the same settings, since v3 is non-distilled, you must have changed settings.
I am also having the same issue
I'm also getting the same on either fp8 or nvfp4
Interesting, I had horrible ghosting with V2, and V3 (with the latest workflow) mostly solves the issue
civi says only UNET instead of ggufs 😭
i can guess from the weights anyways
Already submitted a ticket yesterday ✌️
我V2都没捂热就端上来全新的V3了,太牛逼了
I have been using Draw Things on my MacBook Pro M5 chip. Installing and generating videos with the LTX 2.3 v1.1 no issues.
How Draw Things install safe tensor models, the program will 'recode it into a special file for easy loading - I'm not a programmer, sorry' which makes it compatible for the program.
For your model Dawisa ltx ver3, after initial loading, during compiling the model, (a .journal file was being created) it suddenly aborts. I'm not sure whether anybody else has this issue or not. I have 48gb of Ram and 1TB of HDD, so I don't think is memory overload?
Maybe I'll install via Wan2GP instead. Comfy UI is too mind boggling for me.. Cheers, and keep up the great work!
Dude .. if v3 is not distilled .. you should mention it ! I'm pulling my hair why it doesn't work ..
Honestly, I blame Civitai's poor layout for this. The author actually put a note under 'About this version,' and the second line clearly says 'True Vision (no distillation).'
It is mentioned in the right side under About this version
Yeah it is mentioned multiple times and inside the announcement and in the timeline article....🤷
One question about the GGUF files: could you clarify which quantization level (Q) corresponds to each file size? I believe the 17GB version is Q6 — is that correct?
I did submit the info, civitai does not display, I made a ticket
@Darksidewalker Thanks
Does this model support the LTX Director node from the WhatDreamsCost project?
any ltx 2.3 model should work with it
I tried using your last workflow with this model but I get very low quality video (although very fast generation). I don't understand what should be the settings since V3 is undistilled (I used both distill and non distill option in your wf). With nvfp4 checkpoint, because I have a 5xxx card
I don't know what you mean by low quality, but nvfp4 is not high quality, it is high compression for VRAM poor.
@Darksidewalker I understand, I tried the FP8 model and it's more or less the same (lots of ghosting and blurring). Should I double the steps or sth like that because V3 is truevision ?
You have to use the settings for non distilled or use a distilled Lora
Do you mean that this https://civitai.red/models/2498991/dasiwa-ltx23-workflows-or-i2v-or-flf2v-or-t2v-or-v2v-or-audio is not adapted to V3 ? I used the distilled Lora you suggested.
@Elmer588 my workflow can run it no problem, but you have to use working settings for distilled or non distilled
@Darksidewalker where can I find the recommended settings for this?
Is there a specific distilled Lora you recommend for v3?
The one linked in my omniforge workflow
Can someone answer me?
I have used V2 for a while but eventually gone back to WAN 2.2 since the autonomy was 9 times out of 10 bad.
It most likely has to do with LTX2.3 and not the model but does anyone has advice or does this v3 one a better job?
I'm fairly new to LTX2.3 and searching my questions online is a lot of time muddled with speaking about old versions.
LTX treats input like guidelines more than literal start and stop points like WAN, you are probably not going to get as clean of a output just running it by itself. It helps if you find a style lora for LTX that matches your input - although v3 does a good job of staying fairly faithful to the original format compared to v1.
You can always combine the two as you can use the V2V option to add dialogue/audio/lipsync to a cleaner clip you generated with WAN first
@FirstPrinciples I'll try thanks a lot for this suggestion!
@FirstPrinciples Exactly this! I have tried 100's of LTX23 I2V generations in the past few months and rarely do they even come close to identity accuracy as WAN22.
I'm basing all my flows on Dasiwa's WAN and LTX workflows. But the difference is night and day with WAN.
Do you think we will ever be able to achieve clips greater than 5 seconds with WAN? For me that is the only defining benefit of LTX, with audio being a bonus of course.
Is it just the nature of WAN22 that you're always going to be limited to 5 second clips, and your only choice is to chain videos using last frame generations (using a 3rd party video editor)?
I started my AI journey with WAN2GP and it had a great tool where you could continue your video, using a starting video. I can't find that feature in Dasiwa's WAN workflow, only in his LTX workflow can I see V2V.
video i make with V3 have a black shadow/fade, on moving objects, any idea why? doesnt happen with the same setup with eros10
is this one better with keeping the face of the person instead of changing it to someone elts?
Seems better, but large motion still causes some changes. The image strength settings needs to be towards 1
your getting there man
秒变西方人的脸,感觉对亚洲面孔不太友好
Reach out on the LTX AI Team for training the base like this and if you try the others, well you will be surprised, they are even less "asian" friendly.
Do you have plans of making distilled nvfp4 Version like this one? https://civitai.red/models/2445970/ltx23-fp4?modelVersionId=2751189
I really like your style of animation, but I have only 16gb vram so I need a smaller models
Not a at this point v2 are distilled, but baked distillation has disadvantages over adding lora on different weights. So v3 are without.
Damn man you're so fast at making these models!
Love your work so much!
The workflow seems to not want to work. When I switch something off, it doesn't turn it off. I've tried everything and I'm not exactly new at this. Any advice you can give to help me with this?
Update comfyui and make sure to install all dependencies and custom nodes
Thank you so much. And your WF is great too I'm using a lot. Can you tell me the difference between the versions here in simple words ?
Thank you!❤️
In simple words, they are refined version. Newer should be better.
@Darksidewalker I see thanks ! it tends to give bigger boobs to character's that doesn't have.. but need to test further !
Not having any luck with GoldenLace. SolticeCoin is fine (I'm using your workflow, largely unmodified). I've tried the Civitai mirror of the FP8, Your hugging face mirror (which says that it's the same file as v02, renamed), and I'm currently trying a GUFF of the undistilled with high steps (20-30). Everything that is generates is a noisy mess. Can't explain why SolticeCoin would run fine, and this new one not at all.
Review all your settings, steps alone do not make undistilled work, cfg and other settings may be changed.
Or use distilled settings and a distilled lora, what I would suggest. The model works fine.
@Darksidewalker That nudge actually help. I turned on LightSpeed/Distillation and got a coherent image out. I'll cut if of and try a higher (than 1.0) CFG with 20+ steps and see what I get. It may just be that my copy of SolsticeCoin was distilled and I didn't realize it. Thanks for taking the time to reply.
good
Wow; GoldenLace looks extremely good. Well done as always. :)
would you consider releasing your models in int8 format (https://github.com/BobJohnson24/ComfyUI-INT8-Fast) it is way faster than fp8, gguf's and nvfp4 for amd gpu's and nvidia 3000 series and better quality also. It needs to be converted from bf16 so users can't do it.
I tried, but it always produced pixel clutter. I'm not sure why atm. That said, it is faster but not better quality, it is behind fp8 by any means.
If I find a way to make it, I may do it, but I'm not sure what the problem is or if the custom node is the problem.
I figured that there is no native support for INT8 in comfyui and the development from the awesome programmer silveroxides stopped the int8 development.
So the effort of working around this is just extreme high.
I might not do this, if there is no proper support.
@Darksidewalker dunno how but people did it for the sulphur and eros models, again of course bf16 is needed which you would have.
I encountered an error while loading the VAE files for Audio and Video. I downloaded the corresponding VAE files, placed them in the appropriate folders, and correctly selected the correct VAE file. Please help me.
Error message:
RuntimeError: ERROR: VAE is invalid: None
If the VAE is from a checkpoint loader node, your checkpoint does not contain a valid VAE.
It contains valid VAE if the correct nodes are used (tested in other UI's like SwarmUI), please do not spread misinformation.
The downloaded was not placed or selected correctly, or you got the wrong vae.
@Darksidewalker Thank you for your reply! I was also confused. Normally, it shouldn't throw an error; there shouldn't be any operation that could cause it to malfunction, it's just selecting an option.
I downloaded all the VAE files through the Model links in the workflow, and I also installed KJNodes, and downloaded the plugins in the Requirements.
@Darksidewalker If the downloaded file is fine, could the problem be with the KJNodes version?
@Darksidewalker The problem has been found; it's caused by the version.
Is that much different from wan 2.2 checkpoints? Seems like LTX 2.3 checkpoints didn't go far, maybe even worse at NSFW. Except audio, which is research-level quality and not ready to make good content anyway
you underestimate the fact ltx2.3 makes videos 2x longer, 24fps native and is 50% faster than wan 2.2, wan is better in the smaller actions.
What do you use for the nsfw images ? I was going through the comparison of these models. Btw, V3 is much better than the previous ones.
I use my illustrious, or Anima checkpoints
@Darksidewalker Thanks. I ll try them out. You're making LTX reach WAN 2.2 levels. Appreciate all your work with both models.
@BopStar Thank you😄
I use ZiT MoodyPro model v10 (v13 is ok but less natural penises overall)
@TheLastRemain Thankyou for the suggestion. I ll look into it. I m using Klein ovaNS W , for realism its quite good & fast. Might have to work with the prompt a little for accurate generation.
hi.can v3 works with 16gig vram and 32 ram?
Yes
Tried using this with your workflow but I seem to get ghosting problems, I tried using the original LTX safetensor and it worked well, I didn't change any settings from your workflow.
Why am I getting a thin line of semen (a thin line of semen running from the girl's mouth to her lower body) when I'm depicting ejaculation scenes? I've checked my prompts and negative prompts. Has anyone else encountered this? Or is there something I'm not setting correctly?
this model isnt great at it, but then again ltx 2.3 base thinks cum is pop rock candy for some reason lol
there's a cum lora, use that, its realistic, but lacks the ''shot'' aspect of it as far as im aware
@anth0nx for sure you did a comparison and find a better one. Would love to hear from that model 👍
I would not recommend this over sulphur.
Sulphur is quite overfitted.
I would not recommend sulphur2 over 10Eros.
There is no one-fits-all, like WAN22 this will not happen with a local model, even less likely with clunky LTX23.
So use what fits the situation best. Nobody is forcing anyone to use any model... More options are good tho.
在制作sfw内容的时候,如何避免衣服上出现乳头的问题,是不是因为我用了身体物理增强的lora
Is my assumption right that the v3 needs the distilled-lora enabled where v2 does not?
Swapped v2 out with v3, and the run was blurry and not high quality in the same workflow, enabling the distilled-lora gave good results.
edit: i got the 27GB version from hugginface.
V3 ist without baked distillation
@Darksidewalker Thanks for clearing that up!
One more question as i've been looking for it for a while now, no model or lora seem to accurately have a person take off his/her clothing in a propper manner.
Trying beach clips where people arrive and take off their t-shirt etc. is not working (cloths dropping by themselfs, strange arm movements)
Is it possible to look into that for once? lora or in the checkpoint where undressing is easier to achieve?
Against ToS, so there is unlikely a Lora for this
@Darksidewalker Ah the abuse fears alas ... lots of normal use cases like taking off the characters coat, shot at the laundro-mat taking off the sweater to add it to the wash etc... Ok, got to look for some proper prompting then and see what we can cook up.
is it the model or the lora, or is it your prompting? as i have no problem inpainting nude characters through prompting alone on the base model/no lora for it.
if you prompt like ''the woman takes her shirt off, the woman becomes shirtless'' etc, then ltx often doesnt take the hint, its the same when prompting for a blowjob, wont work when you prompt ''the woman performs oral sex on the penis'' 9 out of 10 times.
you have to guide them in steps, like - the input woman with brown hair wearing a green jacket raises her hands towards the the silver zipper of her green jacket, she uses her fingers to pinch the zipper - she pulls the zipper the zipper down that she has inbetween her fingers to unzip her coat, when the coat is open, she grabs the left side of her open coat with her left fingers and the right side with her right fingers, she takes the coat off and lets it fall to the ground.
stuff like that works fine for me, i keep it superficial like she uses her hands to unzip her coat, if that doesnt work, i go more in depth and so on.
but there's other factors aswell. if you use prompt enhance and it censors your prompt it can botch the generation.
aswell as when you use distilled cfg 1 your prompting gets impacted by it
and what also matters is sigma/steps, scheduler, etc, its a complicated mess as im a novice but ive become aware that scheduler/sampler really matters on what your trying to do, it can be totally different if you want to maintain your input image or alter it allot since with one you dont want allot of noise at the start compared to the other
Bro, I've hardly seen any behavior LoRAs related to taking off clothes — maybe the creators also think this action can be left to chance
@anth0nx Thanks for the tips, will try them out and see how it goes.
About your workflow. I'm trying to get a video to lip-sync with an audio input, (i tried both V2V extend and V2V voice over) but I can't get it to work. No matter what prompt I use, the audio is just there and the person's mouth doesn't move. The audio and the clip's duration match , everything else works fine except that.
highly recommend this lora https://huggingface.co/Kijai/LTX2.3_comfy/blob/main/loras/LTX-2.3-OmniNFT-RL-Lora_bf16.safetensors
it can help a lot with sounds.
which format works for 8gb VRAM + 32 RAM?
any, but nvfp4 or gguf 4-5 would be considered best for 8gb
@Darksidewalker fp4 only if you have 50 series rtx
v4 compared to v3... Game changer, same prompt but such dif result! Motions... Especially background details as V3 I couldn't start moving particles around but v4 does it really good 🔥 Thank you
Thank you for the kind words and support!
