0.9:
This is still i2v trained, but I have noticed with 0.9 that T2V works? Some motions seem like they are coming through on some t2v tests I did and they didn't seem like horror or anything. Slightly improved nudity and motion if you prompt for it. It shouldn't interrupt character loras as it's very weak on upper body and face data and is more of an amalgamation of body movements from many angles. Previews use this workflow. Modified it a bit here, use res_2s sampler from RES4LYFE on first pass, LCM on upscale pass, and my sigma changes and an optional audio latent load to track the output with custom sound voice or music.
Problems already identified: Undertrained, under ranked (should probably be rank 64 for these nsfw concepts) and some tagging changes with angle set ups and pruning of the dataset can be done for 1.0.
I also keep getting a visual bug with the upscaler in to larger vertical formats, if the image goes beyond 1536 vertically sometimes you get color noise at the top and bottom, not related to the lora. Anyone with a fix for that lemme know. Another quirk, or probably advantage, is that this lora has so much motion impact you can occasionally lower the compression on the PreProcess node and get some better adherence to start image, or in some cases turn it off completely and isolate the motion to only the lora. That's not usually plausible without loras, results just zoom in or freeze.
I tried to reform the dataset tagging to bind the base concept under:
doing a sexy move
The model seems to actually pick it's base concept up in context the way I wanted it. If you start a prompt with "a naked woman standing in front of the camera while doing a sexy move" it should start out with it. But you can do things like another action and dialogue, then later on in the prompt: "blah blah blah. Then she starts doing a sexy move." You can bring it in later on in context. At least that's the idea and I gave it clips to lead into that.
You can set these other keywords up in-context with something like:
"she is doing a sexy move with side to side hip movement and butt grinding."
Won't always work, depends on start image and the usual LTX2 finding a good seed and refining prompt until it works. You don't want to just have a short prompt using only these tags, you should still prompt the base model correctly and mesh the data from this into it with a normal long and descriptive prompt. All these tags do is just give a hint to what kind of motion and scene it's trying to do, but this is pretty strongly trained and rigid and if the image is complicit it knows what to do, but also may just start mixing different motions together regardless of prompt. That's the biggest issue to try to iron out. The entire set if mostly static camera as well, obviously for better data clarity.
Angle Description:
Seen from behind
Low angle static camera (crotch focus)
Facing away
From above
(all else is assumed to be front facing camera, probably don't prompt anything)
Motions:
grinding
bouncing
feeling
touching
rubbing
squeezing
dancing
Front:
a naked woman, a semi-naked woman, a girl
vagina and anus
sex, the penis goes inside the vagina, two people having sex
grinding
breast squeeze, grabs her breasts, rubbing, rubs, feels, playing with her breasts
breast bounce, breast jiggle
sexy dance, dancing in-place, bouncy dance
side to side hip movement
gyration, gyrations
Bouncing up and down
Back:
Seen from behind, facing away
Turns to face away from the camera
Looking back over her shoulder
Butt
Butt jiggle
in a doggystyle position/pose
butt grinding
twerking
shaking her butt
feeling her butt
spreading her butt showing her anus and vagina with her hand/hands
Sex: (don't expect much, needs targeted lora motion to partner with for a real output. This is almost just noise to make it aware of concepts. Also is a lot more riding on-top focused.)
explicit motion
she reaches down and spreads her vagina with her fingers
bouncing up and down while having sex
the penis goes inside the vagina
POV
man has sex with the naked woman
very long hard and stiff penis
Anal
Strength: 0.2-0.91, 0.75-0.8 seems the best zone. High strength has adverse effects. Always try it turned off to see if it's even helping, its not magic.
This one is actually still kind of a test of course. It is overfit, but also not? I thought the first one was, but it isn't, just the audio. The 768 bucket seemed like it was going to overfit first, but the 512 long bucket seems to still work to extend and smooth things out. Idk like most others I also find this model has a lot of quirks to training it. The wan2.2 one was trained like this and works fine, but it's all aligned with the same frame count, I never did a split frame count but I don't see any kind of interruptions happening on long generations. I have seen bad outputs and good outputs, but then when it's a very explicit image and I turn the lora off: nothing happens at all, lol. So at minimum you can just use this at half strength to give nude characters in erotic positions movement at least, and it will still have a lot of impact.
This is also mainly a motion support lora, if this was Overwatch or LoL, this lora is Mercy or Lulu. On it's own it will do a few subtle unique motions, but mainly just hone the model in on a single female subject. If you use it with more targeted NSFW loras, make sure you balance this lower, it has sex and NSFW data, but only enough to mesh in with more targeted loras. Do as I say, not as I do. This dataset is borderline untrainable with too much variance, but just enough to transfer a pretty large variety of motions and lifelike quality through scraping on overfitting. You're way better off training if you focus on one thing and one motion with a dataset of that thing in slightly different shapes and settings with very similar composition.
After testing base model more, I changed tagging for this back to the Wan2.2 lora + less vulgar terms for body parts. LTX2 with the default Text Encoder seems to understands anatomy in less vulgar terms. So butt, vagina, breasts, penis, testicles; it knows what parts of the body they are, just not how to fully depict or animate on them. Pussy, ass, cock, dick seem a lot less reliable.
As far as the loss of the sexier voice from the test: I'm trying to put together JOI data, ASMR, and dirty talking during sex clips to make that it's own Sexy Voice package.
V0.9 is expanded with 768 data at 121 frames, audio removed, since the audio training is way too impactful and can conflict with loras that actually need dedicated audio data tied to their concept.
Test 0.5: first rough version with less unique motions and motion clarity, but it has a way sexier voice lines.
Test 0.5 keywords:
Style: realistic - sexy.
Style: realistic - explicit.
Pussy, ass.
The rest of the motion same as 0.9, but use pussy and ass instead (weaker tokens).
Description
512x - 241frames.
FAQ
Comments (50)
I have a feeling that LTX-2 will be known as an I2V model due to its' shit performance in T2V with NSFW concepts. I am done trying to train it for T2V. Way to much effort, all things considered.
Have you checked the official LTX lora trainer? https://www.youtube.com/watch?v=sL-T6dsO0v4
yeah finding the same training nsfw LORAs for LTX-2, works for I2V but T2V sucks. even sfw I get crappy T2V results
The solution is really to fine tune it with at least 720p clips at full size, that is going to take insane memory and cost to do though.
Tbh though, I'm not sure why people want T2V so bad, these models just have the same old image model system that is way worse than dedicated ones underneath that draws the first frame, then flow model takes over. Just use a better image model to have full control over that start frame with I2V.
@Cerner yeah. LTX-2 has absolutely zero knowledge of anything NSFW and the teXt encoder Gemma 3 is also very censored.
@tenstrip dont worry ill just grab some H200's at the corner store for us
@playtime_ai_ body horror in the samples during ltx-2 training is quite scary at times!
@playtime_ai_ I'm not sure exactly how though, it definitely doesn't like "fuck" or nipples, but it says and does dirty prompts. I did 3 tests to compare with prompt the "camera zooms in to a girl's anus for a close up." Using this lora, both encoders did that output, on both a female on all fours, and then a femboy with a penis and balls on all fours. It even added like anus wrinkle details and pubic hair. Then without the lora they both zoomed into her pussy area or usually face on both. So the training and term carried through to both. I hope people training T2V are trying 768 or 1024 and higher resolutions instead of just 384 pixels and calling it off. Gotta actually give that detail, or try Differential Guidance to secure the learning.
@tenstrip yep, im trying differential guidance at 2.0
I disagree with you. The LTX2 has a lot of potential for NSFW T2V, a great example is this LoRa: ( https://civitai.com/models/2298764/prone-face-cam?modelVersionId=2601985 ) All my T2V outings with this LoRa are incredible, the visual quality and movements are fantastic. I think everyone who wants to train for T2V on the LTX2 should learn from this guy's technique.
@AI_2_addicted Yeah it does you can turn it into a straight 1920x1080 30 second porn generator. It's just what's involved with doing that lol, you're basically making a new model at that point.
@tenstrip One lora and you are that confident? lol. To be fair, your model (especially v2) turned out quite nice... but with the effort that it takes to create a good NSFW model, the ecosystem will never develop.
It's also bad for I2V NSFW. It's a complete joke compared to WAN 2.2 in fact.
@aurelius bro wan 2.2 couldn't do anything when it came out. I still have the gens trying to even get a man to thrusts. You're comparing a year of nsfw loras to nothing.
@tenstrip I'm comparing a year of my personal experience with creating NSFW Loras for wan2.2 with my personal experience with LTX-2. I know what I'm talking about
@playtime_ai_ yeah same here lol. Make that training since Sd1.5 embeddings were a thing.
@tenstrip It fails on a basic level. It can't even undress a person because it has like 0 naked people in the training data. Wan 2.2 will undress a person and give you pretty much 1:1 their body shape, body fat % and all. It also turns every face into an uncanny valley plastic/rubber mask. Wan 2.2 actually keeps a person's face looking like themselves/the initial image in motion.
@aurelius I mean undressing and nudifying is becoming literally illegal. Also one of the reasons Wan team may never release their model open again.
@tenstrip So... you agree with us now? Your point does make sense. Wan seperated T2V and I2V models where LTX-2 does not.... it is a single model that does both. That reasoning may hold some weight.
@playtime_ai_ idk who us is, or actually what your point is, or why you were arguing with me when I never called anything you said into question. The model doesn't know nudity at all, like 0 aside from combining male nipples on like a bra-shaped breast area. They trained it mostly on cartoons, it's gonna take a team and a lot of cards to fix that, but I want nudity, I shall use i2v, to each their own.
@tenstrip we aren't arguing...well, I'm not. Just trying to have a discussion... But ok.
@playtime_ai_ @tenstrip Im just eating popcorn and watching the Sexy Move girls dance
@sexgod1979 Too true
@playtime_ai_ Oh, okay. But my point was always with T2V its like a cycle to reinvent the wheel now after having seen it with HY, LTX, Wan, HY again, and then this again. A constant struggle to train nudity into these inferior little flow models that have to be run quantized and half-precision, when there's full porn SDXL finetunes, qwen, chroma, and then zimage to make the perfect scene and starting image. All T2V is to me is a crappy little image model included to generate your start image, which I'd rather do myself.
@tenstrip I 100% agree with you... But LTX is the worst out of that list when it comes to nsfw and nudity. It's much more akin to flux than it is to the others.
@playtime_ai_ I still can't tell if it's censored or just needs the right loras to force it. This one forces out female solo motions at least as a proof of concept, even at like 0.4 strength. That better nudity T2V lora and some of the other ones trained off images do seem to work and add pretty detailed nipples.
@playtime_ai_ I wouldn't say its the worst at least for i2v, I can get pretty good gens using the furry nsfw lora mylo posted. It does great at 2d which I could never really get wan 2.2 to do. Sure wan 2.2 with a bunch of loras is great at NSFW it seems to turn any gen even if 2d into almost 3d in the way the animations are while ltx2 with the 1 furry lora and just the distilled lora it does 2d really good. I think it will get better all around whenever everyone figures out the optimal training and inferencing settings. I do use the full size model I tried an fp8 version and nvfp4 and the quality was quite bad.
@playtime_ai_ I have had some insane results training ZIT, LTX-2.. Far better than I ever got with Hunyuan, and then Wan22.. I think the biggest issue we are faceing, or at least I did.. Is a lack of any real guides into training LLM text encoders, as its all new ground.. with most people trying to use old CLIP captioning and dataset methods.. or even worse, having A.I write huge long prompts for small concept loras. I had tens of failed trainings and thought it was the models. Nearly gave up.. but the last couple of weeks, seeing some of the results I have, completely changed my opinion. Kudos to you playtime, you are an epic contributer. Don't give up hope..
@DaveTheRave5 any captioning tips for ltx-2?
@sexgod1979 I reverted to the way I captioned the Wan 2.2 version. I created one sentence or phrase to describe the different motions, different angles, etc, and just copy pasted it to every clip that it was even slightly applicable too. Some of the descriptions are straight up lifted from testing the base model because they kind of already work. All the subjects were either woman (curvy older) or girl (thinner younger woman look), naked, swimsuit, or semi-naked in a state of undress. Didn't bother deeply describing outfits or scenes, only motion , pose, and angle with context in what order things happen, because with 10 second clips it's a lot of evolution. That's all you're trying to train in i2v, is literally just pixel change and describe exactly why and how those pixels change. Things that I tried to test with the base model that had almost no relevance I did a lot more tagging for, especially when it's an ass focus or a crotch shot, so I gave it "low angle camera", "seen from behind" to those clips. To give things context, I used a lot of "then, x. Then x, while x." Manually captioned of course.
@DaveTheRave5 That is a bold and confident statement from someone who hasn't posted a single image, video or model...
@tenstrip cool thanks! sexy move lora is looking good!
@playtime_ai_ True.. but I have been training since the early SD days, and I have spent the last month doing 12-14 hour days trying to break LLM encoder training.. so maybe this is my time to give something back. Honestly, I only commented, as you are an epic contributer, and it's sad to see you almost giving up on LTX-2 (Like a lot of others).. as I really believe we haven't even scratched the surface on it's potential.. and a lot of that is down to lots of misinformation and misconceptions out there regarding training LLM based models IMO.
@sexgod1979 Serious disentanglement strategies, understanging things like archetypes and how to seperate them is much more important with LLM encoders.. staying the hell away from long A.I generated captions, which are great for big style datasets or core training.. but destructive for small concept loras, due to how LLM encoders will create entaglement accross the board; in a very fragile network in the case of LTX-2.. Prompting styles that used to work with CLIP, for instance "a person sitting/standing/kneeling praying" no longer work in my testing.. Instead, divide archetypes' by varying the terminology..
"A person standing with their hands together praying"
"A human kneeling, arms in front of them and fingers interlocked"
"A character sitting, their palms against each other positioned near their stomach"
There are many other aspects. The biggest shift is getting away from how we conceptualise training, as much of what has become "law", is CLIP based, and not applicable or optimal for LLM based training.
Another example of that, tagging things we want ignored in the captions.. This is far more risky with LLMs, and can have to opposite effect. I've seen two tags, making up just 2% of captions influencing every generation.. as an LLM encoder has no concept of why something is mentioned, and can prioritise it as "important".
Great seeing the lora development take shape for LTX-2, well done!
Heya was this trained on 16 fps videos? my text to Video examples move fine but talk in slow motion even at strength 0.40
24fps. That's usually from the abliterated gemma with the lora at high strength, they sound drunk. The audio part of this lora will be removed completely for 1.0, it would cause too much interference for what it's supposed to do. I'll split the erotic talking off into it's own JOI/ASMR lora.
@tenstrip so was this trained on abliterated gemma? Just curious if i should too for loras I train.
@kronos1959777 No and I wouldn't. Like I said I'm pretty sure you can reference genitals with the proper name "vagina, penis, butt, breasts," even "anus and butthole" seem to work on the normal encoder just fine. Even on base model and normal encoder with no lora, give it an image of a naked woman or man, prompt "zoom in on vagina, she touches the vagina with her hands" that will be an output. If you want to train "pussy" then you're inventing a new concept so it's gonna need a lot of data, but here I'm trying to work with the base model and enhance it only.
Also, the abliterated version of these encoders is just usually removing violence and bomb-making refusals, not really doing much but changing some motion and really destroying the voice outputs on LTX2. What you really want is to somehow get or train a nsfw-gemma that's aware of sex slang and points reference to the genitalia or sex to train them on and hook it into the model, that's beyond my understanding though.
@tenstrip 👌 I would think vulva would also work well? My caption from taggui identified vulva.
@kronos1959777 I'd take an image of a close up on a vagina in the base model, see if you can get anything to happen to it with "x to her vulva" or "the vulva of the vagina, x." If that works, like anything at all. you're good. Otherwise it shouldn't really matter if you tag that and really get a lot of close-up data of "vulva" and where they are and what happens to them, it should work fine for prompting outputs.
why i2v ? the power of LTX2 is t2v
For slop yeah, but not women that look like this without some insane lora stacking. I take the easy way out.
for me, I think taking a few minutes to generate a solid image in z image turbo or flux makes a big difference in the video outcome.
@sexgod1979 Yeah I see the end goal of T2V for nsfw, but tbh you need an entire dedicated trained model on all kinds of porn to pull it off. This lora could work T2V but you get more value by applying it to sexier start images. I mean if you look at top videos by reactions, they're ALL i2v.
The samples are great, some of the best ive seen yet. Nice work!
i am using this LoRA in combination with the Multi Purpose NSFW Lora. The Quality of the NSFW Scenes will be so much better with this! Outstanding work!
Yeah they will. I was actually surprised how it took a model that would just zoom in to picture with no movement and started adding a lot of subtle lifelike stuff as well as the other motions. LTX2 trains very well and I think it's lack of NSFW might actually be a good thing because there aren't random motions or concepts that can interfere and disrupt the lora's influence.
Details
Files
LTX2-i2v-SexyMove-test.safetensors
Mirrors
LTX2-i2v-SexyMove-test.safetensors
LTX2-i2v-SexyMove-test.safetensors
LTX2-i2v-SexyMove-test.safetensors
LTX2-i2v-SexyMove-test.safetensors
LTX2-i2v-SexyMove-test.safetensors
LTX2-i2v-SexyMove-test.safetensors
LTX2-i2v-SexyMove_test000006000.safetensors
LTX2-i2v-SMove-test.safetensors
LTX2-i2v-SexyMove-test.safetensors
LTX2-i2v-SexyMove-test.safetensors
LTX2-i2v-SexyMove-test.safetensors
LTX2-i2v-SexyMove-test.safetensors
ltsm.safetensors
LTX2-i2v-SexyMove-test.safetensors
LTX2-i2v-SexyMove-test.safetensors
LTX2-i2v-SexyMove-test.safetensors