If you want to give runpod credits to help with training, feel free to send a code in dm.
A version of the furry nsfw lora for the 14b wan model. Also works with humans.
This is an img2vid lora, it's meant for the wan 14b img2vid models (both 480p and 720p should work). Attempting to generate txt2vid can yield unexpected results.
Avoid using teacache, if you do use it, keep the threshold low. Teacache causes more artifacts with this lora, sometimes making it do strange things.
(Up to) 5x speedup with causvid
I recommend using lightx2v causvid, combined with the mps reward lora for more movement. When combined, complicated actions continue to have lots of movement. Even on 6 steps with euler/euler a + beta it yields great results with lots of motion and great physics.
Previous causvid lora info below (Personally I prefer v1 at 50% over v2)
Using the causvid lora you can get a speedup of around 5 times (assuming you used 20 steps beforehand) It even works with the wan gguf models. Just make sure to follow these steps:
Load the causvid lora at 0.5 strength (as well as this lora at 1 strength)
Sample at 4 (minimum) to 8 (6+ seems to be the sweet spot) steps with a low cfg (<=3, I usually use 2), use beta scheduler for best results. If you see ghosting, ensure your steps and sampler are set correctly.
On my 3060, this lets me generate a 60 frame video in under 4 minutes. The quality is usually higher than teacache, and it's much faster. Do not combine with teacache.
Purpose
This lora can keep characters consistent, and can handle many positions from pov or similar perspectives. It was trained on many positions, follow the prompting guide. Avoid using with t2v loras, your characters might warp and transform.
This lora is capable of generating (without the need for other loras): cowgirl(+reverse), missionary, doggystyle, blowjob(+deepthroat)
It is also capable of handjobs, titfuck. v1 might need assistance, v1.1 seems to be good on that.
It is effectively a lora for NSFW motions in i2v without changing character consistency. And with a better understanding of furry characters.
V1.1:
Continued training with a new dataset, entirely new captions, more perspectives. It usually yields more motion, and it's easier to tag. It can do many positions, perspectives and motions without the need for a second lora.
Prompting v1.1 is like prompting a t2i model, aside from motions, you can prompt for "moving up and down" or similar. Although v1.1 will usually have motion anyways.
V1:
Note: if you're not getting enough motion/not the right speed
If you don't prompt for motion, you won't get any.
"deep thrusts, fast thrusts", "medium sucking, slow sucking", etc. will adjust the depth and pace.
"The woman moves up and down as she rides the man", "the woman uses her breasts to stroke the man's cock." Should be self-explanatory. You can even prompt for pulling out, varying degrees of success.
Prompting should be similar to prompting the 1.3b version.
Trained on 400 res, frame buckets of [1, 8, 16, 24, 32, 40, 60, 80], context as "multiple_overlapping".
Compared to 1.3b version
This model performs much better at oral, has less stretching artifacts. But might be a little harder to prompt right for the motion, at least on really short videos. Make sure you include a speed and depth in your prompts in the case it doesn't animate enough, this could help.
Trained on img2vid, not recommended for txt2vid
I cannot give any promises about quality when used on txt2vid. Especially furry content, I have not tested it and cannot guarantee quality. It might be able to generate some human content, maybe a little furry content. But I would recommend using a different lora instead for those situations.
Description
Initial release, 10k steps is estimated based on last saved steps checkpoint.
FAQ
Comments (38)
Just like with the previous 1.3b, this one excels at quality and perfection. It's simply the best nsfw lora out there!
What all do you need to actually be able to use this though? I doubt its as simple as downloading like a lora and then being done.
Do you need a specific checkpoint/Lora/Model/ A111/Forge/etc
comfyUI is standard just use a wan2.1 workflow, you need alot of system recourses to run the 14B though at least 12gb VRAM and you might be able to get away with a minimum of 32gb of ram but most likely you need at least 64gb of ram. Wan 14b if you want it to make good generations you need to use the fp16 versions and that will use over 64gb of ram for inference. I tried many models on the standard fp8 people generally have by default and 1 in 10 generations would be somewhat ok, swap to fp16 almost all are good if your prompting correctly.
You need wan 14b img2vid and a tool that can use wan with loras. Like comfyui, swarm, wangp, etc
@basedbase @mylo1337 Thank you both for the replies. I'm fairly sure i have the system resources but i don't use any of the things listed by Mylo. sigh And I don't have time to re-learn entire new systems like comfy.
Is there anything out there that allows you to make animations, atleast short loops, that doesnt want to HARVEST your data, or require you to NOT make NSFW like that one who's name i forget, and preferably run it locally, that can be used on Forge?
I only ask because i HAVE looked but i find nothing other than the ones that do the things mentioned above.
@WorstAirtist sd(.)next (not a link) supports video models, with loras I think. It's based on automatic1111 webui. Personally I use swarmui, which requires comfy but does not require you to learn how comfy works in order to use it.
@mylo1337 Thanks for the info, and sorry for the late replies. Had DND all weekend.
What was your batch size and learning rate?
1 batch size, 2 pipeline stages, up to 80 frames at 400 res. It was pretty much maxxing vram constantly. 15 seconds per step.
Learning rate, by memory was like something like 5e-05
@mylo1337 really 15 seconds per step? Im surprised, im training the 1.3b with a batch size of 2 and gradient accumulation steps of 2, 480x480 res and up to 91 frames my time per step is 40 seconds on a rtx 3090 makes me wonder if im bottlenecked somewhere but I don't see how I would be running a 7800x3d and 128gb of DDR5 even have 3.5GB of VRAM to spare so its not like im at 99% VRAM usage.
@mylo1337 Also what was gradient accumulation steps? Trying to get the settings down for furry content since wan its difficult to train it correctly into wan.
@basedbase Unchanged, so 1
@mylo1337 How many training videos were in your dataset for version 1 and 1.1? Did you change video frame rate to be the same across all videos?
@Remiel Video framerate set to 16 for all videos, 95 videos (some being multiple sections from one video with different captions) were used for v1.1. I don't remember the amount for 1.0, it was slightly less.
@mylo1337 Thank you so much for the training info you provided, and, ofc, the lora itself. It is fantastic! I've noticed a lot of other lora's I've tried introduce various problems such as uncontrollable slow motion or fast forward (due to lack of normalization of frame rate during training), degradation of animation quality into shitty low poly 3d look (reason unsure), melting faces (and other parts), desaturation and change of style into gritty realistic (reason unknown) - none of which are present in your lora. Not only does it not introduce problems, but it improves the output quality of other loras when included alongside them. I think there is much learning potential here in your training method because obviously you are doing it right. Did you include various drawing style in the training data (like realistic, anime, 3d, etc.), is that why it works so well and doesn't change the style of the given image into something else like other loras?
@mylo1337 How did you caption the dataset? Besides the obvious motion info, did you caption subject appearance, background, facial expressions, drawing style (e.g. 3d, anime, realistic), quality? (Asking because you are obviously doing something right so I wish to learn from your success.)
@Remiel v1.0 was a combination of the datasets I used for the 1.3b fun inp model. Some captions contained full details about speed, some of the captions didn't have any information about the character itsself, this was intentional. I was trying to generalize the model as for i2v the actual character is in the input, so it doesn't technically need to be explicitly captioned. I did end up captioning the characters in later versions as it probably helps when the character's face is partially out of frame in the beginning.
v1.1 used captions that were created by first asking an llm (gemma 3 glitter with a vision adapter, in my limited testing it outperformed models that were trained specifically for captioning), then fixing up and slightly improving the llm captions, adding some details about the exact motions (as it used a static frame).
In general all versions had information about the motions, position, perspective, style, etc. v1.1 had significantly less realistic content in the training data, however it didn't affect the model's performance on realistic content, I'd say it probably got better. v1.1 also had more training data with solo characters, non-pov perspectives, and actions that were too limited in the dataset for v1.0. V1.1 outperforms v1.0 at pretty much everything, as it should since it's a continuation of the training.
All caption writing, video cropping, and ffmpeg exports (for cropping and fps change) were done using my captioning tool. LLM captioning was added specifically for v1.1 of the dataset to speed up my captioning process. I'm considering releasing the tool to the public some time soon as well.
@mylo1337 Was v 1.1 a checkpoint you resumed from with new training data/lr
@mylo1337 Thank you so much for sharing. 😘 I'd love to see your captioning tool! Currently, there are very few options available for video captioning and doing it by hand by creating txt files and then opening them in windows explorer is... uh... tedious to put it mildly. Even video cropping, which should be a basic function for a video editor, is either not possible or very inconvenient in most video editors. Shocking, isn't it? I only found out when I actually wanted to crop videos for training purposes. Most video editors require you to set a resolution of the project (which you can't do because you don't know the target resolution beforehand) and you can't crop it by hand like you would easily crop an image in Photoshop. I've yet to find a simple Photoshop-like video cropping app. Some video editors don't let you change framerate at will but only offer some more commonly used fps which doesn't include wan's 16. It's a giant mess making it very difficult to prepare a proper dataset.
@basedbase v1.1 was resumed from v1.0 with a new dataset, some of the videos of the old dataset were also in the new one, but they have new captions.
I didn't have the step checkpoint, so I initialized from the lora weights of v1.0 instead. So the e22 is 22 new epochs on the new dataset, not 22 total epochs.
@mylo1337 I also just did the same with my new single NSFW pose lora recaptioned 5 videos each extended to 5 minutes each, used dense captions I had qwen3 30b a3b rewrite and trained 8 epochs for a much better end result. Used a very low LR for tuning so it would not suffer from catastrophic forgetting LR was 4e-6 for the tune and I used LR of 5e-5 for the base lora. Each used a batch size of 4.
@Remiel Try captioning and editing a 170 video dataset it takes a very long time lol, going forward i'll be having a LLM rewrite all my captions to be denser and more descriptive, in my testing that helps alot with the quality of the outputs the lora makes
@Remiel Just recaptioned and retrained a dataset to use much more descriptive text files and specify if the video is 3D, 2.5D, or 2D and increased batch size from 4 to 12 and used square root LR scaling and got the most consistent lora I've trained to date. Proper dataset captioning is paramount even my tensorboard loss chart shows a consistent drop in loss for train/loss vs all previous runs where it just bounced all over the place.
@Remiel I trained the 1.3b though so if you scale batch size that much on wan 14b you will need at least 3 a6000 48GB GPU's, though you likely don't need to scale that much on the 14B since it learns much better per optimiser step than the 1.3b since you know, 1.3b parameters vs 14b parameters.
Just a heads up civitAI made a change to their TOS that says NSFW posts will be hidden from peoples feeds if the metadata and prompts are not in the video/image post
Yeah I'll see if I can add it somehow. It's swarmui which I don't think civitai normally recognizes. Or maybe I need to explicitly use the json as well.
Civitai only recently started to support slightly more than extremely basic and limited ComfyUI metadata, and apparently the on-site gen uses ComfyUI for the backend, so you can likely guess how well the site will handle anything less common.
Now that I got a chance to try this. I can say without a doubt that this model captures the motion of sex acts really well.
It doesnt need to be furry images, it works extremely well across anime, realism, etc.
Does this Lora only work with the full Wan model, or can it work with a smaller Q8 or Q6 version? I've been using Q6, and no matter what I try (prompt specifying speed and action, photorealistic image, etc.) I can't get any good motion or good output. I'm trying to animate a side view doggystyle picture. Anyone help?
As long as it's the 14b model. You can pick any quantization, or scaled version.
Just so you know, Q8 is just as good as FP16. I've use them both and the difference is almost non existent.
All gens I've uploaded were on Q3, so yes. I wrote some prompting help on the 1.3b version https://civitai.com/models/1429979?modelVersionId=1678119. For some reason all the images are hidden on that model page now, not sure why since I did add metadata.
The more info you write for the movement, the better. There wasn't much side view doggystyle in the training data though, so that might be a tricky one, you might want to include another lora for that at like 0.7 or 0.8 strength.
Please provide prompts whenever and wherever possible folks. I know some like to keep their 'tricks' up their sleeves. But sharing is caring. Share the love. Peace.
I have to say, this is fantastic. Using other concepts it takes quite a lot of tries to get something good, but this Lora I get some very nice results within 5 attempts. Nice job.
If only I had 2 5090's in SLI. Lora is great but it suck's so much RAM I'm surprised my PC didn't blow up
A single 4090 is fine, some of the example vids were genned on a 3060, using a Q3_K_M quant of the model, you can use a gguf quant https://huggingface.co/city96/Wan2.1-I2V-14B-480P-gguf/. This works with comfyui-gguf (https://github.com/city96/ComfyUI-GGUF).
@mylo1337 How much time would you say it takes per video and length? My 2080 with the GGUF at length 33 steps 20 took 18 minutes based on one time test. It uses about 17.5 GB of RAM (27GB of RAM all together). Using non GGUF takes 40GB of RAM, including OS and other stuff. I'm also curious about the wattage on the 4090 and 3060 when using AI. I would estimate that the 4090 total system power would be 500 watts from the wall. I get around 260 watts on average. I was looking forward towards the 5090, but it just sucks so much power and has issues, let alone the price. My room already blows the breaker when my AC and computer are on at the same time (Happens rarely).
@brand175 idk exact times but I know that even on q3, the 3060 is bottlenecked by memory, I get about 20% usage with small spikes to around 50% usage. Not sure about exact wattage
Thats Wan 14B in general, you can use block_swap but I don't have issues running at full fp16 precision on a single 3090 its still slow at least 10 minutes for a higher res generation with any respectable frame length. Try using dpm_2 for the sampler and ddim_uniform for the scheduler that produces the most consistent and best motion generations from all the sampler combos ive tried. Use around 20 steps for the best quality but 6 steps is good for a quicker generation. Samplers and schedulers matter alot when aiming for quality UniPc is honestly not that great.
Details
Files
Available On (1 platform)
Same model published on other platforms. May have additional downloads or version variants.










