🎬 LTX See Motion
Custom LTX-2.3 Checkpoint · Image-to-Video · NSFW-focused · ComfyUI-ready
⚡ TL;DR: A custom-adjusted LTX-2.3 checkpoint, built on the foundation of 10Eros v1.2 and sulphur dev — no LoRAs baked in, no Turbo/distill baked in. Built for image-to-video: the model stays true to your input image. For the fast 8-step / CFG 1.0 workflow, load the official distill LoRA at 0.8 (link below). 🔞 Adults only.
📌 A quick word before you start: this version is tuned for image-to-video — my main goal was that identity, style, and composition of the input image survive the whole clip. Because of that focus, plain text-to-video is not the strength of this version. Video-to-video and audio + image-to-video work well too; video editing I have not tested yet. Dancing, singing, and NSFW prompts work well without any extra LoRA. Lots of example videos are in the gallery, and a workflow showing how they were made will follow. 🎞️
💠 What is LTX See Motion?
LTX See Motion is my personal LTX-2.3 video checkpoint blend, built to make image-to-video generations feel more active, more stable, and more useful for adult NSFW motion work — scenes where anatomy, body contact, and intimate motion need to stay readable across the whole video instead of falling apart after a few frames.
This release is a custom-adjusted LTX-2.3 checkpoint. Two existing checkpoints served as its foundation: 10Eros v1.2 and sulphur dev.
Nothing else is inside:
🚫 No LoRAs merged into the checkpoint
🚫 No Turbo / distill influence baked in
✅ Standard LTX-2.3 structure (fp8 mixed)
✅ Fully LoRA-compatible — stack your own LoRAs without double-stacking
🧩 Required LoRA for the fast workflow
⚠️ Important: This checkpoint contains no distilled (Turbo) weights. Out of the box it needs a normal dev-style workflow (more steps, normal CFG). For the fast 8-step / CFG 1.0 workflow, load this LoRA separately:
👉 ltx-2.3-22b-distilled-lora-384-1.1 — download here · strength 0.8 · goes into ComfyUI/models/loras/
⚙️ Recommended Settings
🐢 Without the distill LoRA (dev-style, official LTX-2.3 defaults)
CFG / guidance:
4.0Steps:
40(community range:30–50)Use a fixed seed when comparing settings
🚀 With the distill LoRA loaded (strength 0.8)
CFG / guidance:
1.0(keep it there)First pass:
8 stepsUpscale / second pass:
4 steps
⏱️ Performance reference
RTX 5060 Ti · 16 GB VRAM · 15-second video @ ~1 megapixel → ~300–400 seconds
💡 If results become unstable, don't raise CFG first — adjust the prompt, seed, image quality, or motion intensity instead.
🎯 Recommended Use
Best used for image-to-video, especially adult/NSFW image-to-video — the model stays true to the input image. Video-to-video and audio + image-to-video work well too. Dancing, singing, and NSFW prompts work well without any extra LoRA. The starting image should already define the subject, pose, composition, and style — your prompt describes what moves.
🎥 Camera movement
💃 Body, hair & clothing motion
😊 Facial motion
🔞 Intimate motion & contact stability
💡 Lighting changes & environmental motion
Avoid purely static scene descriptions — if the prompt only describes what's already visible, the model has no direction for movement.
📺 Video Guides / Walkthroughs
Step-by-step walkthroughs for every path this workflow can do. Each one is a short screen recording with on-screen chapters.
Image to Video — the basics: turn a single image into a video with natural motion.
https://civarchive.com/images/136919778Video to Video — extend any existing clip to any length you like.
https://civarchive.com/images/136922039Audio to Video — feed a 30 s song (with vocals) + one image and get a 30 s music video, lip-synced to the track.
https://civarchive.com/images/137063968IC-LoRA to Video — drive motion, poses or depth from a control video using an IC-LoRA (like the one shown in the clip).
https://civarchive.com/images/137075598LoRA + Image to Video — multi-subject / character-sheet consistency. Needs the LTX-2.3 Licon MSR V2 LoRA.
Guide: https://civarchive.com/images/137093666
LoRA: LTX-2.3-Licon-MSR-V2.safetensors
Multiple-Subject-Reference LoRA by LiconStudio — keeps multiple characters and the background consistent across frames. Load it alongside the distilled LoRA.
✍️ Prompting Guide
LTX image-to-video prompting works better as a short directing paragraph than a tag list. The input image defines the scene — the prompt tells the model what happens next.
Prompt shape
[subject/action in present tense]. [camera movement]. [body/contact motion].
[secondary motion: hair, cloth, lighting, environment]. [stability note].🌸 Example — Soft Motion
The character slowly turns toward the camera and breathes softly. The camera makes a gentle push-in while the hair and clothing move with a light breeze. The expression stays stable, the face remains consistent, and the motion is smooth and natural.🎥 Example — Camera Movement
The camera slowly pans from left to right while the character keeps eye contact with the viewer. The body shifts subtly, hair moves naturally, and the lighting flickers softly across the scene. The character remains anatomically stable with smooth frame-to-frame motion.🔞 Example — Adult I2V
One adult anime woman kneels between the legs of an adult anime man, her hands braced on his thighs. She looks up at him with direct eye contact, opens her mouth, and slowly takes his erect penis fully inside her mouth, her lips stretching visibly around the shaft as she lowers her head. She pulls back until only the tip remains between her lips, then pushes down again, establishing a slow, deliberate rhythm. Saliva glistens visibly on her lips and his shaft with each stroke. Her cheeks hollow slightly with suction, her throat visibly adjusts as she takes him deeper, and her eyes water slightly but stay locked on his face. He groans and grips her hair gently, guiding her pace without forcing. Her shoulders, chest, and hips shift naturally with each motion. Add wet sucking sounds, soft gagging catches, his low groans, her muffled moans, and heavy mutual breathing. Detailed throat and lip animation, visible saliva physics, cinematic anime style.💡 Tips
Write one flowing paragraph, not a tag wall.
Use present-tense verbs: turns, moves, breathes, leans, holds, sways.
Describe the camera relationship: push-in, slow pan, close handheld angle, static shot.
Say what should stay stable: face, anatomy, hands, contact, identity.
Don't overdescribe what the image already shows.
One main action per generation — if it gets chaotic, simplify instead of adding more.
📊 Known Behavior
✅ More motion-biased than a neutral LTX-2.3 checkpoint
✅ Stays true to the input image — identity and style hold up across the clip
✅ Dancing, singing, and NSFW prompts work well without any extra LoRA
✅ Video-to-video and audio + image-to-video work well
✅ Strongest results in adult / NSFW image-to-video
✅ Depicts adult anatomy (breasts, penis, vulva) more reliably than the neutral base
✅ Handles explicit categories (oral, anal, vaginal) well when the source image is clear
⚠️ Can exaggerate movement if the prompt is too aggressive
⚠️ Works best when the input image is already strong
⚠️ Plain text-to-video is noticeably weaker — this version is tuned for image-to-video
⚠️ Video editing is untested so far
⚠️ Needs the separate distill LoRA for low-step / CFG 1.0 workflows
🚫 What This Model Is Not
❌ Not a clean official base model
❌ Not a from-scratch fine-tune
❌ Not a distilled / Turbo checkpoint
❌ Not a text-to-video model — use it with an input image
❌ Not a perfectly neutral LTX-2.3 checkpoint
❌ Not intended for SFW-only use cases
❌ Not guaranteed to fix every hand, face, or temporal artifact
❌ Not for minors or illegal/non-consensual content
🛠️ Suggested Workflow
Start with a clean, readable image.
Load the checkpoint in your LTX-2.3 ComfyUI workflow (
ComfyUI/models/checkpoints/).Optional: load the distill LoRA for the fast workflow.
Use a short, motion-focused prompt.
Without LoRA:
40 steps, CFG4.0· With LoRA at0.8:8 steps, CFG1.0, second pass4 steps.Test a few seeds before changing the whole workflow.
🎞️ Coming soon: a full ComfyUI workflow showing exactly how the gallery example videos were made will be added to this page.
🙏 Credits
Lightricks — LTX-2.3
TenStrip — 10Eros LTX-2.3 work
The creator of the sulphur LTX-2.3 checkpoint
The ComfyUI & LTX-2.3 workflow community
Adjustment / release: SeeSee21
📜 License
This model is a derivative of LTX-2.3-based checkpoints. Please respect the licenses of the original LTX-2.3 model and all source resources. LTX-2.3 is provided by Lightricks under the LTX-2 community license — commercial users should check the official terms, especially the revenue threshold requiring a separate commercial license.
🔗 Links
🎬 LTX See Motion — a personal LTX-2.3 adult motion blend for ComfyUI video generation.
Description
🎬 LTX See Motion
Custom LTX-2.3 Checkpoint · Image-to-Video · NSFW-focused · ComfyUI-ready
⚡ TL;DR: A custom-adjusted LTX-2.3 checkpoint, built on the foundation of 10Eros v1.2 and sulphur dev — no LoRAs baked in, no Turbo/distill baked in. Built for image-to-video: the model stays true to your input image. For the fast 8-step / CFG 1.0 workflow, load the official distill LoRA at 0.8 (link below). 🔞 Adults only.
📌 A quick word before you start: this version is tuned for image-to-video — my main goal was that identity, style, and composition of the input image survive the whole clip. Because of that focus, plain text-to-video is not the strength of this version. Video-to-video and audio + image-to-video work well too; video editing I have not tested yet. Dancing, singing, and NSFW prompts work well without any extra LoRA. Lots of example videos are in the gallery, and a workflow showing how they were made will follow. 🎞️
💠 What is LTX See Motion?
LTX See Motion is my personal LTX-2.3 video checkpoint blend, built to make image-to-video generations feel more active, more stable, and more useful for adult NSFW motion work — scenes where anatomy, body contact, and intimate motion need to stay readable across the whole video instead of falling apart after a few frames.
This release is a custom-adjusted LTX-2.3 checkpoint. Two existing checkpoints served as its foundation: 10Eros v1.2 and sulphur dev.
Nothing else is inside:
🚫 No LoRAs merged into the checkpoint
🚫 No Turbo / distill influence baked in
✅ Standard LTX-2.3 structure (fp8 mixed)
✅ Fully LoRA-compatible — stack your own LoRAs without double-stacking
🧩 Required LoRA for the fast workflow
⚠️ Important: This checkpoint contains no distilled (Turbo) weights. Out of the box it needs a normal dev-style workflow (more steps, normal CFG). For the fast 8-step / CFG 1.0 workflow, load this LoRA separately:
👉 ltx-2.3-22b-distilled-lora-384-1.1 — download here · strength 0.8 · goes into ComfyUI/models/loras/
⚙️ Recommended Settings
🐢 Without the distill LoRA (dev-style, official LTX-2.3 defaults)
CFG / guidance:
4.0Steps:
40(community range:30–50)Use a fixed seed when comparing settings
🚀 With the distill LoRA loaded (strength 0.8)
CFG / guidance:
1.0(keep it there)First pass:
8 stepsUpscale / second pass:
4 steps
⏱️ Performance reference
RTX 5060 Ti · 16 GB VRAM · 15-second video @ ~1 megapixel → ~300–400 seconds
💡 If results become unstable, don't raise CFG first — adjust the prompt, seed, image quality, or motion intensity instead.
🎯 Recommended Use
Best used for image-to-video, especially adult/NSFW image-to-video — the model stays true to the input image. Video-to-video and audio + image-to-video work well too. Dancing, singing, and NSFW prompts work well without any extra LoRA. The starting image should already define the subject, pose, composition, and style — your prompt describes what moves.
🎥 Camera movement
💃 Body, hair & clothing motion
😊 Facial motion
🔞 Intimate motion & contact stability
💡 Lighting changes & environmental motion
Avoid purely static scene descriptions — if the prompt only describes what's already visible, the model has no direction for movement.
✍️ Prompting Guide
LTX image-to-video prompting works better as a short directing paragraph than a tag list. The input image defines the scene — the prompt tells the model what happens next.
Prompt shape
[subject/action in present tense]. [camera movement]. [body/contact motion].
[secondary motion: hair, cloth, lighting, environment]. [stability note].🌸 Example — Soft Motion
The character slowly turns toward the camera and breathes softly. The camera makes a gentle push-in while the hair and clothing move with a light breeze. The expression stays stable, the face remains consistent, and the motion is smooth and natural.🎥 Example — Camera Movement
The camera slowly pans from left to right while the character keeps eye contact with the viewer. The body shifts subtly, hair moves naturally, and the lighting flickers softly across the scene. The character remains anatomically stable with smooth frame-to-frame motion.🔞 Example — Adult I2V
An adult couple moves in a steady intimate rhythm. The camera holds a close cinematic angle with a slight handheld feeling. Body contact remains consistent, anatomy stays stable, hands remain coherent, and the motion is smooth without flicker.💡 Tips
Write one flowing paragraph, not a tag wall.
Use present-tense verbs: turns, moves, breathes, leans, holds, sways.
Describe the camera relationship: push-in, slow pan, close handheld angle, static shot.
Say what should stay stable: face, anatomy, hands, contact, identity.
Don't overdescribe what the image already shows.
One main action per generation — if it gets chaotic, simplify instead of adding more.
📊 Known Behavior
✅ More motion-biased than a neutral LTX-2.3 checkpoint
✅ Stays true to the input image — identity and style hold up across the clip
✅ Dancing, singing, and NSFW prompts work well without any extra LoRA
✅ Video-to-video and audio + image-to-video work well
✅ Strongest results in adult / NSFW image-to-video
✅ Depicts adult anatomy (breasts, penis, vulva) more reliably than the neutral base
✅ Handles explicit categories (oral, anal, vaginal) well when the source image is clear
⚠️ Can exaggerate movement if the prompt is too aggressive
⚠️ Works best when the input image is already strong
⚠️ Plain text-to-video is noticeably weaker — this version is tuned for image-to-video
⚠️ Video editing is untested so far
⚠️ Needs the separate distill LoRA for low-step / CFG 1.0 workflows
🚫 What This Model Is Not
❌ Not a clean official base model
❌ Not a from-scratch fine-tune
❌ Not a distilled / Turbo checkpoint
❌ Not a text-to-video model — use it with an input image
❌ Not a perfectly neutral LTX-2.3 checkpoint
❌ Not intended for SFW-only use cases
❌ Not guaranteed to fix every hand, face, or temporal artifact
❌ Not for minors or illegal/non-consensual content
🛠️ Suggested Workflow
Start with a clean, readable image.
Load the checkpoint in your LTX-2.3 ComfyUI workflow (
ComfyUI/models/checkpoints/).Optional: load the distill LoRA for the fast workflow.
Use a short, motion-focused prompt.
Without LoRA:
40 steps, CFG4.0· With LoRA at0.8:8 steps, CFG1.0, second pass4 steps.Test a few seeds before changing the whole workflow.
🎞️ Coming soon: a full ComfyUI workflow showing exactly how the gallery example videos were made will be added to this page.
🙏 Credits
Lightricks — LTX-2.3
TenStrip — 10Eros LTX-2.3 work
The creator of the sulphur LTX-2.3 checkpoint
The ComfyUI & LTX-2.3 workflow community
Adjustment / release: SeeSee21
📜 License
This model is a derivative of LTX-2.3-based checkpoints. Please respect the licenses of the original LTX-2.3 model and all source resources. LTX-2.3 is provided by Lightricks under the LTX-2 community license — commercial users should check the official terms, especially the revenue threshold requiring a separate commercial license.
🔗 Links
🎬 LTX See Motion — a personal LTX-2.3 adult motion blend for ComfyUI video generation.
FAQ
Comments (27)
I'll upload more examples and videos later 😊
Update! I'll be uploading a few more with workflows, such as video-to-video, etc.
📺 Video Guides / Walkthroughs
Step-by-step walkthroughs for every path this workflow can do. Each one is a short screen recording with on-screen chapters.
Image to Video — the basics: turn a single image into a video with natural motion.
https://civitai.red/images/136919778
Video to Video — extend any existing clip to any length you like.
https://civitai.red/images/136922039
Audio to Video — feed a 30 s song (with vocals) + one image and get a 30 s music video, lip-synced to the track.
https://civitai.red/images/137063968
IC-LoRA to Video — drive motion, poses or depth from a control video using an IC-LoRA (like the one shown in the clip).
https://civitai.red/images/137075598
LoRA + Image to Video — multi-subject / character-sheet consistency. Needs the LTX-2.3 Licon MSR V2 LoRA.
Guide: https://civitai.red/images/137093666
LoRA: LTX-2.3-Licon-MSR-V2.safetensors
Multiple-Subject-Reference LoRA by LiconStudio — keeps multiple characters and the background consistent across frames. Load it alongside the distilled LoRA.
We're gonna need more VRAM lol
Seriously AMD and Nvidia need to stop being cheap assholes and just release AI-driven GPUs next gen with at least 48gb of vram...
If they want to keep selling shitty gpus with 8gb of vram for the lowest class, go ahead, but we need MORE VRAM in consumer gpus at affordable prices... 48gb of vram costs nothing even with the insane prices. They could easily do it.
I completely agree with you. That said, I can tell you that the model runs with 8GB of VRAM if you have enough system RAM. But there’s nothing wrong with what you said.
Runs perfectly fine on my 3080 12GB card, although i wouldn't push the resolution pass 720p, but anyway great work on this, hope you can get a int8(convrot) version of this one.
This is quite a cool project, i just don't have the machine to create these, my docker containers already use up too much RAM :(
https://github.com/Starnodes2024/comfyui-starnodes-modelconverter
you seem to specialize in creating your low-effort AIO models, often not even normal merges, and have been using your favorite LLM to write everything instead of you, so awkward
Okay, now that’s what I call criticism that hits hard. 🙂
First of all, this is not an AIO model. But you are right about one thing: I do enjoy creating AIO models. That is far from the only thing I make, though—you can find plenty of other examples on my profile.
As for the texts, yes, I use AI to help me write them in English because my English is not fluent enough to express everything as smoothly as I can in my native language. I write what I want to say in my own language first, and then I use AI to translate and polish it.
Should I be ashamed of that? I don’t think so, and I’m certainly not.
Best regards to you as well.
下次直接用中文写,管他们个屁的
@SeeSeeLP i dont even understand the initial attack i'm just proppin' you for such a responsible and mature response. Downloading your content now just because you clearly deserve the support. cheers.
hi could you please add quants like nvfp4 and int8 ? fp8 is super slow and too big ...
I don't have any scripts for that yet, but I'd be happy to take a look at it.
can you put your bf/fp16 model for download. Want to create a int8/int4 version from it if possible
Sir,Excuse me, could you clarify the claim in the description that this model can run on 16GB of VRAM? The model file size is 31GB.
Thanks for asking! 🙂
The 31 GB refers to the checkpoint file size on disk, not the amount of VRAM required to run it.
LTX-2.3 supports CPU/RAM offloading, which means the entire model does not need to fit into your GPU's VRAM. Parts of the model can be kept in system RAM and loaded as needed during inference.
I've personally tested this model on an RTX 5060 Ti with 16 GB of VRAM, and it runs fine. It is even possible to run it on 8 GB GPUs if you have enough system RAM and use aggressive offloading, although it will be significantly slower and is not the recommended configuration.
So, the checkpoint being 31 GB does not mean you need 31 GB of VRAM. They are two different things.
I hope that clears things up! 😊
@SeeSeeLP Quick question related to this: does it not matter what the model size is(any model), and can the model still be partially loaded? When I tried loading FLUX.1-dev (22GB), it was silently crashing again and again. Is this expected behavior?
@Wendy_Earth Not really—there isn't a universal answer to that.
It depends on three things:
The model itself. Some models support partial loading/offloading much better than others.
The software you're using. Recent versions of ComfyUI introduced Dynamic VRAM, which significantly improves memory management and allows many models to run on GPUs that previously didn't have enough VRAM.
Your system. If your GPU runs out of VRAM, the remaining weights are offloaded to system RAM. If you don't have enough RAM (or your page file is too small), crashes or silent exits can happen.
So the checkpoint size alone doesn't determine whether a model will run. A 22 GB model may work on a 16 GB GPU, while another model of a similar size may not—it all depends on how the model is implemented and how efficiently your software can offload it.
If FLUX.1-dev is silently crashing, I'd first check:
How much system RAM do you have?
Which ComfyUI version are you using?
Are you using the latest version with the new Dynamic VRAM optimizations?
Those details usually make a much bigger difference than the checkpoint size itself.
@SeeSeeLP Crystal Clear. Thank You
@SeeSeeLP SIR,Thank you for the explanation. My PC specs are 16GB VRAM and 64GB system RAM,its OOM,OMG!!!!. Maybe,I need new workflow,
Could you break down your workflow for me?
o(TヘTo)。
@wwwmanmoon103 Thanks for the info! 😊
I'm already working on a few ComfyUI workflows for this model. If everything goes well, I'll try to upload them before the weekend.
In the meantime, I'd also recommend downloading a fresh copy of the latest ComfyUI. Then simply copy your models folder and your user folder into the new installation and try running the workflow there.
A surprising number of memory-related issues are caused by outdated custom nodes or older ComfyUI builds, so a clean installation often fixes problems like this.
Hopefully one of those approaches gets you up and running! 👍
@SeeSeeLP THANKS!!!SIR。
Base on 10Eros!? Does this mean this is an automatic ComfyUI lock in such that this so called "LTX 2.3" model can even be used by the official LTX 2.3 pipelines?
I was looking for other NSFW base LTX 2.3 models that were truly open and not locked into one SD gen tool.
No extra installation is required. You can use it directly with the official LTX 2.3 workflows and pipelines in ComfyUI, just like a normal LTX 2.3 checkpoint.
Hi.
I've just done a few tests in my environment: 896x1152 10sec, single pass, full body shot (less pixels for face), both first frame & first-last frame
and your model indeed shows better face consistency than 10eros v1.2.
I tested with 4 different distill loras; rank72 condsafe, rank111(kijai), DMD256(tenstrip) and rank384(official)
Interestingly, the setting you recommend (rank384 - 0.8) actually shows quite decent result, while the rank384 lora is normally least recommended when it comes to maintain the face consistency.
Is official lora recommended for some specific reason? (If yes, Tenstrip's DMD256 might be a better choice?) Or rank72 condsafe with 0.5 strength is still the best choice for the face consistency?
Hi, thanks for testing it!
I tested several LoRAs with my LTX-2.3 model and kept coming back to the official rank384 at 0.8 because it gave me the best overall results.
I haven’t tested your exact combinations yet, especially DMD256 and rank72 CondSafe at 0.5, but I’ll definitely compare them.
@SeeSeeLP Hi.
I wondered if the model is specifically trained or configurated under official lora.
rank72 condsafe would definitely show better id consistency since it brings less change, but it means less motion, emotion, facial expression (especially when strength is set to 0.5).
📺 Video Guides / Walkthroughs
Step-by-step walkthroughs for every path this workflow can do. Each one is a short screen recording with on-screen chapters.
Image to Video — the basics: turn a single image into a video with natural motion.
https://civitai.red/images/136919778
Video to Video — extend any existing clip to any length you like.
https://civitai.red/images/136922039
Audio to Video — feed a 30 s song (with vocals) + one image and get a 30 s music video, lip-synced to the track.
https://civitai.red/images/137063968
IC-LoRA to Video — drive motion, poses or depth from a control video using an IC-LoRA (like the one shown in the clip).
https://civitai.red/images/137075598
LoRA + Image to Video — multi-subject / character-sheet consistency. Needs the LTX-2.3 Licon MSR V2 LoRA.
Guide: https://civitai.red/images/137093666
LoRA: LTX-2.3-Licon-MSR-V2.safetensors
Multiple-Subject-Reference LoRA by LiconStudio — keeps multiple characters and the background consistent across frames. Load it alongside the distilled LoRA.