He finishes in her mouth, and you actually see his cock spasm and contract while it happens — over and over, not one twitch. Then it runs out the side of her mouth and down the shaft.
This is the oral creampie / CIM finish specifically — not the money shot where she tilts her head back with her tongue out and he finishes onto her from outside. Here it happens inside, with him still in her mouth, and the contractions are the whole point.
Different thing, different tool. If the money shot is what you're after, there are good LoRAs for that and this isn't one of them.
Video + audio. It generates its own sound.
⚠ Two versions — pick the right tab at the top
version base model trigger v1.0 — MiniMax H3 MiniMax H3 CUMOUF v1.0 LTX 2.3 / 10Eros v1.5 CUMOUF
Same footage, trained separately for each base. They are not interchangeable — a LoRA only works on the model it was trained on.
The two models want completely different prompts. Civitai shows one description for the whole page, so both guides are below. Scroll to the one matching the version you downloaded.
▶ MINIMAX H3 VERSION
Quick start
Base model MiniMax H3 Trigger CUMOUF — first word of SHOT 1 Strength 0.5 Guidance 1.0 — H3 is guidance-distilled, raising it breaks the image Length 12 seconds Start image one already mid-action
Tested side by side at a locked seed: 0.5 and 0.7 gave identical motion, but the cum went unnaturally thick at 0.7. Nothing is gained above 0.5.
Above about 1.0 it distorts badly. H3 LoRAs adapt the feedforward layers as well as attention, so extra strength redraws texture instead of strengthening motion — at 1.5 you get twisted anatomy.
⭐ Read this first: write the prompt to match your picture
Look at your start image. Describe what is actually in it — the camera angle, where her eyes are, whose hands are in shot, what is behind them, his colouring. Then say what should change.
Anything you describe differently from the picture is an instruction to redraw it. And when H3 redraws, it re-derives the people too. That is when the face changes, the skin tone shifts, or the whole shot flips to another angle.
Four ways to break it without noticing:
you wrote picture actually shows what happens "close side view" a three-quarter view the camera flips and everything is re-derived "looking up at his face" her eyes are closed her face gets rebuilt "he puts his hand on her head" no hands in shot a hand has to be invented, and the people change with it "a bedroom, plain bedding" a couch the background is rebuilt, and it drags the rest along
Copying someone else's prompt is exactly how this goes wrong. It was written for their picture. Change the look line to yours and keep the action.
Structure
A look line ending in "The environment is constant throughout.", then SHOT 1: with the action, then Audio: last.
Close three-quarter view in a bedroom, warm lamp light, plain bedding behind them, tight framing, shallow depth of field. The man and woman from <Picture 1> in their original scene: <describe them>. Her face, her hair and his skin tone stay exactly as they are in <Picture 1> for the whole clip. The camera holds this one angle throughout. The environment is constant throughout.
SHOT 1: CUMOUF. He drives his hips forward and fucks her mouth, thrusting in and out over and over, hard and fast and deep, setting the rhythm himself. She stays where she is and takes it. He thrusts in deep and she gags hard around him, then he draws back and she breathes. He cums, still thrusting, and you see his cock spasm and contract over and over as it pumps into her mouth. The spasm is hard and repeats as he continues to cum. A light dribble of white cum spills from the corner of her mouth and runs down his shaft.
Audio: fast wet slapping, the bed creaking under them, his heavy breathing and low groans, wet muffled sucking from her stuffed mouth. Quiet room tone, no music.
<Picture 1> is H3's own reference tag. It works with plain image-to-video, not only the reference model.
What one small change does
change result He holds still anywhere in the prompt Nothing happens. She sits motionless for the whole shot. Never write stillness The spasm is hard and repeats as he continues to cum Much stronger contractions. This is the intensity dial. Leave it out for something subtler Thick white cum runs out and spills over her lips, sliding down the shaft Gallons of it. Use deliberately if you want it over the top A light dribble of white cum appears on the side of her mouth Realistic texture and volume a light dribble of clear cum Renders clear. Most men are closer to clear than white — this is the realistic option Drop the colour word On darker skin tones it drifts toward caramel She keeps her lips sealed and swallows on its own No cum at all — and it kills the spasm too the same line plus a dribble line Works fine. It is the swallow alone that suppresses it She moves her head back and forth along his cock She does all the work, he stands there He drives his hips forward and fucks her mouth He drives it. Make him the subject of the action sentences Gagging in the Audio: block She gags continuously, out of sync with everything Gagging in the prose, tied to a thrust It happens at that moment She coughs Can loop She coughs once Happens once
Three rules worth memorising
1. Audio: is continuous. Prose is timed. Anything in the audio block plays across the whole clip — that is where ambience belongs. A sound that should happen at a moment goes in the prose, next to the action that causes it.
2. Events happen in the order you write them. Speech, the finish, the aftermath — put them in sequence and H3 times them that way.
This is the best thing about H3. A cabin chime written into the same sentence as the finish landed on the spasm. On LTX I have re-rendered the same clip twenty or thirty times trying to get a line to land before the action it describes instead of after. Here it is one render.
A loud electronic cabin chime sounds twice as he cums, still thrusting, and you see his cock spasm and contract over and over as it pumps into her mouth.
Avoid spelling out noises — write "a loud electronic chime sounds twice", not the noise itself, or it may get spoken aloud.
3. Say whose sound it is. An unattributed "heavy breathing" gets given to whoever is most visible — which produced a woman breathing heavily with her mouth full. Write "his heavy breathing and low groans" and "wet muffled sucking from her stuffed mouth".
Never negate. "Less cum", "no spilling", "without going deep" all produce the opposite. Ask for the small version positively instead.
Setting and ambience
Only the look line and the Audio: block change. The action stays identical.
Audio: ocean waves rolling in and breaking on the sand, steady sea breeze, distant seagulls calling, fast wet slapping, his heavy breathing, wet muffled sucking from her stuffed mouth. No music.
If ambience buries the wet sounds, drop the wind first — it masks everything else.
Speech
In quotes, in the prose, at the point it should happen. It lip-syncs on its own.
He says, "swallow my load babe"
Known behaviours
Don't ask for something that isn't in the start image. If the prompt says he puts his hand on the back of her head and there is no hand in frame, H3 has to introduce one — and introducing an element is itself a transition, so it re-derives the whole picture to make room. It usually strikes a second or two in, exactly when the new element would have to appear.
Anything it has to redraw, it may re-derive. A camera move, a hand coming off, a pull-out. Anchor with <Picture 1> and hold the camera:
The camera holds this one angle throughout.
If the first second or two is a different person and the rest is good, just trim the front off. Not worth re-rolling.
Audio can push the picture. H3 runs audio and video through the same attention. A loud rhythmic sound cued to a moment can drive motion at that moment — an engine rev on the finish made the whole thing vibrate. Useful or not depending on what you want.
Foley grunts are weak. Breathing and speech come through well. That is the source footage, not the prompt. It is not a sound LoRA.
H3 training
Base MiniMax H3, pruned int8 Clips 1,734 × 39 frames @ 24 fps Rank / alpha 16 / 16 Steps 4,000 — about 2 epochs Trainer ai-toolkit Hardware one RTX 5090
H3 reaches the behaviour in far fewer passes than LTX did — a third of the training. Rungs above 4,000 got worse, not better.
▶ LTX 2.3 / EROS 1.5 VERSION
Quick start
Base model 10Eros v1.5 Trigger CUMOUF — first word Strength 1.0 CFG 1.2 – 1.3 Length 8 – 15 seconds Distilled LoRA ltx-2.3-22b-distilled-lora-384-1.1 Sigmas the longer schedule below — this matters
First-pass sigmas. Worth using: switching to this schedule improved the spasm, the detail and the audio noticeably over a shorter one.
1.000, 0.955, 0.893, 0.812, 0.715, 0.603, 0.482, 0.241, 0.121, 0.0
The prompt used for the example renders:
Performance: [CUMOUF. Close side view of his cock in her mouth, her lips wrapped
tight around the shaft. He pumps slowly into her mouth, then holds deep as he starts
to cum. You see his cock twitch and spasm over and over, arching and pulsing as it
pumps into her. Thick white cum fills her mouth and spills out over her lips, sliding
down the shaft.]
Sounds: [Man grunts loudly as he cums. Woman coughs and sputters trying to swallow it all]
Dialog: [(off camera) Man says: Take my load in your mouth. Swallow my cum.]
Note the three separate blocks — that matters, see below.
Three things that will waste your time if you don't know them
1. Don't ask for a man's voice without saying he's off-camera.
LTX-2.3 lip-syncs. Ask for a line and it needs a face to put it on — so it drags one into your shot, and once there's a second mouth near a cock it does the obvious thing with it. Cost me an hour to work out.
Dialog: [(out of frame) Man says: swallow it]
2. Sound goes in its own block, or it gets spoken out loud.
I wrote "he grunts loudly" into the visual description and got a render where the man said the words "he grunts loudly". A grunt is a facial action, so describing it in the picture asks for a face doing it.
Performance: [ the action — no sounds, no speech ]
Sounds: [ wet mouth sounds, grunting from the man, moans from the woman ]
Dialog: [ (out of frame) Man says: ... ]
3. pumps means ejaculating here, not thrusting.
It's in 82% of the training captions as "as it pumps" — the cock pumping cum. So "he pumps fast in and out" asks for the climax while you're trying to ask for the build-up. Use thrusts for motion.
Words it knows
Measured across all 1,525 captions:
word in cock 100% spasm 79% contract 79% hard (intensity) 48%
Zero occurrences — these fall through to the base model: penis, dick, load, spasms (with an -s), twitches, jerks, arch.
That last group matters. spasm and spasms are different tokens. The trained phrase is "spasm and contract over and over" — use it as written.
It moves in for the finish — that's the point
Start from a normal, wider shot. When the finish begins, the camera pushes in on its own.
That is the behaviour, not a side effect. It's a finisher, and this is what a finisher should do — the same move an editor would make. You stay in the scene through the act, and when it matters the shot tightens to show you the thing you came for.
There's a physical reason it has to. A spasm is a few millimetres of movement. From any distance it's sub-pixel and simply doesn't exist on screen — you can't see one from across the room. Every clip I trained on was cut tight for exactly that reason, so the model composes the same way: it frames for its subject.
To get more of the move: start wider, and state the spasm phrase twice. To get less: start closer, and state it once. A start image already cropped tight gives it nowhere to go, and puts more pixels on the spasm from frame one.
Your start frame decides what his cock looks like
Raised in the comments, and it's a good catch — thank you.
The model only knows the shape of something from what it can see in frame one. If it's already halfway into her mouth, the head was never visible, so when she comes off him it has to invent one — and it invents an ordinary one. If yours is distinctive at all, that's where the detail goes.
So start with it near her mouth, not yet in it. One clean look at the shaft and head is enough and it holds for the rest of the clip.
The trade-off: a frame already mid-act starts moving sooner, because this continues an action far better than it starts one. About-to-go-in is the sweet spot — clear anatomy, and something already happening.
It will pull back at the end. That's trained in, not a fault: a lot of the dataset is the finish continuing after she comes off him, cum running out of her mouth and down the shaft. To push against it, keep the render at 8 seconds so it ends during the finish, write that she holds him deep and swallows, and say nothing about it spilling down the shaft. None of that is reliable — showing it clearly in frame one is.
To get it out of her mouth: start from a frame where he's already out, or say she is looking at it and holding still.
Strength — 1.0, and the range is narrow
Alpha equals rank, so 1.0 is genuinely the LoRA as trained.
0.7 clean but flat 1.0 use this 1.5 stronger, but body morphing appears
If the spasm is weak at 1.0, more strength won't fix it — that's a data limit.
Known limits
The audio is trained, not the base model's. It generates its own sound — wet mouth sounds, breathing, and grunts that came from the source footage rather than the gritty default. Use the longer sigma schedule; it makes a real difference to the audio too.
It's a finisher. 1,169 of 1,525 clips are the climax itself; only 93 are in-and-out motion. It won't drive face-fucking or a long build-up — use another LoRA for the act and bring this one in for the finish.
Anatomy holds up but isn't perfect. Nine continuous seconds in her mouth and it comes back recognisably the same — that used to render as chewed-up meat. It still stretches sometimes on the way out, and it can only keep what your start frame actually showed it.
Long renders drift. Trained on 2-second clips. 8–15 seconds is the tested band.
LTX training
Base 10Eros v1.5 bf16 Clips 1,525 × 49 frames @ 25 fps Sources ~14 couples plus solo footage, every window timed by hand Bucket 640×384×49, rank 8 / alpha 8 Branches video + audio, trained jointly Hardware one RTX 5090 Steps trained to 12,000; released checkpoint is step 11,000
Later is not automatically better. I tested the whole ladder and 11,000 was the best trade between spasm strength and stable anatomy — the last checkpoint was not the one I shipped.
How both were made
Every clip was cut to hand-written timestamps, and only to seconds where the spasm is genuinely visible. Where I couldn't see it — a hand in the way, a body blocking the shot — that footage contributed zero spasm clips even though the event was happening.
Every clip in a group carries one identical sentence for the event. Variety comes from the scenes, not the wording. That's the opposite of the usual advice, and it's what made this version work after eleven that didn't.
Full method notes, traps and measurements are published alongside this — please copy any of it. I learned a lot from other people's write-ups while building this, so everything I found is written down and free to take.
Description
Quick start
Update: Verified works well with LTX 2.5
Base model 10Eros v1.5 Trigger CUMOUF — first word Strength 1.0 CFG 1.2 – 1.3 Length 8 – 15 seconds Distilled LoRA ltx-2.3-22b-distilled-lora-384-1.1 Sigmas the longer schedule below — this matters
First-pass sigmas. Worth using: switching to this schedule improved the spasm, the detail and the audio noticeably over a shorter one.
1.000, 0.955, 0.893, 0.812, 0.715, 0.603, 0.482, 0.241, 0.121, 0.0
The prompt used for the example renders:
Performance: [CUMOUF. Close side view of his cock in her mouth, her lips wrapped
tight around the shaft. He pumps slowly into her mouth, then holds deep as he starts
to cum. You see his cock twitch and spasm over and over, arching and pulsing as it
pumps into her. Thick white cum fills her mouth and spills out over her lips, sliding
down the shaft.]
Sounds: [Man grunts loudly as he cums. Woman coughs and sputters trying to swallow it all]
Dialog: [(off camera) Man says: Take my load in your mouth. Swallow my cum.]
Note the three separate blocks — that matters, see below.
Three things that will waste your time if you don't know them
1. Don't ask for a man's voice without saying he's off-camera.
LTX-2.3 lip-syncs. Ask for a line and it needs a face to put it on — so it drags one into your shot, and once there's a second mouth near a cock it does the obvious thing with it. Cost me an hour to work out.
Dialog: [(out of frame) Man says: swallow it]
2. Sound goes in its own block, or it gets spoken out loud.
I wrote "he grunts loudly" into the visual description and got a render where the man said the words "he grunts loudly". A grunt is a facial action, so describing it in the picture asks for a face doing it.
Performance: [ the action — no sounds, no speech ]
Sounds: [ wet mouth sounds, grunting from the man, moans from the woman ]
Dialog: [ (out of frame) Man says: ... ]
3. pumps means ejaculating here, not thrusting.
It's in 82% of the training captions as "as it pumps" — the cock pumping cum. So "he pumps fast in and out" asks for the climax while you're trying to ask for the build-up. Use thrusts for motion.
Words it knows
Measured across all 1,525 captions:
word in cock 100% spasm 79% contract 79% hard (intensity) 48%
Zero occurrences — these fall through to the base model: penis, dick, load, spasms (with an -s), twitches, jerks, arch.
That last group matters. spasm and spasms are different tokens. The trained phrase is "spasm and contract over and over" — use it as written.
It moves in for the finish — that's the point
Start from a normal, wider shot. When the finish begins, the camera pushes in on its own.
That is the behaviour, not a side effect. It's a finisher, and this is what a finisher should do — the same move an editor would make. You stay in the scene through the act, and when it matters the shot tightens to show you the thing you came for.
There's a physical reason it has to. A spasm is a few millimetres of movement. From any distance it's sub-pixel and simply doesn't exist on screen — you can't see one from across the room. Every clip I trained on was cut tight for exactly that reason, so the model composes the same way: it frames for its subject.
To get more of the move: start wider, and state the spasm phrase twice. To get less: start closer, and state it once. A start image already cropped tight gives it nowhere to go, and puts more pixels on the spasm from frame one.
Your start frame decides what his cock looks like
Raised in the comments, and it's a good catch — thank you.
The model only knows the shape of something from what it can see in frame one. If it's already halfway into her mouth, the head was never visible, so when she comes off him it has to invent one — and it invents an ordinary one. If yours is distinctive at all, that's where the detail goes.
So start with it near her mouth, not yet in it. One clean look at the shaft and head is enough and it holds for the rest of the clip.
The trade-off: a frame already mid-act starts moving sooner, because this continues an action far better than it starts one. About-to-go-in is the sweet spot — clear anatomy, and something already happening.
It will pull back at the end. That's trained in, not a fault: a lot of the dataset is the finish continuing after she comes off him, cum running out of her mouth and down the shaft. To push against it, keep the render at 8 seconds so it ends during the finish, write that she holds him deep and swallows, and say nothing about it spilling down the shaft. None of that is reliable — showing it clearly in frame one is.
To get it out of her mouth: start from a frame where he's already out, or say she is looking at it and holding still.
Strength — 1.0, and the range is narrow
Alpha equals rank, so 1.0 is genuinely the LoRA as trained.
0.7 clean but flat 1.0 use this 1.5 stronger, but body morphing appears
If the spasm is weak at 1.0, more strength won't fix it — that's a data limit.
Known limits
The audio is trained, not the base model's. It generates its own sound — wet mouth sounds, breathing, and grunts that came from the source footage rather than the gritty default. Use the longer sigma schedule; it makes a real difference to the audio too.
It's a finisher. 1,169 of 1,525 clips are the climax itself; only 93 are in-and-out motion. It won't drive face-fucking or a long build-up — use another LoRA for the act and bring this one in for the finish.
Anatomy holds up but isn't perfect. Nine continuous seconds in her mouth and it comes back recognisably the same — that used to render as chewed-up meat. It still stretches sometimes on the way out, and it can only keep what your start frame actually showed it.
Long renders drift. Trained on 2-second clips. 8–15 seconds is the tested band.
LTX training
Base 10Eros v1.5 bf16 Clips 1,525 × 49 frames @ 25 fps Sources ~14 couples plus solo footage, every window timed by hand Bucket 640×384×49, rank 8 / alpha 8 Branches video + audio, trained jointly Hardware one RTX 5090 Steps trained to 12,000; released checkpoint is step 11,000
Later is not automatically better. I tested the whole ladder and 11,000 was the best trade between spasm strength and stable anatomy — the last checkpoint was not the one I shipped.
FAQ
Comments (18)
Nice job. This made me stop using MiniMax for a second. It works well. I posted a video but it was flagged for review so I guess will see it later. LTX2.3 definitely won't go away.
Ill look at doing same thing for MiniMax. Thanks!
i know, right????????????????????????????????? I think this is the first oral creampie on Civit AI, actually congrats!!
@The_Last_Goblin_King I made it for that reason. I couldnt fine one myself :) Thank you!
Despite what you or other think. LTX2.3 ad Wan22 will remain the main model for NSFW content. for one simple reason, it's a lot faster. Even with Lightx2v lora, H3 is slow for genrating prduction ready HD videos.
So, nohing will change this in the futur, LTX2.3 and Wan22 will remain the faster way to genrate HD NSFW content for PAtreon, OFM PPV.
not to mention, that native Wan22 face/body consistency reamain unbeaten for now.
@Pat3dx I used LTX2.3 exclusively. I haven't used it all except to make this lora's video since Minimax was released. The only time I'll use LTX2.3 is to trying out interesting Loras. I don't need to debate it though. Just look at stablediffusion reddit to see what's preferred. MiniMax videos posted number about 100 to 1 vs LTX2.3. When a good model arrives on the scene people use is (Like Krea2). Why debate though? if LTX2.5 trumps MiniMax I'll use it. I've made 10's of thousands of videos with both Wan and LTX. (hunyuan before that). MiniMax is on another level. I have no "hate" for any model. I use what I like as should you. Wan still has some kinky fetish stuff not out on MiniMax yet.I support all open weight models but my god MiniMax is good. I'm looking forward to the next LTX release.
@iodrg244 Another level is an understatement. This is game-changing.
@Pat3dx Lol? I guess you haven't tried H3 enough...
same for H3 please 🙏🙏🙏🙏🙏
I’ll release it tonight! On the last bit of training now
@LORAGEEK 😍😍😍😍 you are the goat! doing the lords work right here!
Hey Azra, its out. Check the writeup for how to use the H3 version vs LTX. Weird they cant have same page but different "tabs"
@LORAGEEK hey, yes i already downloaded. will definetly try it out! thanks a ton, you da best!
Great lora! hope you can make one for a creampie as well
Very good lora, well done :).
The only issue i noticed, is that even when the start frame the cock is already in her mouth, at the end, she move her head back and the cock comes out.
This is a critical issue, as LTX don't know what was the shape and lengh of the cock, the glans is very bad.
Example, my first frame the cock is already in her mouth, this is a futa huge black cock, and half of the cock is in the mouth.
At the end, she moves her head back, cock comes out, but the cock is no longer a huge black cock.
So, my advice, never start with a frame with the cock already inside the mouth, if you want perfect consistency with the shape of the cock/glans.
Start with the cock, near her mouth and not yet inside, that way LTX2.3 perfectly know the shape/lengh of the cock, especially the glans.
And at the end when the girl move back her head and cock comes out, it will be perfect consistency.
Or is there a way to make the girl at the end, don't move back her head, and keep the entire cock inside her mouth?
Thanks and that's a really good catch. You're right about the cause.
LTX only knows the shape of something from what it can actually see in frame one. If half of it is already in her mouth, it never saw the glans, so when she pulls back it has to invent one and it invents an average one, not yours. Nothing about the start frame tells it "huge and black" if the huge black part is hidden.
So your advice is right and I'll put it on the page: start with it near her mouth, not in it. One clean look is all it needs and it stays consistent the whole way through.
On keeping it in her mouth at the end the pull-back is partly trained in. A good chunk of the dataset is the cum running out of her mouth and down the shaft, so it wants to end with it visible. You can push back against it:
- shorter render (8s) so it ends during the finish rather than after it
- say she holds him deep in her mouth and swallows, and leave out anything about it spilling down the shaft
- drop the strength a little
But your own fix will give you better consistency than any of that. I'm noting it for v2 a group of clips where she stays down through the whole finish would make "stay in" an actual thing you can ask for.
@LORAGEEK yes, this is not a particular issue specific with LTX2.3, wan 22 and H3 have the same issue, and i think, all futur model will also have this issue, it's a normal issue. Model can't know something it can't see.
But my first test where the girl have half of the glans inside her mouth, was a success. Even at the end when she pull her head back, the consistency and glans shape was almost perfect. :)
Check it, i just uploaded the video on your page.
@LORAGEEK thanks for the H3 version, will test it this afternoon :)