PinkCherry MM H3 (Beta)
this is the fl2va model, first frame last frame. plugging it the ref2va reference workflow will yield poor results likely, use the comfyui ref2va model for that. it wasnt trained on ref2va data
Beta 0.6
im using this workflow https://huggingface.co/Plaguekind/Minimax-H3
** beta improves t2v and motion, but its not perfect. still needs further training
** I post new snapshots on my huggingface periodically for testing. for those that want latest checkpoint we are testing. on beta-0.6 at the moment
** Aug 09 2026 - Still deep in training this model, next update should be a good one with lots of major fixes. but will not likely be ready until later this week
** Aug 08 2026 - received some great feedback/comments on weak areas in the checkpoint, and ill work on targeting those weaknesses with further training. thanks for all the input and please keep it comming so I know what you are seeing/hearing.
** Aug 8 2026 - both unpruned (32GB) and pruned (20GB) versions of 0.5-alpha are uploaded. Wont be any new versions of this model for at least a week. as im working on the ref2va model now
Join our discord: https://discord.gg/GCrr4Cj3D
Loras: any lora works with this base. We have had good luck using pussy synth, riding and coach bates penis lora at about str 0.4 to address the faults in this alpha release base.
I2V/T2V: Your biggest win will be with image to video, not text to video at this stage of training this model.
Model Type: First Frame/Last Frame FL2VA , however it worked fine in the reference workflow for me. Ill train ref model later
Model Training: Full fine tune of fl2va base model
Training Dataset: 5000+ 4k videos of various concepts
Model Training Target: FFN/Attn blocks only
Training System: Custom Training Pipeline
no guardrails or censorship has been bypassed in this model
This is early still in training model I am working on. Treat it as alpha level still. its the first last frame model not the reference model.
Fair Warning: Expect to find dicks on foreheads, exploding breasts and various other issues typical of alpha models. If that's not your thing, wait for beta.
Training Captions used: cock, pussy, vagina, penis, doggystyle, missionary, slut, moaning, orgasm, wet, hairy pussy, spreading, asshole, tight, fucked, riding, cowgirl, anal, reverse cowgirl, pounding, thrusting, hanging tits, fondling, massage, squeezing, dildo, sex toy, penetrate, squats, flexing, bent over, sucking, licking, tongue, blowjob, arousal fluids, jiggling, swaying, sex position, prone, nude, clothing, saggy, breasts, big tits, legs spread, panties pulled to side, cum, semen, creampie, saliva, masturbating, fingers, fingering, eyes closed, perky tits, rubs, lustful, areolas, pink nipples, mouth open, topless, exposed, voyeur, public, bounce, labia, folds, slut, whore, POV, overhead, low angle, high angle, close up, dripping, precum, round ass, amateur, glistening, facing camera, standing sex position, vulva, slutty, tattooed, pounds, from behind, hanging, stroking, jerking, pants down, panties, cock head, cock tip, wet cunt, passionately , rough, gangbang, group fuck, lesbian, threesome, handjob, double blowjob, throbbing, rapidly, rhythmically, wet lips, cleavage, jiggling breasts, cupping, tittyfuck, teasing, erotic, romantic, making out, blindfolded, thick, crotch, deep, steady, aroused, erect, soft flesh, slamming, messy, slapping, pink anus, bathroom
Description
better fucking
FAQ
Comments (80)
Any other samples for a 31gb model?
yes ill work on some, but it will be a few hours
Is it better than your Lora?
yes, much. They are not related models at all. The lora was trained on ai toolkit on a distilled base and is a bit "well weird". this is a fine tuned model. but its alpha and still being trained as we speak, so treat it as such, keep expectations low :)
Wow this looks incredible! Thank you for the work as always!
I just noticed the filesize though.
Any chance for a version of the model that can fit on 24GBs of VRAM?
I use the Pruned version of the H3 because it can fit on 24GBs of VRAM (like 16GBs size), so maybe the same technique could be applied??
you can run it on 24GB in ComfyUI, comfyUI has block swapping it will swap some of the blocks to your CPU. I can run it on mine. In terms of the pruned model, something ill have to investigate but cannot right at the moment.
@sexgod1979 aye keep up the good work and keep hydrating.
@sexgod1979 Oh interesting! Ill givs it a download and check it out. Thanks!
And as the other person said: Stay hydrated 😂 !
Thanks for all you do
I am convinced SexGod1979 is an actual god. :) Amazing work sir!
thats what my wife keeps telling. kidding, I gave this name to myself
@sexgod1979 lmao
Minimax is a good director; he creates interesting camera movements and shot transitions. But this —I'm sorry, this looks really bad.
I did say it was alpha level quality at this stage :)
Have you tried it? You don't lose any of the benefits of Minimax H3, but it adds some great motion and while still WIP finer details, it still knows the concepts. Prompt as you would for the base model, but include words like cock, pussy, etc. See what happens :P
soon i'm going to have to sue. my meat has never been more sore then today and you release more models to keep me from being productive.
I think you win the internet award for today my friend
Buddy, you probably got the sites mixed up; you need to go to Grok.
lol
looking forward to r2v
Not Pruned and nvfp4 missing... but waiting for ref2vid anyway
thanks for your work, eventhough its alpha, its getting good!
How long can videos be when generated with MiniMax H3? Longer than with LTX 2.3?
As i know after 15 sec quality may start to decay
it is trained on clips of 5-15 seconds
Since the official support is 15 seconds, I have still seen cases with more than 20 seconds
i made 20 second clips already without any issues in i2v and ref2v
Thanks a lot for your efforts. Since you're about to train a fresh checkpoint, may i ask for including not only unnatural big "digggs"? Only if possible of course. Would like to test more "realistic" videos. :-)
Pretty sure MM are going to start complaining to Civit if people keep using their company name for NSFW stuff, so you should probably just use "H3" in the name...
ok
Minimax is complaining about NSFW content with their name attached to it, you should just change this to H3 and drop the minimax to avoid problems for the whole community
ok
@sexgod1979 ok
Source?
@ProvenFlawless Yeah, I'd like to see a source too.
its funny cause the only reason AI Image and Video generation is still relevant is because of porn. Even grok subscribers jumping the ship.
@ProvenFlawless @EpicStuffs
There is no problem making NSFW loras etc for H3, problem is if the company name Minimax is used because p0rn is illegal in china they do not want to associate their company name to any NSFW. so this is ok : "amazing H3 nsfw lora v1" this is not ok: "amazing minimax H3 nsfw lora v1".
Sources:
https://www.reddit.com/r/StableDiffusion/comments/1vfwijz/minimax_are_issuing_takedowns_on_decensorexplicit/?rdt=46089
https://huggingface.co/ethanfel/Qwen3-VL-32B-Ultra-Heretic-H3-ComfyUI-INT8-ConvRot/discussions/7
Any pruned version?
"I wont be releasing a pruned version of this model at this time for alpha."
@spidersweater I will at some point, just trying to get the model in better shape first
@sexgod1979 sorry, I didnt read the full intro, looking fowward to you later nice works!
@Baiyixi there is a pruned version as of today on 0.5-alpha. there isnt for previous versions though
I absolutely love how quickly the community got straight to work on this stuff. Can hardly wait for it to download.
yeah I still have a lot of training to do on this one, but im willing to put the effort in. training is running right now...
thank u, This model also works well for R2V
yeah nice one
The movement looks weird
working on it :)
Not bad for alpha
You trained a finetune on 5000+ 4k videos in this time? I have strong doubts
The model is under trained at this stage, I only stated how many videos are in the dataset. Sorry for the confusion :)
someone is butthurt they were not first.
I've tried Alpha v0.2 and v0.3 and I must say that the motion is limited in speed and strength.
It is definitely SLOWER and less POWERFUL in the action compare to the normal Model, try Missionary or Doggy and you will notice that even if you try as much as you can in the prompt to describe FASTER and HARDER / STRONGER impact... there is a limit to the motion speed which is NOT good enough.
SUGGESTION:
If you can add more dataset for more INTENSE, FASTER, POWERFUL / STRONGER actions and not just slow and gentle, please add it, thanks ahead! 🙏
thank you, and I agree with all the above. working on it now :)
we have been using the synth pussy and riding lora at like 0.4 str and it pumps up the action until I get this issue resolved
@Virtual This might help out in the meantime, maybe in lower strengths, since the result looks a little crispy to me:
https://civitai.red/models/2840146/faster-harder-shake-harder-or-h3-motion-booster?modelVersionId=3205912
@Jellai great idea. Im working on motion now in the base model
@sexgod1979 That's awesome!
Please also make sure MORE NATURAL BOOBS physics and NOT FAKE because it trains better physics ❤️
@JellaiThis is VERY mechanic and unnatural, I tried it... sorry not a big fan.
The BASE of what we already have in PinkCherry is already AMAZING! it's only adding extra NATURAL BOOBS and more POWER to the strength and FAST since MiniMax-H3 can accept it VERY WELL much better than LTX 2.3 which have lots of physics issues that even hard train didn't solve... the important thins IMO is to train on NATURAL BOOBS and not FAKE because physics may look like PLASTIC or JELLYFISH just like the LoRA you showd.
Currently, the results are similar to the LTX 2.3 model. Next time, please show videos that showcase the H3's strengths: camera handling, changing shots, and overall consistency of the video sequence in a more complex scenario.
good idea. yes, I am still learning how to maximize this model.
yes that is what I also love about it. it's not just 1 scene anymore. you can create little stories with multiple shots and more complex interactions in a single video =)
this guy cums with gud idea
too big cabrón 🤌
I've been using your model, and while it's good, I want to echo what someone else said on here ---- ive noticed h3 has specific things it can do that absolutely kills ltx 2.3, and certain areas where ltx kills h3.
h3 is far too safe a model, and i don't mean uncensored, uncensored is besides the point. h3 tends to be very happy, upbeat, by default. And it resolves things in a too cleanly, this means that content can lack of energetic aggression that ltx 2.3, 1.1 has. h3 never has that frantic, sweaty, aggressive, aspect. It always dails down things to make them smoothed down.
That smoothing down is a large reason why your ltx checkpoint worked, because ltx was too course, it would fail 90% of generations because things would blister and wart and generally fall apart. You kept things rock solid and stable.
I believe this model needs the exact opposite treatment to how you worked with the ltx model, instead of grounding things this model needs to be made more aggressive, unhinged, and generally frantic. That will pull it away from it's worse tendencies, which is to always make things more passive and safe in terms of energy.
I truly do believe this training requires a completely different approach in order to be incredible. But the room for that incredible is large, because h3 has a lot going for it.
thats a pretty interesting point, what sort of dataset would you envision would do that? I assume what you are getting really is that its too timid and safe, it needs a shake up. Its not gritty or wild. really interesting honestly. Tell me more
@sexgod1979 Well, run ltx 2.3 1.1 (im sure you have already due to building from it), and take things up a notch, turn that config up, put some complex sexually psychopathic prompt in there, do things like focusing on hatefucking, or general impersonal sex, and you'll notice that things tend to get really wild in ltx, not always in a always controllable way, but there's beats of sweat, peoples faces either close down completely, distort into statue like emotionless figures, yet still are capable of sexuality. These juxtapositions and conferencing aspects LTX can really pull off. Better than even porn itself in many ways.
If you do the same thing with h3, it goes to about 25%, and never spills over, there's this cap of emotiveness and wildness. So she'll always be a bit too loving, or a bit too goofy (if your starting picture is just a normal picture) there's none of that hyper organic dark edge to things.
Except -- when your starting image is that tonality, that tends to shift the model to adapt to it well, which is why the ref>vid, and vid > vid is so good in h3, i believe the way it interprets references is a large part of how the entire model shades things. I'm not sure if that works the same with trained images, but i assume it's not far off.
That might be a good starting point with how to test data. I would start with image tonality being in the more wild, grim, almost distorted end of the spectrum. None of that safe "flux 1" over rounded imagery, instead go for the most organic and broken looking starting content, because my assumption is that the model will still want to keep bringing things back to that "rounded down point" if the data is on the more wild, intense, aggressive, you know, hair all over the place, spit, cum, wild looking eyes, mouth being stretched, hand around throat, perhaps some of the images even having a pass over them to push that further than the original starting image (you know, edit models, so you can keep the original realistic soul of the image, but then add extra details via the model, klien9b is good for this) that kind of thing taken to it's upper edge, then h3 will be normalizing from an elevated reference, and it would allow the model to not fall back to it's worst tendency.
Now, i'm not exactly sure how this would pan out. But it would be at the very least, interesting. To me at least. To see how things respond, and potentially build on that, if it's rewarding.
https://civitai.red/models/2839440/hand-in-panties-masturbation?modelVersionId=3205014
look at the comparison image of this masturbation example in this lora, notice how the second one has a darker shade, is more aggressive? more lifelike? I think that's what the end result might be if the model was leaning more like that, except in all areas, not just the specific lora example. h3's major sin is it's low energy.
@whitespider9999 I think it might be prompt-related because I certainly have not shared that experience
@BBBAAA2 could be a few things, i've been doing pretty high level prompt testing across a variety of different methods. So it could be a sensitivity thing, for example many people are fine with "flux" type of hyper clean resolves and really cleanly resolved and dialled down motion. Alternatively you could just prompting a lot better than me.
I can say h3 is - from my experience, very clean/non organic from image or text source. this suits people that like that "ai look", but its far from ideal for people that don't (scan the comments on h3 on the interwebs and even in this comments section)
It's video > video is a different story however, and i find that it's quite capable.
@whitespider9999 I think the LORA you linked too makes it look more fake. I guess to each their own. I find H3 will reproduce what you prompt it for. If you want more energetic action you need to explicitly prompt for it.
@whitespider9999 A ton of the people in this group are the kind of people to use turbo lora's and call it overall wins. They don't know what they are talking about. Agree with most of what you said, with a few exceptions, agree with things being safe with h3, agree with motion. I don't think the lora you included was the best example. But yes, I've always had a problem with how h3 tends to make things slow, safe, and rounded down without any quirk or intensity. I also realized that a bunch of people on civai tend to like that rounded down overly polished overly sterile AI look, they think it makes things look more realistic and sexy. I've been using the full model on topview ai, and it's the same way, it filters away all the quirky edges. Great that it's uncensored, but it's not the model that some say it is, yet anyway. It needs exactly like this checkpoint to come along and shake things up.
chance of a lora version? for both flf and ref modes?
looking into that, there is a pruned version of alpha 0.20 that is smaller. Let me look into the fl2va lora and ref2va request
can you add jiggle physics and motion physics like GROK ?
default jiggle is good