Minimax H3 INT8/INT4 Convrot
Sampler settings I currently use are: er_sde/res_multistep simple
Just uploaded both FL2VA & REF2VA w4a8_mixed models quantized by Kijai. (Makes my mixed int4 models obsolete) True int4 model size with int8 activations, near int8 quality, same speed as int8. For RAM/VRAM-constrained systems. Update your ComfyUI to the latest version for support of the new w4a8 (Should be included in Stable ComfyUI v0.31.0 https://github.com/Comfy-Org/ComfyUI/pull/15308
FL2VA - first last (frame) to video / audio
REF2VA - ref_images / ref_videos / ref_video_audios / ref_audios: up to 9 reference images, 3 reference videos (each may carry its own paired soundtrack), and 3 standalone reference audio clips
Both models can generate t2v (Text to video), i2v (Image to video), v2v (Video to Video), a2v (Audio to video), and multiple references (image/video/audio). But were further fine-tuned/trained for higher-quality outputs for the intended use.
Description
FAQ
Comments (128)
158 sec to make a: 0.5 megapixel, 5 seconds video with a 5800.
cool.
rtx5080?
@Novastar27 Ah, I got fat fingers. RTX 5080, correct.
and 1040 seconds 0.3 megapixel 6 seconds on a RTX3060 12gb
@gambikules858 300s for 0.3mpx 6s with a RTX4070super 12gb
~180sec for a 10sec video at 0.4mpx, 20 steps. RTX 5070 TI 16GB
@BonerSoup More Vram is better
8 minutes for 1mpx video of 5 seconds on my RTX 4070, classic one, 12GB vram. (I have 128GB of RAM). 20 steps, with sage attention.
im Gonna have fun on my RTX 6000. this is very good quality
dude ivde been using my 6k all fuckign day with this. I know it's illegal to use this model in the USA, but fuck that noise; this is awesome!
@The_Last_Goblin_King Minimax h3 is illegal in the USA?
@mddw no it's not. he's tweaking
6 minutes for 10sec at 12fps 320x600p on laptop rtx-A2000- 8GB vram , but loaded the 64GB ddr5 ram to 58GB/64GB! The results are definitely worth waiting for, compared to both LTX and WAN.
Yes, in my opinion, this is the most powerful local video model available right now. I'm incredibly impressed by the animation quality, the anatomy, and the facial stability. This model is literally better in almost every way compared to the LTX and the WAN. Perhaps the LTX is slightly better in terms of realism. But oh my god, the LTX's facial and finger stability are so poor; it takes a lot of effort to get a halfway decent result on the LTX.
@Voxe1 Yah, and it simply does what it is told!..if you craft your prompt well, it sticks to it.
Anyone here with a 5090? I have one and 64gb ddr5 with gen 5 crazy fast ssd, I dont think im getting the crazy fast speed si should be getting.
0.3 megapixel at 10s generates in about 180 seconds with 20 steps res_multistep.
Are any of yall getting that too?
use spectrum node
and install sage attention 2+
I have a 5090 laptop w/64gb ram. 15 sec video takes me about 12 mins. I'm using the MiniMax Sage Attention node.
I'm using a 5090 with 96 gigs of ddr5 and a gen5 ssd, I'm getting 5 second 480p videos in around 1.30 minutes, I've used the nodes people recommended that gets it down to 40 second but the quality loss is not worth it
My GPU turned into an airplane engine
yeah! Same here. WOOOSH WOOOSH
I've never used nvfp4. Am I right that it's only for the RTX 5000 series? Or am I confusing it with something else? I'm talking about the text encoder.
Can I use it on my RTX 4070?
I've run it on my RTX 4070 super and it works so you should be good
@Shuttergrenade thanks
Any Blackwell, so RTX 4000 series, DGX Spark (GB10), RTX Pro 6000, etc. It is still a 4bit quantization (it's just nVidia's optimized version of it that has hardware support) so the results won't look as good as fp8. Maybe good enough for what you're after. Certainly good for prototyping or if you plan to upscale. I used nvfp4 and then switched to fp8 after I was happy with the faster previews.
I use the nvfp4 text encoder on a 3060
@tsolful it'll use software emulation and be much slower. You might find fp8 faster.
@HackAfterDark RTX 4070 isn't Blackwell, the RTX Pro 4000 is, but the non-pro consumer GPUs that are Blackwell are in the RTX 50x0 series. The RTX 4070 and other RTX 40x0 boards are Ada Lovelace.
There's support on AD though, I've run it myself but switched to FP8 for better quality. Especially where finger animation is concerned.
@HackAfterDark Yeah i was thinking of using the int8 version but the size is quite large for my 3060 12gb vram 32gb ram rig, might look for a int4 version
It's optimized to run on the latest nvidia cards, but can be emulated on other gpus as well. i even managed to use that on amd!
hey there.
i'm a bit of a noob on video gen.
which version would u rec for a 3090 and 96gb ram?
INT8 for sure
Obviously you can run up to wan 2.2 easily, with the new dynamic Vram mode, you can run any wonder with those 96 of ram.
@tsolful hello why do I get the errror " Load Diffusion Model Failed"
@tsolful thanks for answering.
appreciate it.
@chamo9009 thanks for answering.
appreciate it.
@amirmorad6468 No problem. If you want to generate random things, I recommend Ltx, if you want to generate sex I recommend Wan 2.2 with high quality Loras, or flat the nsfw version of ltx which would be Eros or sulphur 2. You can also run miniMax, as far as I can see, it's one of the finest quality models.
@loneillustrator Have you updated your comfy, usually get the error when im behind in updates and comfy doesn't recognise the new model
@tsolful I did update it
@tsolful do you have a different loader for your diffusion model? maybe you can passs me a workflow?
@loneillustrator Send me the full error log, I use the default template workflow and loader with the subgraph unpacked, Heres a custom node that handles int8 loading (Obsolete with the native implementation) https://github.com/BobJohnson24/ComfyUI-INT8-Fast. And if you download this latest generation i uploaded, drag and drop the file into comfy, should have the workflow used embeded https://civitai.red/images/138864870
Very nice work, those are the only minimax h3 models i found so far offering a good quantization while keeping a reasonable quality.
Thanks!
Thank you 💚 , Kijai has worked some magic with a native implementation of w4a8 and different quantization method here https://huggingface.co/Kijai/MiniMax-H3-experimental Im going to run some tests in a new comfy install but the example looks amazing for the model size. ill probably upload here when it gets merged to comfy as the size is 12gb with near int8 quality
@tsolful Very interesting information, thanks. i'm trying that right now
I'm using a 4070 super (12 GB VRAM) with 32 GB of ram, which variant of this model do you recommend? Are any of them compatible?
I don't think you'll be able to run it. It is very likely that you will get an OOM error, because, even if you have 32 of ram, The model uses about 35 gbs of Vram, if you don't get an error (thanks to the Dynamic ram mode), your latency time will be eternal, perhaps, depending on your configurations, up to 10 minutes per video, although I haven't seen if there is a distilled version, the distilled version would suit you.
I'm running fine out of the gate the int8 convort variants with 5080 16 gb and 32 gb ram. Without any finetuning 0.8 mpx 9:16 aspect ratio (672x1216) 5s is 330s and 10s around 12 min. No oom or anything. I don't mind it. Every generation is just soooo good. Tried a bit different schedulers and that straight up drop minute from 10s clips without noticing any huge quality drops.
Just starting to check faster ways to use it sageattention etc.. You should be fine and I can say Chamo9009 is just wrong with his comment.
@GrimGr1m @VicePalette Yea your system should be able to handle the int8 as comfy has great model management of offloading and streaming from ssd if/when necessary
@chamo9009 i'm running int8 just fine on 12gb vram and 16gb RAM. 5s videos take ~10min at 720p
set up Swapfile (or Pagefile on windows) and run it on a Nvme. 50Gbs of swapfile should be enough. if you don't know how to set it up, you can ask AI
@BekkaT8 It sounds interesting, and I'm so glad you were able to run it. I was able to run locally and in Google colab Boogu Image edit in int8. I was very impressed that it could be done with 16 of Vram and 10 of Ram. Maybe I can run this checkpoint, could you please pass me your workflow?
I am using an NVIDIA RTX A4000 Ada 12GB paired with 32GB of RAM.
I successfully managed to run INT4BQ—processing a 5-second 720p video takes about 25 minutes.
However, whenever I try running an 8 to 10-second video, I run out of VRAM (VRAM explosion).
@chamo9009 ofc, i only use the default workflow with a little tinkering. i can't post my wrkflw on here, i can send it through chat if you want. it's very generic, nothing fancy
i use sage attention, you can disable it if you don't have that set up. it'll take a few minutes longer to generate (about x1.7 faster with sageattn2 after testing. but will trade off a bit in quality)
@BekkaT8 You could publish it in an article, I have the chat disabled, it doesn't work well for me even if I activate it .
@chamo9009 you might need to wait. it's been an hour and still pending at SFW classification. also because i'm so bum idek how to upload a workflow hahah
edit: here it is
https://civitai.red/models/2842347/a-very-generic-workflow-low-vram-12gb-low-ram-16gb?modelVersionId=3208751
should the 4Q still use qwen3vl_32b_minimax_h3_nvfp4_awq.safetensors text encoder?
Yes the vae's and text encoder are the same across the minimax h3 model,
Why is fb32 500M and fb16 is 5G for recommended VAE? Is the fb32 one wrong?
fp32 is the audio vae, the fp16 is the video vae
Model is great. Update comfy and other stuff. Run latest stuff. Everything just works.
I'm running fine out of the gate the int8 convort variants with 5080 16 gb and 32 gb ram. Without any finetuning 0.8 mpx 9:16 aspect ratio (672x1216) 5s is 330s and 10s around 12 min. No oom or anything. I don't mind it. Every generation is just soooo good. Tried a bit different schedulers and samplers that straight up drop minute from 10s clips without noticing any huge quality drops.
Now to do other fine tuning.
Sageattention is another speedup by 25-50%, There are so many caching methods out there idk which one is best as of yet, but multiple people have recommended this node https://github.com/lihaoyun6/ComfyUI-MiniMaxH3-Cache
Can you clarify?
Have you run 17-19Gb models on your 16GB video card?
Minimax works like LTX2.3, no matter how much the model weighs, it will expand in memory?
@Kitsune_Yukino I've run the int8 model and the nvfp4 text encoder on a 12 GB 3060 + 32 GB RAM. I had to edit my pagefile so it can stream some of the weights from my NVMe SSD drive, with the total committed RAM showing 46 GB. And it's because of Comfy's model management that you can do this with any large model, as long as you have enough VRAM + RAM + SSD space. The speed of your SSD matters as well for how fast the weights are read. Hope that clarifies
@tsolful I have 5060ti 16gb and 64gb RAM. So it should be enough without a swap file. Thank you, you answered my question. So it intelligently allocates memory like LTX2.3, and that's very good.
@Kitsune_Yukino Yes, it should be enough for sure 💚
For the LIFE of me how are you guys downloading these big files, every time mine finishes, the download either repeats or says file cannot be found on site. Been going through this for days now :(
I believe it has to do with your security settings or virus protector. The download repeating thing happens to me all the time as well with large files, but after the second time it finishes. It's very annoying and I need to figure out what's causing it, but most likely the virus protector because it's thinking the .safetensor files are bad.
Google Chrome didn't work for me, I tried using Microsoft Edge "save as" and it worked.
I generally try a few things:
- opening up the "Downloads" window (Crtl+J i think) will sometimes increase the stability of the download.
- if that fails, as other uses specified, other browers may work
- I'll use a CLI utility to download the file (like curl or httpie or something)
- I've never done this, but if it kept failing, next i'd try downloading via the civit or hugging face cli tooling specifically
Thought I was going crazy, glad I'm not the only one
Honestly that's what I had to do, unfortunately I just couldn't download the model here. I had to do a roundabout way by using comfyui and through the hugging face node, download someone else's ver of minimax H3, it's super frustrating. I'll take sage attention installation over this "hmmm 15 gb done...actually nah we gonna try again" ass download, was about to just stick with wan 2.1 all my life and call it a day lol. @makiaeveli
Yea man, i was furious for 2 days wondering wth was going on. @DrainBamage
@madaraxuchiha88 Funny enough i did check my virus defender and disabled it, and the download gave up at 1 mb. I'm not joking. And I have good internet. someone does not want me having this model.
I appreciate the help guys
@TheRamPricesAreTooDamHigh I would not be surprised if this particular model is getting downloaded at a crazy rate. It's huge. Internet is only so big. Just try tomorrow. :)
@TheRamPricesAreTooDamHigh
i've had the same issue. downloading anything over 10GB always fails.
i've been using the CMD with the curl command to download them directly.
1. first you go in civitai settings, scroll down to API keys->add api key->save. now store this key somewhere safe (don't share it).
2. now go to the folder where you want to download the file and open CMD there.
3. then type curl -L -H "Authorization: Bearer (your API key here)" "(here you will paste the download link, get it by right clicking the download button and copy address)" --output (file name here, can be whatever you want, that will be the name with which the file will be saved).
the final command will look something like this :
curl -L -H "Authorization: Bearer fkldfr3fkfndsklnf" "https://civitai.red/api/download/models/3193337?fileId=3074134" --output minimax_h3_fl2va_pruned_int8_convrot.safetensors.
hope this helps. as far as i know the api key thing is supposed to tell the servers that it is you who is downloading so it doesn't fail halfway. you can try without it, but this is how i've been doing it for a while.
@Hycros Wow the fact that isn't talked about more is a crime, thank you!
Try disabling "Safe Browsing" in Chrome Settings : chrome://settings/security
I found that Chrome was scanning the downloads, triggering a redownload. Disabling "Safe Browsing" seems to work.
I had this constantly when comfyui was actively running (not just open).
I make sure comfy isn't running when I D/L now and it all works perfectly.
The file was failing verification because the system either ran out of memory or processing power and it just reverts to 'failed'. It always worked on the second pass though.
Anyway, make sure your machine has all it's resources available to it as the download completes and you should be good, if yours is the same problem as mine anyway.
Good luck.
Guys I'm having an issue I've never had before, I'm running the pruned int8 on my 5090 with 96 gigs of ram, for a 480, 5 second video it's taking up almost 70 gigs of ram and 30 gigs of vram. What is going on?
What? I mean... that's correct? your gpu has 32gbs of vram. the model is 20gb, the textencoder is like 20gb, the vaes are big. I even asked google and it says:
"When you sum the absolute minimum weights required to sit in memory just to execute the pipeline, you get ~42 GB of model data. [1]. The absolute math of the weights + the execution overhead demands roughly 70 to 80 GB of total combined memory.
Because the RTX 5090 caps out at 32 GB of physical VRAM, ComfyUI packs the card right up to its limit—using about 30 GB of VRAM to stay stable and avoid an out-of-memory crash. It then offloads the remaining 12 GB of model weights and the massive 20+ GB generation overhead into his system RAM. [1, 2, 3, 4]"
@makiaeveli I get that, 70 gigs of ram plus 30 gigs of Vram is 100 gigs of combined memory for the minimum video quality for me, Also people were talking about how they're running the int8 models on much smaller rigs but I guess it is what it is.
Also are you guys having problems with the I2V face consistency specially on the fl workflows? For some reason the face and body stays much more consistent for me using the ref model
@JoeyDiaz You may also need to update cuda or pytorch or something else. The thing is though, you'll be able to run much higher maximums. Just cause they can make the video, it will take 2-5 times as long as yours. You could probably make 20 second videos. Mine OOMs at like 15-16 seconds after a few. It's also a new model with different computation expectations. It could improve in the future.
@makiaeveli It's so weird I turned off my pc a couple hours ago and I came back to what seems like a completely different model :)), It sits on around 70 gigs of combined memory, I'm using the plaguekind workflow with the spectrum node added and it's giving me pretty consistent results, I am truly impressed with the prompt following and off camera world generation capabilities but the audio generation needs very detailed prompts to come out as intended even tho there is no distortion like we had on ltx,
I'm getting 10 sec 1080p in around 5 minutes and 20 sec 1080p in 16 minutes.
@JoeyDiaz I mean those are pretty crazy output times, so that should make up for your memory "woes" lol
@makiaeveli I actually had a question about this, I'm using the spectrum and sageattention nodes which make the speed difference day and night and I haven't seen much of a difference in quality, but my problem is the workflow keeps randomly bypassing those nodes so something that normally takes 3 minutes, every 5 times that hit run there is a chance that it would take 50 seconds, but I dont know how to keep that speed constant
try disabling pinned memory and see if your ram use drops. Comfy messed up dynamic vram but whether that affects you depends on your system, worth a try
@MrTitsworth Omg wtf bro I cant believe it, Is there a side effect? it got me from 100 gigs of combined ram to 45, You save my life
@JoeyDiaz Yah, I'm honestly not sure other than my assumption is pinned memory is just wasting memory and somehow comfy gets confused. It didn't do this when dynamic vram was officially released by ComfYUI team but they messed something up around the time they released INT8. After they first released dynamic vram, with LTX I finally had my RAM sitting comfortably and surprisingly I was only using 50% or so of my RAM where as before that it would peak my RAM and I'd be writing to my paging file... Then they broke something and it started peaking my RAM again and writing to the paging file. So I tried many things and then tried disabling pinned memory which I'd seen other people do for whatever reason but hadn't tried and boom... back to the same memory use as after official dynamic vram release..
So.. I don't know exactly but yeah it seems that the pinned memory is sitting with nothing of value in it wasting memory and then it loads your actual models on top of that whch causes the stupid ridiculous unnecessary ram use.. I haven't noticed any speed decrease or downside from this so yeah... happy to help
@MrTitsworth Yup I can say with confidence that it has not effected the speed or quality in any way but it just unlocked the bf16 model for me. Reading your experiencing was so relatable I ran into the same problems in the exact same timeframes but after days and days of troubleshooting comfy I couldn't bother to look for a solution
what do you suggest for a 4060 8gb vram? and please dont say "upgrading your gpu"
if you have enough RAM, you can run any of them. But your GPU should be new enough to run the int8 convrot models with the speed improvements. But if you don't have like 64gb of RAM, its probably not going to work right away. You'll be able to get it to work with the right memory configuration though, at least you should be able to.
What @makiaeveli said is true. I was able to run the full model of flux on a 2080 8gb with 64GB of RAM when the requirements said 4090 24GB. Your CPU needs to be from 2020 or higher and have 16 cores and 32 threads or the generation time will be horrible.
@makiaeveli It is not necessary to have a very modern gpu, in Google Colab they offer a free T4 for 5 hours and look that you can run Boogu Image edit Int8 with its 9gb encoder. A beauty if you ask me.
With an 8GB graphics card and 64GB of RAM, the maximum capability—to my knowledge—is generating 12 seconds of 480p video; if you are using a 30-series card, selecting FP8 ensures that this 12-second 480p video can be successfully generated.
Thanks for replying guys. so to my understanding, and correct me if im wrong, i should use the FL2VA int8 pruned version of this for my 4060? i have 32 gb of ram
@Aiddicted 32GB isn't enough with a 8GB card. The model uses around 70GB of total RAM. You probably could but the generation time would take at least 2 hours for a 3 second clip
@brand175 thanks for the insight. i think i'll pass on this then
У меня 3050 8гб, 32гб RAM. Вполне норм запускаются квантованные версии. 0.4-0.5 мп - в районе 13 минут - портретное видео на 10 сек. Плюс турбо-лора на 6 шагов. Ну и Comfy и сами модели конечно на ssd.
@Aiddicted Please don't give up; I spent the afternoon experimenting and found that 8GB of VRAM is sufficient to generate a 15-second 480p video in just 10 minutes. You should use the 21GB int8 quantized model. You need to upgrade your environment to the latest version—PyTorch 2.10.0 + CU130. If you have any questions, consult Gemini; it guided me through the whole process. Before I upgraded my environment, I could only generate a maximum of 12 seconds of video and was forced to use the FP8 model; I later discovered the issue was the environment, and the int8 model is faster.
@Silicon_Mirage How much RAM do you have?
@Silicon_Mirage thanks for the insight! I’ll look into it!
@ToughActive3278369 @Aiddicted My computer has 64GB of RAM, but the new PyTorch 2.10.0 + CU130 version doesn't rely heavily on system memory; it primarily uses dynamic VRAM allocation to generate the video. I was mistaken earlier when I mentioned PyTorch 2.9 + CU130; the correct version combination is actually PyTorch 2.10.0 + CU130. My apologies.
@Silicon_Mirage Thank you for the response!
Has anyone made a good prompt for shooting cum? everything i tried it keeps resisting it
Hi everyone! Could someone please recommend some good Workflows? I would truly appreciate it from the bottom of my heart! Many thanks in advance! 🤗✨
Take a look at the ones embedded in my videos.
DaSiWa MiniMax H3 Workflows | T2VA | FL2VA | REF2VA
tried dasiwa but all videos are ultra slow motion
@eddmoe that just means the frames and the interpolation are off. If the workflow has interpolation settings, it's not matching the outputted frames
there are literally default workflows on their github. And they're very simple, probably the simplest among video models, no need to look here on civitai where people tend to overcomplicate and do their own stuff that is just hard to follow
@Aratoum there are even simpler custom workflows better than the default workflow, you are also missing built in upscalers, nodes that give massive speed boosts and other goodies!
I just used the generic workflow. Download my video below it has a workflow. This is the best thing in the world; I'm going to have so much fun. Shit, I may even start a YouTube channel lol
hi guys which ones of these models are the best for 16gb vram
pruned/curved Q6 gguf
int8 pruned runs well on my 3060 12gb, with the inference speed benefits over gguf @gambikules858
Hi, whats this appears with a I try to download every file? {"error":"Unauthorized","message":"The creator of this asset has disabled downloads on this file"} 🥲
Same here but you can get them from hugging face: https://huggingface.co/Comfy-Org/MiniMax-H3/tree/main
@Eliz99 @fog777 I changed nothing on my end, could be the download all components bugging, if it still fails ☝️
@fog777 Yes! Thanks so much! 🤩✨
In comfy, it doesnt appear you can use loras or add custom nodes to the default "Image to video" template, just to the more complex reference to video template, since it only has a single node for the model and all other settings. Any way to use turbo LORAs with the fl2va workflow?
Right click the Image to Video (MiniMax H3) subgraph, and click Unpack Subgraph, this is what i do to all template workflows with subgraphs, and edit them as i need Load diffusion model > Load Lora > Basic Guider/Scheduler
Just uploaded a simple workflow ive been using to generate turbo outputs, Found in components section
I think i just literally added image processing to the ref2vid workflow lol
This model's outputs are significantly better than older ones but is so unbelievably slow and uses so much memory that I don't think is viable for regular use.
it does eat up a lot of memory but i will say it is by far superior to the other models out there right now anyway. Most people just go to the lowest quality setting generate a bunch of videos, if you have sage attention and the turbo lora so you can go to 8 steps instead of 20, then its even faster. find a seed you like and then generate that one at higher quality with a slightly different result at the end because a difference in quality change.
For this model, type of cpu-ram is far more important than gpu..a ddr5 is a must for speedy generations.