Minimax H3 INT8/INT4 Convrot
FL2VA - first last (frame) to video / audio
REF2VA - ref_images / ref_videos / ref_video_audios / ref_audios: up to 9 reference images, 3 reference videos (each may carry its own paired soundtrack), and 3 standalone reference audio clips
Both models can generate t2v (Text to video), i2v (Image to video), v2v (Video to Video), a2v (Audio to video), and multiple references (image/video/audio). But were further fine-tuned/trained for higher-quality outputs for the intended use.
14.81 GB INT4BQ Balanced-Quality leaning: int8_ratio 0.41 by count, 0.47 by params; int8mm_ratio 0.58
17.27 GB INT4Q Quality: int8_ratio 0.73 by count, 0.75 by params; int8mm_ratio 0.27
Description
FAQ
Comments (22)
158 sec to make a: 0.5 megapixel, 5 seconds video with a 5800.
cool.
rtx5080?
@Novastar27 Ah, I got fat fingers. RTX 5080, correct.
and 1040 seconds 0.3 megapixel 6 seconds on a RTX3060 12gb
@gambikules858 300s for 0.3mpx 6s with a RTX4070super 12gb
~180sec for a 10sec video at 0.4mpx, 20 steps. RTX 5070 TI 16GB
@BonerSoup More Vram is better
8 minutes for 1mpx video of 5 seconds on my RTX 4070, classic one, 12GB vram. (I have 128GB of RAM). 20 steps, with sage attention.
im Gonna have fun on my RTX 6000. this is very good quality
dude ivde been using my 6k all fuckign day with this. I know it's illegal to use this model in the USA, but fuck that noise; this is awesome!
6 minutes for 10sec at 12fps 320x600p on laptop rtx-A2000- 8GB vram , but loaded the 64GB ddr5 ram to 58GB/64GB! The results are definitely worth waiting for, compared to both LTX and WAN.
Yes, in my opinion, this is the most powerful local video model available right now. I'm incredibly impressed by the animation quality, the anatomy, and the facial stability. This model is literally better in almost every way compared to the LTX and the WAN. Perhaps the LTX is slightly better in terms of realism. But oh my god, the LTX's facial and finger stability are so poor; it takes a lot of effort to get a halfway decent result on the LTX.
@Voxe1 Yah, and it simply does what it is told!..if you craft your prompt well, it sticks to it.
Anyone here with a 5090? I have one and 64gb ddr5 with gen 5 crazy fast ssd, I dont think im getting the crazy fast speed si should be getting.
0.3 megapixel at 10s generates in about 180 seconds with 20 steps res_multistep.
Are any of yall getting that too?
My GPU turned into an airplane engine
I've never used nvfp4. Am I right that it's only for the RTX 5000 series? Or am I confusing it with something else? I'm talking about the text encoder.
Can I use it on my RTX 4070?
I've run it on my RTX 4070 super and it works so you should be good
@Shuttergrenade thanks
Any Blackwell, so RTX 4000 series, DGX Spark (GB10), RTX Pro 6000, etc. It is still a 4bit quantization (it's just nVidia's optimized version of it that has hardware support) so the results won't look as good as fp8. Maybe good enough for what you're after. Certainly good for prototyping or if you plan to upscale. I used nvfp4 and then switched to fp8 after I was happy with the faster previews.