New ref distilled lora
https://huggingface.co/Kijai/MiniMax-H3-experimental
I debated releasing this as its not much more than the basic comfy workflow
Why use minimax
- Great prompt listening
- Combine multiple subjects easily.
- Quality lip sync and voice clone
- 5-15 second video clips 300 frames at 1MP.
- You want to create a story.
This uses A LOT of resources, crashed a few times, soft locked too when i went too big, but 1MP is not a tiny video when i ran about the same for LTX and much smaller for vace. Its slow, but its doing more than most.
I think it can extend like vace could, and without a contrast shift. Its down to learning the prompting to extend from the last 8 frames of a video.
Minimax is a beast for storytellers.
(fyi my gpu is stuck on PCI3 atm due to my cpu)
From clean load of comfy loading models and render of 1MP 311 frames at 8 steps took with my setup 23 min 21 seconds 158s/i
Output is 1376x768 at 12.75 seconds
Description
FAQ
Comments (7)
Thank you for the workflow. On my RTX 4080 Super 16 GB and 128 GB of RAM, generating a 6‑second video (30 fps) without using the Turbo LoRA takes about 4 minutes, and 60 fps takes about 10 minutes. However, if I try to generate at 60 fps for more than 5 seconds, ComfyUI crashes and closes. Still, the results are not bad. I think upscaling can be done separately through Topaz, since Minimax is a very resource‑hungry model. I also noticed that Minimax is sensitive to the prompt and prefers shorter prompts. When writing long prompts like you would for LTX, it tends to steer the image toward realism or semi-realism.
Interesting, does the 60fps videos have proper audio and pacing? I had thought the audio would go out of sync as its trained on 24fps and the switch i used only changed the frame counts. Minimax doesnt let you set the FPS.
LTX can upscale minimax first/last frame videos but ltx needs both first and last frames. Using ref2va to place a subject doesnt have first/last frames for ltx to use as refs unless you use the actual frames. This can work if using 1mp minimax sizes.
@sy0ww4bb1984 For 60 fps, I only used test prompts without any words or phrases, for example 'a girl dancing a rhythmic dance' - there was just random music playing. In the example with your workflow, I simply wrote a phrase in the prompt that the girl was supposed to say ('Hi, how are we lifting, boys?'), that was at 30 fps, but even there is a slight desync. I haven't experimented yet with attached audio or with references from multiple images - I think I'll give it a try. By the way, a new model LTX 2.5 has been released - I wonder if it will support LoRAs from version 2.3; that would be great.
@incubusinmyhead LTX 2.5 works exactly as 2.3 did. Replace the models and the clip loader with a single loader and all loras work from 2.3. Union control works the same with slightly better quality and better mouths.
Ltx 2.5
@sy0ww4bb1984 Thanks for the info! I'll try to test it out in the next few days.
Good day sir! Ltc 2.5 just dropped more faster then 2.3, can we expect a 'Fastest v2v workflow on planet earth' from you. P.s. please check inbox 📥 sir have a great day.
:P tomorrow, To use LTX 2.5 you can change the clip loader to a single loader and load the models the same way into the flow as 2.3. Loras seem to work fine.
Will take a look at your other request. lol.
