V2.1 UPDATE
New version naming. New nodes for better I2V generations (previous one had cases where it would create a bad first frame), Chunking node for VRAM optimization, new node in first pass for better micro-details. New nodes in workflow to install via ComfyUI Manager: https://github.com/yichengup/ComfyUI-YCNodes-MiniMax-H3
***- PLEASE leave some kind of feedback after trying -
Base Information
Now actually the fastest HQ audio + video generation workflow. See for yourself in the examples. I can currently create these 8s, 1 megapixel videos in 3.5 minutes which is INSANE if you ask me, while keeping good audio!
I'm still experimenting with different scenarios, tweaking for the ultimate best balance of quality vs speed. There might be more updates very soon.
WHY 2 MODELS AND REFINEMENT AT ALL??
Yes, valid question. My answer: creating high resolution videos with the Fast H3 model takes much longer than this workflow, and the quality of the short Taomate turbo workflow is outright bad with low steps. The Solution: We use the power of the audio quality and base movement of the Fast H3 model while then using the ultra-rapid speed of the Taomate 3 step turbo lora to refine and enhance the video to High Quality!
PLEASE let me know how this works for you!
This is my very first attempt at making a FAST workflow that runs on my 16gb + 32gb System RAM Setup and I'm looking forward to see if you guys have the same results like me.
Please try it out and let me know what you think!
INSTRUCTIONS:
Use the green nodes to configure your resolution, aspect ratio & duration
Keep the aspect ratio the same for the first pass and second pass resolution selector
Add your desired LoRAs in the purple node
Enable or Bypass your desired 1st Frame or Last Frame input Image Groups on the left side of the workflow
Adjust to your desired VAE's, text encoders that you might already have (I'm using a new KJ video vae that is smaller)
Don't forget to adjust the filename and path in the Save Video node to your prefered syntax and location (the one in the workflow creates this type of filename syntax with MP, date and time of generation: 1MP_Video_210926-150209_24fps_00001_.mp4)
MODELS
Base FLF2VA / T2V model: https://huggingface.co/Comfy-Org/MiniMax-H3/blob/main/diffusion_models/minimax_h3_fl2va_pruned_int8_convrot.safetensors
Text Encoder for 50xx Series Nvidia GPUs: https://huggingface.co/Comfy-Org/MiniMax-H3/blob/main/text_encoders/qwen3vl_32b_minimax_h3_nvfp4_awq.safetensors
Text Encoder for older GPUs: https://huggingface.co/Comfy-Org/MiniMax-H3/blob/main/text_encoders/qwen3vl_32b_minimax_h3_int8_convrot.safetensors
Audio VAE: https://huggingface.co/Comfy-Org/MiniMax-H3/blob/main/vae/minimax_h3_audio_vae_fp32.safetensors
Taomate 3 Step Turbo Lora: https://huggingface.co/drbaph/MiniMax-H3-Turbo-Lora-ComfyUI/blob/main/experimental/minimax_h3_taomate_fl2va_3step_ema_comfyui.safetensors
H3 Latent Upscaler: https://github.com/LBH-123-AI/Comfyui_Minimax_h3_latent_Upscaler
Custom Nodes Links
GENERATION SPEED:
On 5070Ti with 32gb System RAM:
8 seconds - 1 MP resolution: ~3:30 minutes
8 seconds - 0.5 MP resolution: ~1:45 minutes
10 seconds - 1 MP resolution: ~4:30 minutes
10 seconds - 0.7 MP resolution: ~3:10 minutes
Description
New node for better seemless I2V generations + added Chunking for VRAM management to enable higher resolution video + better micro-details nodes.