The Workflow was setup to have a clean "GUI" showing only parameters that matter, so you might want to toggle off Link visibility.
V1.0 Minimax H3: IMAGE or TEXT or REFERENCE to Video with Ollama
Workflow features:
single Image, First Frame/Last Frame, Text or Reference to VIDEO
can use up to 4 images, 1 video, 1 audio input as Reference to generate videos
uses Ollama with dedicated system prompts to enhance simple user prompts
applies RTX Video Super Resolution to upscale to a final resolution very fast
can toggle the following accelerators: Turbo Lora, Sage Att., EasyCache, Spectrum, FirstBlockCache, Sol. Att.
Downloads:
Models (IT2V and REF2V, int8_convrot recommended): https://huggingface.co/Comfy-Org/MiniMax-H3/tree/main/diffusion_models
Alternative Model (recommended), IT2V, Ref2V, Turbo Lora, NSFW Lora (Mystic) merged into 1 model: https://huggingface.co/MATLOWAI/minimax-h3-fused-turbo-int8-convrot/tree/main/diffusion_models
Textencoder (nvfp4) : https://huggingface.co/Comfy-Org/MiniMax-H3/tree/main/text_encoders
VAEs (audio & video) : https://huggingface.co/Comfy-Org/MiniMax-H3/tree/main/vae
Turbo Lora (resized): https://huggingface.co/Kijai/MiniMax-H3_comfy/tree/main/loras
LightX2V Turbo Loras: https://huggingface.co/lightx2v/Minimax-h3-Turbo/tree/main
Winner of IT2V Arena (minimax_h3_turbo_v4_step600_ema): https://huggingface.co/larryvrh/MiniMax-H3-Turbo-Lora/tree/main
Ollama model (tested w/ 26b): https://ollama.com/huihui_ai/gemma-4-abliterated or any other model ideally with vision capabilities, like (9b) :https://ollama.com/huihui_ai/qwen3.5-abliterated
Save locations:
π ComfyUI/
βββ π models/
β βββ π vae/
β β βββ minimax_h3_video_vae_fp16.safetensors
β β βββ minimax_h3_audio_vae_fp32.safetensors
β βββ π diffusion_models/
β β βββ minimax_h3_fl2va_pruned_int8_convrot.safetensors
β βββ π text_encoders/
β βββ qwen3vl_32b_minimax_h3_nvfp4_awq.safetensorsOllama help:
Install Ollama fromΒ https://ollama.com/
download a model: Go to a model page, chose a model , then hit the copy button, i.e.Β https://ollama.com/huihui_ai/qwen3-vl-abliterated
open terminal and paste the model name, i.e.: ollama run huihui_ai/qwen3-vl-abliterated
model will be downloaded and can be selected in green comfy node "Ollama Connectivity". Hit "Reconnect" to refresh.
Custom Nodes used:
Description
V1.0 of Minimax H3 Workflow for Image,Text or Reference to VIDEO with Ollama
FAQ
Comments (5)
That's an absolute fantastic and helpful workflow. The support from the LLM for the prompt is much more needed with Minimax H3 compared to LTX2.3 - and here we get a nice, flexible solution for all Minimax scenarios. Thanks!
Thank you, glad you like it
best workflow ever ! very fast and ollama works perfectly !
There is a model merge avail, that works with 4+ steps, both IT2V and Ref2V + Turbo+other Loras merged into one model:
https://huggingface.co/MATLOWAI/minimax-h3-fused-turbo-int8-convrot/tree/main/diffusion_models
"Base: the pruned fl2va transformer with a rank-1024 SVD of the (ref2va - fl2va) weight delta fused in, so a single partition serves both first/last-frame and reference conditioning...
Merged in: lightx2v FL2VA Turbo 8-step v1.0 at 1.0 (rank-24 resize from Kijai/MiniMax-H3_comfy) and Mystic v2.0 at 0.7 (Civitai model 2856467, merged for motion smoothness)."
There is a link where user could blind vote on Turbo Lora results (Arena), see link below.
This seems to be the winner:
(minimax_h3_turbo_v4_step600_ema.safetensor): https://huggingface.co/larryvrh/MiniMax-H3-Turbo-Lora/tree/main
more Info:
https://www.reddit.com/r/StableDiffusion/comments/1w64k18/first_results_from_h3_acceleration_arena/