This workflow has been replaced with my new MINIMAX SEED HUNTER WORKFLOW: https://civarchive.com/models/2881362/minimax-seed-hunter-workflow-optimized-fast-latent-upscaler-speedups
Please use that instead.
Watch the tutorial video on how to build an anime-style short using this workflow:
A no-nonsense high level T2V / I2V / FFLF / REF2V workflow for Minimax H3 with lots of options & togglable quality of life features.
Toggle between the [T2V / I2V / FFLF] model & the [REF2VA] model easily
Added Forced Custom Audio: Lipsyncing to custom dialogue/music is now easy for T2V/I2V/REF2VA!
Temporal Upsampler 2nd Pass Option! [https://github.com/matlowai/ComfyUI-MAINodes]
New Image Loader allows cropping/setting max megapixels directly inside node: https://github.com/obvpm/comfyui-obvpm (Must install via unzipping to custom_nodes folder or through comfy-manager's Install via Git option, it's too new to be on the comfy-manager's index)
Kijai's Preview Override (put https://huggingface.co/Kijai/MiniMax-H3-TAE/blob/main/vae_approx/taeh3.safetensors in /vae_approx/ and set as the custom vae for non-pixelated previews)
Speedup: Sage-Attn + New Sol-Attn [https://github.com/kijai/ComfyUI-SolAttn_triton]
Speedup node: EasyCache / Spectrum [https://github.com/xmarre/ComfyUI-Spectrum-MiniMax-H3]
Turbo LoRA (t2v/i2v but works with ref2va if you don't mind [some] audio degradation)
New (8/11/26) Lightx2v 1.0 release [https://huggingface.co/lightx2v/Minimax-h3-Turbo/tree/main]
Film VFI Frame Interpolation (24 fps -> 48 fps)
Easily togglable reference fields: 4 pictures, 2 audio, 1 video
Use this workflow if you:
Are an AI Filmmaker who values character/scene consistency
Want to be able to use FFLF effectively, seamlessly extend videos, or use character reference sheets.
Want to edit videos, or copy & use motion from reference videos
Want fine-tuned control over your shots visuals and audio.
If you wish to generate simple one-off video clips like Will Smith eating spaghetti, the t2v/i2v model that you can toggle to in this workflow will do that for you.
Kijai's int8 convrot video vae: https://huggingface.co/Kijai/MiniMax-H3-experimental/blob/main/minimax_h3_video_vae_int8_convrot.safetensors
UPDATE YOUR COMFY CUDA VERSION TO 13.0. If you start comfy and see cu130:
[INFO] pytorch version: 2.13.0+cu130Then you are good to go! But if you are using cu126, then ALL your gens with the best version of Minimax's model (INT8 convrot) will be 2x slower than they should be due to inefficient comfy-kitchen operations! Keep in mind if you update your cuda, you will need to reinstall Sage Attention! Download the correct wheel from here https://wildminder.github.io/AI-windows-whl/
ALL MODEL / NODE LINKS ARE IN THE WORKFLOW NOTES OR EASILY INSTALLABLE THROUGH COMFY-UI MANAGER.
If you appreciate what I'm doing, please consider following/subscribing on my patreon (free) which gets access to all my work early.
https://www.patreon.com/cw/foxfuressence
Also my youtube where I make AI filmmaking tutorials plz & thx:
Description
fix Sol Attn not being plugged in
fix so Forced Custom Audio option can be used with both fl2va and ref2va models.
Adds T2V / I2V (so now you can swap easily between t2v, i2v, FFLF, and ref2va), new speed up node options, cleans up workflow
FAQ
Comments (110)
Probably one of the best creators right now. FoxFur is real gold for this community!
High praise and hugely appreciated!
Thank you for great workflow! and Do you write your own prompt? or let LLM help you?
I'm extremely old fashioned and handwrite EVERYTHING. The only AI I use is local open source stable diffusion models. Handwriting prompts forces you to learn and adapt!
You can put the "get video components" default node between your load video node and the ref inputs. Probably fix your audio error.
I'll try it. But so far, anything that 'draws' the audio output from the VHS load video node when the loaded video has no audio, results in an error from the node itself. I'd use a different video loading node but that one is just so damn useful, with its ability to select start frames/frame cap/set FPS/etc.
Thank you. For example, let's say I've attached a video and an image. If I want the character in the video to be replaced with the character in the reference image, how should the prompt be configured?
From the official documentation:
"video editingAn existing source video is directly modified"
So something on the lines of: "[video editing] Replace the girl in <Video 1> with the blonde haired girl from <Picture 1>" then go on to describe the actions in the video, reaffirming through description that the girl from the video looks like your reference. Sometimes if reference looks too similar to video, the model doesn't bother trying since it thinks 'close enough'
You forgot to note that Sol-Attn and Sage Attention really degrade the quality
Depends on what you're going for. Action scenes? Sure. Talking head? Not so much. That's why they're included in my wf as toggable options! And why I list this as an 'advanced' workflow; I expect people to know the trade-offs!
Sol crushes, but I haven't seen any definite issue with sage.
at first i was like
MAN THIS VIDEO MODEL SUCKS ASS 11 MINS FOR 5 SECONDS OF .5 MP GARBAGE WHEN I GOT WAY BETTER SPEED AND RESOLUTION IN LTX 2.3??
then i took the leap installed a separate version of comfyui portable with sage attention and nothing else but what is needed to run this wf
and now im making .8mp 5 second vids in 2 mins
cheers to you friend!
Glad to hear it! Keep it up
What GPU are you running to achieve those generation times? I'm curious because I just generated a 10-second clip at 0.4MP and it took around 5 minutes. Before that, I ran a 5-second clip at 0.5MP and it took about 3 minutes.
Do you think doing a clean install actually makes a noticeable difference in generation speed, or is it mostly dependent on the GPU, drivers, and workflow optimization?
@Wallyhojo i have a 4080 ti and doing 5 secs at .8MP I2V averages around 2:22
takes between 190-220 seconds including interpolating to 60fps
the reference to video takes quite a bit longer though
the clean install helped me for the reason that my current comfyui portable is the very first one and has ship of Theseus itself into what it is today lmao, i tried to install sageattn into it before but it just broke everything so i didnt even attempt it
Edit: sorry I just realized I was completely wrong. I wasn't doing 5 second .8mp videos in 2:22
That was 10 second long videos at .8mp and 190 - 220 seconds total after upscaling and frame interpolation to 60fps
5 second long .8 mp videos take 1:10 to produce and a total of 115 seconds for upscale and frame interpolation to 60fps
Hi, here’s my feedback. I always speak my mind—if it’s not good, it’s not good—but when it’s okay, well, that’s true for this workflow. Not only is it okay, but the animation is also super fast—in my case, barely 6 minutes—so thanks for sharing, “foxydits.” There you go, you know the drill. Looking forward to another test ;)
Lol interesting review but I'll take it. I'm dropping a new version today that is pretty exciting!
@foxydits That's awesome! I'm having you test your new version. Thanks again :)
I'm looking forward to testing another workflow as well. I just ran your workflow a few minutes ago, and honestly, I'm blown away by how impressive the results are. The quality is noticeably better than LTX 2.3. While LTX is definitely faster, waiting around 3 minutes for a 5-second clip is completely reasonable considering the level of detail and realism this produces.
@Wallyhojo Couldn't agree more.
Black Magic. Have done a seperate portable installation and the generation time is only 25%.
Thanks for the workflow
Added RTX VSR node just before rife and change the placement of nodes & images
Otherwise, works nicely.
Great workflow again !
Solid workflow man, able to pump out 5 second clips in less than 2 minutes. Yeah Sage/Sol Attention degrade detail a bit, but if I get a clip I really like, I just run it through an upscaler anyway. The render speed is a worthwhile tradeoff IMO.
How do you upscale it? What do you use?
As for version 1.5, I’ve got no complaints: it’s brilliant, even though I’ve disabled three or four options because I couldn’t find any alternatives – but I’ll take the time to look for some later :)
crazy good WF!
Please tell me you will add loras in the future.
Yeah, when I made the workflow there were no LoRAs. I'll update today with a spot to start clicking in LoRAs. For now you can just add your own "Load LoRA" node and connect it somewhere in the model link pipeline.
Great WF, what are people using to generate prompts? Not having a lot of success with Ref2V. I have tried passing VIDEO_PROMPT_WRITING_GUIDE to an LLM along with a general idea of how I want the final output. The LLM generated prompt does not yield any better results. This problem is not with the WF as I also had this problem with the out of box one supplied by Comfy
It really depends on what you're trying to do. I'm building a textpad file full of little prompt nuggets for ref2va that have worked for me.
@foxydits I am trying to replace a character in the reference video with a character I have input under Picture 1 (basically a deepfake). For some reason it always starts off using the scene in picture 1 and then at the end finally transitions to the video but fails to replace the character
@udizjxvo
[video editing] The target video is an edited version of <Video 1>. Replace the girl in <Video 1> with the exact girl in <Picture 1> fully_preserved. is a good starting point. It helps to describe the new deepfaked character too, and the background. Like if I were replacing a girl on the beach with an anime waifu, I would do:
[video editing] The target video is an edited version of <Video 1>. Replace the girl in <Video 1> with the exact 2D animate blonde haired girl in <Picture 1> fully_preserved. Preserve <Video 1>'s background environment in the target video. The girl dances around on the beach playfully, her blonde wavy hair flowing in the breeze as waves crash behind her.
This will motivate the model to maintain the video reference's environment/setting.
@foxydits I used pretty much exactly that and it didn't work. Good to know that it should have though
@udizjxvo If ref character and video character look too alike, I've noticed the model ignores editing it entirely. I think it says "good enough" lol. That's the only time this method hasn't worked for me. As long as your prompt spells out the differences; how your reference character looks different than the one in the video.
sol triton node not working , it needs a second one or
It's not required, it's just a speed-up helper for high resolution/high step count gens. you can disable it. it requires triton to be installed to work. if it's not working, it typically means you don't have triton installed or have an old version.
Spectrum-MiniMax-H3 causes an error during execution and the workflow fails. If I disable it, the generation works normally. Do you know what might be causing this?
If you don't tell me what the error says I can't really help.
On their github it says:
"The node adds no third-party Python dependency. It uses PyTorch and ComfyUI modules already present in a normal ComfyUI installation."
So the only thing that could be an issue unless it's installed wrong is that your pytorch is out of date. Check your run logs for outdated cuda, too.. should be cu130 (cuda 13.0)
Ah, I see now, the most recent comfyUI update broke the node and the author hasn't updated it yet. You can use EasyCache in the meantime, it's built into the workflow.
@foxydits My current setup is ComfyUI 0.30.0-17? 2.9.1+cu130.
It seems very likely that a node is broken. I'll use EasyCache for a while. Thanks!
@ge0078 @foxydits I reinstalled Spetrum to latest release 0.1.8 and Set the bootstrap_first_forecast to false in node. That fixed it for me.
@Sam110177 I also see that Spectrum updated the node 3 hrs ago so I'm sure they addressed the missing variable due to Comfy's update. Should be set now.
What about minimax h3 cache, mem eff sage attention nodes? what's the consensus on that?
The mem efficient node is in my current version of the WF that will be updated today. Minimax h3 cache just drags behind the others. The only reason EasyCache is included is because it's comfy core and doesn't require a special node. Spectrum and T8 Block are better. I an considering switching the option to t8 since Kijai says it handles audio better.
@foxydits hope you'll add rtx upscaler too
@drfaker911219 I can do that, but in my experience testing it, I felt it wasn't that good. Like, I feel it made videos look weird in an uncanny sort of way. But sure, as an option, why not.
the sound is very bad.. I added a nsfw lora after the turbo lora, unsure if related
I think moving updating comfy ui 0.30 or 1.30 messes up sound
@dragonite9000263 <- is correct. It's comfyUI's fuckup.
I think it was the wrong lora for my pruned minimax model. switched to 4step_ckpt850_pruned turbo lora
the 4 step loras are wip , but no one ever reads the docs
Great workflow! I see add audio, but a how do you do voice cloning?
You would use the ref2va workflow, load 10-30 sec of your character's voice into <Audio 1>, make sure the "Use <Audio 1> for T2V/I2V" toggle is OFF (it will force your audio latent to be 1:1 with the audio file).
In your prompt, under the subject_definitions section:
"<Subject 1> is the girl in <Picture 1>
<Audio 1> is the voice-timbre reference for <Subject 1>"
In the summary section where you are writing out your shots, at that point all you would need to do is write:
"<Subject 1> speaks towards the camera: <d>Are you gonna finish that muffin?</d>"
If your sampler/scheduler/step count are all good, and you're not using too many speedups / cache nodes, you should get a fairly convincing cloned voice.
Quick question, I am using ref2va how many steps should I use with the turbo lora turned off?
20-25 with Spectrum or EasyCache is good. It'll seem really slow at first but then around 10 steps in you'll start seeing the cache node skipping unnecessary steps.
Could a person on the god's green earth, for the love of great god tell me how do i even get over this error? I've been searching around everywhere, did a fresh install and still get the following error when i enable sageattention
No module named 'sageattn3'
and get the following error when i disable only sageattention and try to run it up :
sageattention is not new enough version or could not determine CUDA architecture, cannot apply MiniMax H3 Memory Efficient Sage Attention Patch.
note: everything is fresh, 5080 gpu, 32gb ram,r7 9800x3d cpu
By the way, i just generated a 10 sec. video on .5 mp on 16:9, and it took 261. sec to generate, with a previous workflow i remember it used to take like 7 to 8ish min on 480*832 res. with 20 steps
it looks great, but can i do even better if i handle the sageattention thingy ?
Sorry for flooding again,but, it is an insane workflow by the way.
Godspeed man!
Honestly I would just use Easy Installer. It makes installing sage attention really easy.
cluade can
Claude has helped me fix every error i got with comfyui, even those with obscure custom nodes.
Just paste the console lines with your error in it or the whole thing to him like this:
<Help me fix this error i have with comfyui "error">
I have 2.10.0+cu130
Is that alright? Im new
cu130 is great, idk about pythorch. I was on 2.12 and after upgrading to 2.13 I saw no difference in sampling speed
In the current workflow, it looks like Sol Attention might not be plugged into anything.
Thank you for sharing this, your workflow is so good ive tried quite a few now and this is the best one I keep coming back to. I was just wondering do you have any plans of adding an upscaler to this in the future?
There isn't a good upscaler, simply put. RTX upscaler, LTX latent upscaler.. bleh. Just kinda waiting for minimax's 2k upscaler release.
This is the best Minimax workflow so far. My only issue is the lip-sync is really off with custom audio files. I'v been trying to figure out the solution.
If the gen length doesn't match the audio file duration closely, that can happen. It can be like 0.5 sec off but it really should be close. It works really well for me most of the time, with both t2v/i2v, and ref2va.
Update. If your audio is out of sync make sure Rife is OFF when your at 24fps. Use Rife for the 60fps
First of all, this workflow is absolutely astonishing, you did a wonderful job. I do have one question. In the video in the description, the sampler shows a preview of the video while it is being generated (while the sampler is working on generating the video, you can see what he is doing), for some reason this doesn't appear on mine. What was your comfyui instance's version when you recorded the video? I know this "preview" thing is possible because it appeared on one LTX 2.3 workflow I downloaded from civitai. I am using Comfy desktop app.
I don't know how comfy desktop works with this (you might just ask chatGPT), but for me comfy manager has an option to choose preview methods, and I start my comfy with the flag --preview-method latent2rgb
@foxydits Yes, you can change the preview method under "execution" on settings and I have tried exactly latent2rgb and other one before, but it only showed a still image. The default option on preview is just called "default" and for some reason it did show the video being generated on the ltx workflow (using the default option). I might check if there is an extra option on the manager later.
does workflow have any lenght limitation?
hi foxy! thank you for the v16 update, would just like to ask if you have tried kijai's vae convrot? if yes , which is better? the official h3 or that , tysm
I get black screens from kijai's vae. I gave up lol. I use the official.
so this workflow is basically moot now considering most of the custom nodes it depends on to be better than other workflows are no longer available. Love that i wasted my time and lost my previous comfy session to give this a go.
No longer available?.. what are you smokin' dude?
@foxydits For example the git link for sol attention is little more than just an article with no downloads or get commands to run, and the node for it doesn't exist in custom node manager or anywhere i can find online. This workflow might be great if you already have everything offline from when they were available, but for people who use cloud gpus and have to build everything from scratch with each session, some of these nodes are no longer available.
@Fealow I don't know anything about the way cloud gpus' handle custom nodes, but Sol-Attn is completely downloadable from github. It's Kijai's newest node, and he's still working on it because this is a cutting edge brand new model that only released like 7 days ago. When you put reviews like this out there saying "the workflow is moot because the custom nodes aren't available", you imply that would be true for everyone.. but it's not. Considering thousands of people use this workflow, it's clear this idea of nodes not being available is only a problem for you.
@Fealow ready made comfyui clouds usually don't have these nodes, you have to install them locally or on clouds with rented GPUs
Love the workflow. Best one I have tried so far. Just a couple of suggestions: 1. Can you add an optional video upscaler - RTX,LTX, etc. 2. When you make a generation you get a first frame, a video with no sound and a video with sound. I dont see an option to control this, for example if you need only the final video. Is it possible to add one. 3. I dont see a Sigma shift node (i apologize if there is one and I dont see it). Changing sigma shift for audio for example currently fixes audio problems with some turbo loras.
Hey since you listed out in numbers I'll reply to each.
1. I wasn't satisfied with RTX or LTX's upscaler, so I didn't add it. I spent a while trying to make the LTX latent upsampler worth the time, but it just wasn't. I went deep, looking at pixels, and found that yes the resolution was suddenly 1080p, but it was no more detailed or higher quality, and the motion/detail actually would go down. I will consider adding it as a subgraph now that my workflow already uses LTX nodes anyway for my custom audio hack.
2. I'm not sure what you're saying here actually. Sorry. Are you saying that video references can't have audio in my current workflow? I have a little note explaining the video-audio ref situation at the bottom right of the workflow.
3. That's true, I should add that.
@foxydits Hey thank you for the swift response. Regarding 2. What i meant is that currently when i do a generation i get 3 files generated in the output folder of comfy: 1 image with the first frame, 1 video with no sound and one finished video with the sound. I dont see if there are any toggles to choose not to generate the image and the video without the sound, so there will be only one file, since i dont need the rest and have to delete them every time.
@josharkness449 Oh, that has nothing to do with my workflow. That's your comfy settings. I had a similar thing at one point a long time ago and the solution was to turn off intermediate files in the settings. I don't know the setting name exactly, but you should be able to find it. Maybe in VHS settings. ask chatGPT if you can't find it.
Hi, I still have red boxes/missing nodes in the workflow.
Everything is installed except one entry called “Unknown pack”.
When I expand Unknown pack, it shows 5 missing nodes, and all 5 are rgthree nodes.
I believe this Unknown pack is what is causing the red boxes. The normal rgthree package is installed, and I have already completely reinstalled it.
I also verified that FastGroupsBypasser is actually present in the installation:
fast_groups_bypasser.js
and the file contains:
rgthree.FastGroupsBypasser
So the problem is not that the rgthree files are missing. The problem is that ComfyUI still treats these 5 rgthree nodes as belonging to an Unknown pack and therefore marks them as missing.
What package/version provides these 5 nodes, and why are they being detected as “Unknown pack”?
I have already reinstalled rgthree-comfy, but the Unknown pack remains and the workflow still has red boxes.
Thanks
GPU: NVIDIA GeForce RTX 5060 Ti
PyTorch: 2.13.0+cu130
CUDA: 13.0
OS: Windows
ComfyUI: Windows Portable
Hmm. Very strange. It is indeed all just rgthree-comfy nodes. Nothing magical or outside the normal. I guess you can update comfyUI / python dependencies if you haven't, and when you boot it, look at the logs. Find where it tries to import rgthree-comfy and let me know what it says, there might be something like a "FAILED TO IMPORT" message.
@foxydits Hi, I spent about 2 hours with GPT going through everything. I reinstalled rgthree, checked all the versions (Cu, Python, with and without Sage, etc.), made sure everything was up to date, and also tried different versions of rgthree.
I think the issue might be those boxes where you select the options, which appear completely empty. I believe they might be called Fast Muter (or something similar).
The error message is “Unknown Pack.” Next to it, it shows 5, and when I open it, it just lists “rgthree” five times, with nothing else.
Rgthree is installed and working, which makes the whole thing even more confusing.
@bernt8379944 Indeed, it is confusing. You're the first person in many thousands who have used my workflows to have reported issues with rgthree like this. It is a very popular and well-regarded essential node-pack for many workflows. I simply have no idea what could be going wrong for your installation compared to the others, EXCEPT that typically nodes not loading even though the pack is installed typically means python dependencies are wrong.
@foxydits
Well, nothing new there — it’s always me. 😄
It might actually be worth trying a complete reinstall, keeping my models and reinstalling everything from scratch.
And yes, it’s really appreciated. Very well done!
Did not help to install all from scratch.
first off...WOW! great work!
I am doing 15 seconds vids with 0,7 mp in 11 minutes ...including spectrum.(RTX 5070 Ti 16gbvram)
that is insane!
though I am using a global sage and am bypasing kj sage node...the turbo lora isnt doing anything for me either , in fact its dragging gens down a few seconds.
also kjs experimental vae gives me black vids...so I am sticking to the regular vae.
Yes, it gives me black vids too. Minimax generative audio is still rather bad. It's one of the main reasons why I included that hack (toggled on with the "Force <Audio 1> latent" option on the top left of the workflow) that lets you inject your own custom audio, which should come out clear. Just make sure your audio file is of similar length to your clip for that one though.
One warning about global sage, (--use-sage-attention) is that it breaks some models, like Krea2.
Does anybody know where to get the loadaudioui node? Nothing in the description. Nothing in the notes included in the workflow.
Sorry about that, it's because it should be easily installable through the comfy-UI manager which doesn't require much effort. You'd just go to comfy manager -> missing nodes and it should pull up WhatDreamsCost's node suite. If not, here you go:
https://github.com/WhatDreamsCost/WhatDreamsCost-ComfyUI
thx, couldn't find it via "installing missing custom nodes"
@foxydits also couldnt find it via manager. their Database sucks ass and after 2 years of comfy use i still dont understand half of it
@foxydits If I could have found it that way I wouldn't have asked.
anyone else getting oom errors with v 1.6? 1.5 was working perfectly on my machine, but as soon as I switched to the new one, I can't get anything going due to out of memory errors. :(
Nothing that should affect memory usage was changed. The only things changed are node cleanups, the forced custom audio option now working for ref2va as well, and Sol-Attn being plugged in instead of unlinked on accident. So it's likely being cause by some other factor.
Thank you for the wonderful WF.
There’s probably something wrong with v1.6.
Since I’m using a ref model in v1.5, I’m not using LightningLora; instead, I’m checking the video’s motion at low step and low MP before trying high resolution.
Low-resolution videos in v1.5 run very fast, but under the same conditions in v1.6, low resolution is nearly three times slower.
Nothing that affects speed was changed between 1.5 and 1.6 besides Sol-attn being correctly linked (see below). My guess is you updated comfy or something is different without you noticing it between the time you used 1.5 and the 1.6 update.
The only thing I can think of between the versions is that in 1.5 the Sol-Attn wasn't connected to anything so it never worked, which was fixed in 1.6.
If you are using ref2va, it could be as simple as your input images/ref video being large (2+ MP) and that the minimax ref2va node is set to 'max' instead of 'match' which uses the full resolution references. That can lead to slower results. But the reality is, my workflow didn't change so much as everything else did; comfy updates, the nodes being updated, settings you're maybe not aware of.
This is nice. Thank you.
What do you think about incorporating some ideas from here?
very nice! installing some nodes, and it worked perfectly out of the box. tested around 5 videos with 4070 super 12vram on 0.5 megapixel / 8 sec , finish around 3 to 4 mins
examples are freaking lovely and quality rocks. I'm gonna test !!!thanks
How can I add the updated MiniMax H3 Preview Override to this workflow?
I am adding it to the v1.8 version releasing in a few minutes.
@foxydits Which one is the new image loader node because it says Unknown: LoadImageCrop
@josharkness449 There's a note inside the workflow that gives links to all the nodes. It's bright red on the left side.
