V3 Update
Now actually the fastest HQ audio + video generation workflow. See for yourself in the examples. I have intentionally created shots that show distant faces, because that is still a Minimax issue and no workflow issue currently. I can currently create these 8s, 1 megapixel videos in 3.5 minutes which is INSANE if you ask me, while keeping good audio!
I'm still experimenting with different scenarios, tweaking stuff like first pass generation (0.2 - 0.4 MP), aswell as step count for the first pass
WHY 2 MODELS AND REFINEMENT AT ALL??
Yes, valid question. My answer: creating high resolution videos with the Fast H3 model takes much longer than this workflow, and the quality of the short Taomate turbo workflow is outright bad with low steps. The Solution: We use the power of the audio quality and base movement of the Fast H3 model while then using the ultra-rapid speed of the Taomate 3 step turbo lora to refine and enhance the video to High Quality!
PLEASE let me know how this works for you!
What changed: Completely reworked the 2nd pass to use the base T2V model from H3, you can of course try and use other ones. It is important to use the Fast H3 model for the first pass.
Base Information
This is my very first attempt at making a FAST workflow that runs on my 16gb + 32gb System RAM Setup and I'm looking forward to see if you guys have the same results like me.
This workflow uses 2 passes with 2 different models - the Fast H3 model + another normal T2V/I2V model. The first generates a smaller resolution video + the final audio, using the Fast H3 base model. The second pass refines the video into a higher resolution using the Taomate 3-Step Lora. Also Attention nodes are in use as well as Previews for the first and second pass seperately, so you can already decide if you want to keep generating the higher resolution video or not. This combination ensures good movement aswell as the good audio from the first pass while enabling high resolution generations in only 3 refinement steps at 0.35 denoise.
Please try it out and let me know what you think!
INSTRUCTIONS:
Use the green nodes to configure your resolution, aspect ratio & duration
Keep the aspect ratio the same for the first pass and second pass resolution selector
Add your desired LoRAs in the purple node
Adjust to your desired VAE's, text encoders that you might already have (I'm using a new KJ video vae that is smaller)
Don't forget to adjust the filename and path in the Save Video node to your prefered syntax and location
GENERATION SPEED:
On 5070Ti with 32gb System RAM:
8 seconds - 1 MP resolution: ~3:30 minutes
10 seconds - 1 MP resolution: ~4:30 minutes
10 seconds - 0.7 MP resolution: ~3:10 minutes
Description
First Upload to see what you guys think.