Lipsync with Minimax H3 Ref2VA.
There is no transcript or lyrics needed in the prompt. The model is smart enough to lipsync, as long as you prompt it to.
This is the default ComfyUI workflow for Ref2VA, with:
Kijai's LightX2V LoRA and Kijai's sage attention patch enabled
Math to calculate the length of the audio track to match the number of frames
Due to request from community, I made a version with 2 speakers for a podcast. This is mainly a matter of prompting for the speakers to take turns and to have their lips closed when there is no speech in the audio.
What generally works is to give the official prompt guide (https://github.com/MiniMax-AI/MiniMax-H3/blob/main/.agents/skills/h3-prompt-writing/references/ref-en.txt) to a large language model and ask it to generate a prompt to do what you want.
Description
Removed MelBandRoFormer as podcast do not require audio stemming.
To prevent the lips from moving when there is silence in the audio track is a matter of explicitly stating so in the prompt.
FAQ
Comments (11)
Very nice thank you for sharing! ❤️
I can't make the character STOP speaking even when the AUDIO length is finished and the video continues, or when there are 1-2 seconds gaps in the same audio (not for SONG but for SPEAKING).
I want it to be accurate after with the lipsync also when there are SILENCE gaps.
What happens is that the character ALWAYS talking, can you update the workflow to do that, it will be amazing if you can make it work!
Also IDEA:
Can you make another version with 2 characters speaking (I failed with it because I'm a noob).
see the new version, called podcast. to have 2 speakers take turns without any lip movement when there is no audio is a matter of prompting for it. hope this version helps.
@PixelMuseAI Thanks for the update I'll give it a try later on ❤️
I was wondering why lip-syncing wasn't working with Minimax 3. This workflow comes at exactly the right time and works great. Thanks!
Thanks for your workflow with MelBrand. It's pretty fast, but as the audio file progresses, the lip-sync isn't always accurate. I've tried several Lora Turbo, Steps, Guide, and samplers without success. Best of luck!
how long is your audio track? so far i've only tested up to 15 sec and the lipsync seems to be accurate.
I am not using the audio output by the model but replacing it with the original audio input track, so that is where the desync might happen. if you can provide me with your prompt, starting image and audio track, i can try to replicate the desync you are seeing.
There is a custom node for enable and lock lip-sync, even the scene is changing the target still lip-sync perfectly.
Hey guys im using audio in R2VA workflow and wondering why the output will sometimes say what I ask it to say twice? Usually right off the bat, at the start of the video and then sometimes but not always at the time I specify
you are asking this question in general or specific to this workflow?
if you can provide me with your prompt, ref image and audio, i can try to replicate your issue.
@PixelMuseAI I was trying it on a different workflow! General question
@tyrannnyyy it's probably due to the way you are prompting. Please follow the prompt guide here: https://github.com/MiniMax-AI/MiniMax-H3/blob/main/.agents/skills/h3-prompt-writing/references/ref-en.txt