Turn a portrait and any-length audio into a talking video
Open this ready-to-run workflow on Floyo. No install needed.
HOW IT WORKS
Step 1. Upload your portrait. A front-facing headshot of the person you want to animate, well-lit, with the mouth and jaw visible. Works great with: portraits · headshots · AI-generated faces · character renders
Step 2. Upload your audio. Speech, narration, or singing, at any length. The workflow separates the vocals from background noise on its own before it starts. Works great with: voiceovers · podcast clips · dialogue · singing
Step 3. Write a short prompt. Describe the performance and mood, like "a man talking calmly with slight head nods." This steers the motion around the lip-sync. Keep it about the delivery, not the room.
Step 4. Hit run and download. InfiniteTalk generates the video in overlapping windows and tiles them across the whole track, then combines them into one MP4 at 25fps with your audio embedded. Ready for: Premiere · DaVinci Resolve · After Effects · YouTube · TikTok · podcasts
First time? Upload a front-facing portrait and an audio clip, leave every setting as-is, and hit run. The defaults handle the rest.
Overview
This workflow turns a portrait and an audio file into a lip-synced talking video using Wan 2.1 14B with the InfiniteTalk adapter. The point of it is length. Most lip-sync tools cap out at a few seconds, because they generate a fixed number of frames in one pass. InfiniteTalk processes your audio in overlapping 81-frame windows and loops until the full recording is covered, so a 10-second clip and a 5-minute monologue run through the same pipeline with no duration limit. It separates vocals from background music, reads the speech with a Wav2Vec2 encoder, and drives the mouth from the sound, so the sync works across languages. You upload a portrait and audio, write a short prompt, and get the full-length video back. Generation time scales with audio length, and a short clip takes about 2 minutes 22 seconds. No setup, no nodes to wire.
Who it's for: creators, podcasters, and course makers who want a talking avatar for long audio in ComfyUI without wiring the windowed lip-sync pipeline from scratch. Not for: a fast one-off clip or a profile-angle portrait. It is a heavy render, and the mouth and jaw need to be visible and front-facing.
Why Floyo
Floyo is the only ComfyUI platform built for teams in the browser.
Made for teams. Share run history, files, and models across your whole team. A teammate opens your exact run and picks up where you left off. No file handoffs, no version confusion.
No install, no setup. Every workflow and model is preloaded. Open it in your browser and run. Nothing to download, nothing to configure.
No local hardware. Workflows run on H100 NVL GPUs, so heavy models run fast without a card of your own. Your VRAM stops being the limit.
Open and closed models in one place. Floyo runs open-source workflows and API models side by side.
How to use (in your browser on Floyo)
Upload your front-facing portrait and an audio file of any length.
Write a short prompt describing the performance, then run.
Download the talking video with your audio embedded, ready for any editor.
Expectations The models are preloaded, so there is nothing to download. Running it needs a free Floyo account. This is a heavy render, and generation time scales with your audio: a short clip is about 2 minutes 22 seconds, and longer tracks take proportionally longer. A few things decide the result. Front-facing portraits with a clear, unobstructed mouth sync best, while profile angles and hair or hands over the lower face do not. Clean, isolated vocals give tighter sync than audio buried under music or reverb, even though the workflow separates vocals first. To push the mouth harder, raise the audio_scale value. On licensing: InfiniteTalk builds on Alibaba's open-source Wan 2.1, but the adapter, LoRAs, and supporting models each carry their own terms, so check each component before commercial use.
Use Cases
Podcasts & Narration. Turn a full episode or a long narration into a talking avatar video, giving audio-only content a face for YouTube and social.
Course & Explainer Content. Generate a consistent instructor from one portrait across full-length lessons, no filming.
Singing & Music. Feed a vocal track and a character image to produce a singing performance synced for the whole song.
Long-Form Social. Make talking clips longer than other lip-sync tools allow, covering full walkthroughs, story times, and monologues.
What can InfiniteTalk turn into a talking video?


