Make a portrait speak or sing in sync with your audio
Open this ready-to-run workflow on Floyo. No install needed.
HOW IT WORKS
Step 1. Upload your portrait. A clear, front-facing image with the mouth, jaw, and chin visible. This becomes the character in the video. Works great with: headshots · anime characters · AI-generated faces · character renders
Step 2. Upload your audio. Speech, narration, or singing. The workflow separates the vocals from any background noise on its own and trims to 8 seconds.
Step 3. Write your prompt. Describe the character and the performance, like "anime mecha pilot speaking with focused determination, accurate lip sync, natural blinking." Describe the performance, not camera moves.
Step 4. Hit run and download. Three passes generate and upscale the clip while your audio drives the mouth. You get back a 10-second 1080p MP4 at 24fps with the audio embedded. Ready for: Premiere · DaVinci Resolve · After Effects · YouTube · TikTok
First time? Upload a front-facing portrait and an audio clip, edit the prompt to describe your character, and leave everything else as-is.
Overview
This workflow turns a portrait and an audio clip into a lip-synced video using LTX Video 2.3, an open-source audio-visual model from Lightricks. It reads your audio, not just a prompt. A separation model pulls the vocals out of the track, the LTX audio encoder turns them into features, and the model generates mouth movement, timing, and jaw motion that match the speech. A static camera LoRA locks the frame so the face stays centered the whole time. A three-pass pipeline builds the clip at a small base size, then upscales twice to reach 1080p, with detail LoRAs sharpening the face along the way. Because the sync reads sound, not words, it works across languages. You upload a portrait and audio, write a short prompt, and get back a 10-second 1080p clip in about 2 minutes 46 seconds. No setup, no nodes to wire.
Who it's for: creators, animators, and course makers who want audio-driven talking characters in ComfyUI without wiring a multi-pass lip-sync pipeline from scratch. Not for: a fast one-click clip. This is a heavy render, and the source portrait and clean audio decide most of the result.
Why Floyo
Floyo is the only ComfyUI platform built for teams in the browser.
Made for teams. Share run history, files, and models across your whole team. A teammate opens your exact run and picks up where you left off. No file handoffs, no version confusion.
No install, no setup. Every workflow and model is preloaded. Open it in your browser and run. Nothing to download, nothing to configure.
No local hardware. Workflows run on H100 NVL GPUs, so heavy models run fast without a card of your own. Your VRAM stops being the limit.
Open and closed models in one place. Floyo runs open-source workflows and API models side by side.
How to use (in your browser on Floyo)
Upload your front-facing portrait and an audio clip.
Write a short prompt describing the character and performance, then run.
Download the 1080p lip-synced clip, ready for any editor.
Expectations The models are preloaded, so there is nothing to download. Running it needs a free Floyo account. This one is heavier than most: each clip takes about 2 minutes 46 seconds, because it generates and upscales across three passes. A few things decide the outcome. Front-facing portraits with a visible mouth sync best, profile or extreme angles do not. Clean, isolated vocals give tighter sync than audio buried under music or reverb. Keep camera-movement words out of the prompt, since they fight the static camera lock. On licensing: LTX Video 2.3 is open-source from Lightricks, but the LoRAs and supporting models carry their own terms, so check each component's license before commercial use.
Use Cases
Talking-Head Videos. Turn a voiceover and a portrait into a 1080p talking character with lip-sync, natural blinking, and subtle head movement.
Singing & Music. Feed a vocal track and a character image to generate a singing performance with matched mouth movement.
Course & Explainer Content. Generate a consistent instructor from one portrait across lectures and explainers, no filming.
Character Animation. Give illustrated or AI characters dialogue-driven lip-sync for shorts, trailers, and proof-of-concept reels.
What can LTX 2.3 lip sync bring to life?
