CivArchive
    Text-to-Speech with Voice Clone in LTX 2.3 - v1.1
    Preview 126545198
    Preview 126545204

    Use LTX 2.3 as a Text-to-Speech (TTS) model.

    Note: As it was pointed out in the comments below, there is an outstanding bug in ComfyUI that may prevent you from using the new voice cloning node in this workflow. See the discussion on github for more information.

    This workflow is designed to generate audio only output from a prompt and a speech sample. The generated audio will clone the sample voice and apply it to the specified dialogue and prompt. If you want to make more than one video with a consistent character voice, then you can't rely on LTX's random voice assignment, so this workflow will give you a consistent voice that you can use to create whatever spoken script you want.

    Voice cloning is made possible with the ID-LoRA models and new ComfyUI node, "LTXV Reference Audio". In my experience, the video generated with this LoRA isn't very reliable or high-quality, so I had better results from creating an audio file and applying that pre-made audio track to a new LTX 2.3 video based on a starting image. Those results haven't been 100% perfect either, but the success rate was higher than any other method I tried.

    IMPORTANT: If you haven't updated your ComfyUI installation after about March 25, 2026, you will need to run an update to get the new ComfyUI native node for ID-LoRA.

    Please see the ID-LoRA github page for important guidance on prompt formatting and usage. The default values in this workflow have worked well for me, but all the nodes are clearly exposed and labeled so you can tweak and experiment to get your own favorite results.

    Github: https://github.com/ID-LoRA/ID-LoRA

    One more piece of advice for generating audio tracks with LTX 2.3: The length of the generated audio clip is very important, probably more important than any other setting in the workflow. If the time is too short, LTX will rush through some of the script with almost no pause in between sentences and the result doesn't sound as natural as LTX is capable of doing. If the time is too long, LTX will stretch out pauses and sometimes repeat sections of the script. My recommendation is to say your script out loud to yourself, in a normal conversation pace, and time yourself doing it. Use that time as the duration for your clip and then adjust it longer or shorter as needed.

    Finally, don't be afraid to generate multiple audio clips if you have a long script with several breaks or pauses and LTX can't seem to get the pauses right. It's much easier to combine audio files into one track than it is to combine video and there are lots of online tools to help with that. When you assemble your own final audio track, you can insert pauses as long or short as you want.

    Description

    Bugfix version: correcting a mistake in the workflow where the cloned voice model was not connected to the following nodes as it should have been.

    FAQ

    Comments (15)

    katdarnell98993Apr 6, 2026
    CivitAI

    I've followed all the instructions and the final output voices are not cloned from the reference audio. I'm not sure what else I need to do or tweak. Any suggestions?

    darkroast175696
    Author
    Apr 6, 2026

    Are you using the 1.1 version that I posted last night? If you are, and you still don't get the cloned voice, you may be experiencing the bug in comfyui's new cloning node that is discussed in the github issue thread I mentioned in the model description below. At least one other person in that thread said he was able to get the node to show up in his workflow, but that the cloning didn't work. If that's the case with you, all I can say is to wait for a new update of comfyui and see if that fixes it.

    katdarnell98993Apr 8, 2026

    @darkroast175696 I have tried the fix on github and the node is now showing. Still no cloning though :(

    darkroast175696
    Author
    Apr 8, 2026

    @katdarnell98993 and you're using version 1.1 of the workflow, right? Because version 1.0 had a bug.

    katdarnell98993Apr 9, 2026

    @darkroast175696 Yes, 1.1. It generates and processes everything all the way to completion, but there is no similarities to the input voice unfortunately.

    darkroast175696
    Author
    Apr 9, 2026

    @katdarnell98993 I double-checked my version 1.1 to make sure I uploaded the correct workflow file and it's correct. There is another comment thread on this page started by "InsidiousOne" where someone posted a link to a discussion about the bug and someone had fixed their installation by replacing more files than just the one python script. You could read the comment there and see if that process helps in your case.

    darkroast175696
    Author
    Apr 15, 2026

    The github thread I mentioned earlier has a new message indicating the bug may now be fixed, in case anyone wants to give it a try. You'll need to update your comfyui to the newest version to test it out.
    https://github.com/Comfy-Org/ComfyUI/issues/13194#issuecomment-4249039663

    kamil_kamillia578May 22, 2026
    CivitAI

    The output sounds like a low bitrate mp3 despite being flac, and there are artifacts every now and then. Is this just the best LTX can do, or is it a matter of tweaking the settings?

    darkroast175696
    Author
    Jun 9, 2026· 2 reactions

    There are two different data models that were released with this feature (https://github.com/ID-LoRA/ID-LoRA#-dataset-preparation) and they may give you different results. Yes, it could also be a matter of settings, but also try using the other voice model and see how that works.

    poisenberyMay 23, 2026
    CivitAI

    Thanks for posting this. I'm having so much fun cloning character voices and making them say sus things.

    RobopsychoJun 19, 2026
    CivitAI

    Trying to run this on a 16 gb vram lappy -- keep getting out of memory errors. weird thing is it worked twice with very little setting tweaks. Ran a 20 step 10 sec clip once. So I know it works.

    darkroast175696
    Author
    Jun 19, 2026

    In spite of being included in comfyui core, I think this voice cloning node is still very much experimental. I've had some mixed success with it especially since recent comfyui updates.

    llsporkSep 2, 2026
    CivitAI

    I just downloaded the workflow and installed the nodes, no errors but the output is not right. It may be voice but I can't understand it and it's very quiet. I tried a simpler prompt with "clear voice saying..." removing the background sound prompt but no luck. I'm on a fresh portable install of comfyui (8/2026). I couldn't find the input file used so I used a 1.5min clip of some voice-only talking from some other AI tool. Any tips or ways to debug? I don't have any nodes in red and no errors are shown when I run it. Is there a way to get the relax.mp3 so I can have the exact setup as the demo?

    darkroast175696
    Author
    Sep 2, 2026

    Honestly, I haven't tried running that workflow in quite a while and there have been a lot of updates to comfyui since then. If I get the time, I'll give it a try on my side using all the original files, but it's very possible that it just doesn't work well any more.
    I also wonder if something similar could be done with one of the newer video models to come out since they're always improving the voice processing.
    Anyway, I'll put the relax sound file up on a fileshare or something so you can give it a try, though I don't think it should matter much which voice input you use.

    llsporkSep 2, 2026

    Wow, thanks for the quick response! It's not urgent, I'm just getting into text-to-voice. If you have a workflow on civit that you think I should try that would be great and I'll try the input if you share it and see if that fixes it. I'm going to see if somehow I messed up the workflow when I was adding my downloaded models and I'll update here if that fixes it.

    Workflows
    LTXV 2.3

    Details

    Downloads
    1,216
    Platform
    CivitAI
    Platform Status
    Available
    Created
    4/6/2026
    Updated
    9/20/2026
    Deleted
    -

    Files

    textToSpeechWithVoice_v11.zip

    Mirrors

    HuggingFace (1 mirrors)