CivArchive
    LTX-2.3 Whisper & Soft-Spoken Audio LoRA - v1.0

    LTX-2.3 Whisper & Soft-Spoken Audio LoRA

    Base model: LTX-2.3 · Type: Audio-style LoRA · Rank: 32

    ---

    ## What this does

    LTX-2.3 can generate dialogue, multi-speaker scenes, and full dynamic range audio including screaming — but it cannot whisper. This LoRA adds two quiet vocal registers to the model:

    - Whispering — devoiced, breathy, close-mic delivery

    - Soft-spoken — voiced but low-volume, intimate, relaxed

    The LoRA targets only the three attention modules that write to the audio branch audio_attn1, audio_attn2, video_to_audio_attn). Video output is provably unchanged — no visual fighting, no style drift.

    ---

    ## Usage

    Load at strength 1.0. The register is controlled entirely by the manner keyword in your prompt — no special strength tuning needed.

    ### Trigger words (none, use natural language)

    | Whispering | (woman, whispering) | (man, whispering quietly) |

    | Soft-spoken | (woman, speaking softly) | (man, speaking softly) |

    > Note: Male whisper may requires the extra word quietly to tip the model over. (man, whispering) alone produces soft-spoken, not true whisper.

    ### Prompt format

    Follow the LTX-2.3 dialogue caption style:

    ```

    a [scene description], ([gender], [manner]): "[what they say]", intimate ASMR

    ```

    Examples:

    ```

    a woman sitting close to a microphone in warm dim lighting, (woman, whispering): "close your eyes and listen"

    a man at a desk late at night, (man, speaking softly): "I've been thinking about this all day"

    a woman doing a skincare routine, (woman, whispering quietly): "this is my favourite step"

    ```

    ### Without manner keywords

    Using the LoRA without any manner keyword defaults to soft-spoken — a subtle volume-softening effect on whatever the base model would have generated. Useful as a gentle "quieter audio" modifier.

    ---

    ## What it can't do

    - No intra-clip register mixing. You can't have one character whisper and another speak normally in the same clip. The register applies to the whole generation. For mixed-register dialogue, generate each part separately and cut them together.

    - No magic above the vocoder ceiling. The audio chain passes through a mel spectrogram bottleneck. Breathy whisper HF energy gets partially smoothed. Expect intimate and quiet, not studio-crisp ASMR.

    - Video is untouched by design. If you want the visuals to also feel ASMR (soft lighting, close-up framing), describe that in the scene prompt — the LoRA won't help or hurt.

    ---

    ## Training details

    | | |

    |---|---|

    | Base model | LTX-2.3 dev |

    | Steps | 2000 |

    | Rank / Alpha | 32 / 32 |

    | Target modules | audio_attn1, audio_attn2, video_to_audio_attn |

    | Training resolution | 192×192, 97 frames (~4s @ 24fps) |

    | Dataset | 74 clips, 8 voices (4F / 4M), 2 registers each |

    Clips were 4-second segments sourced from ASMR content across 8 speakers — 4 female (2 soft-spoken, 2 whisper) and 4 male (2 soft-spoken, 2 whisper). Captions used Whisper ASR transcription in (gender, manner): "transcript", intimate ASMR format.

    Description

    78 audio clips spanning 8 voices, both male and female, supporting whispering and softly spoken audio.

    FAQ

    Comments (23)

    kronos1959777Jun 14, 2026
    CivitAI

    Does this work combined with character loras?

    plz12345
    Author
    Jun 14, 2026

    I don't see why it wouldn't. It's purely audio-trained, so it didn't touch the video layer at all. However, it may fight with a character LoRA if that was trained with both video and audio.

    OneBulletJun 14, 2026

    you can also use "LTX2 Lora Loader Advanced" from Kj-Nodes. it lets you disable certain blocks (video, audio other) so you can prevent the lora from affecting i.e. video generation.

    plz12345
    Author
    Jun 14, 2026

    @OneBullet I don't actually use Comfy (MacOS here)

    bennyboy_77Jun 14, 2026
    CivitAI

    Thanks so much. This is a much needed lora. I've only just started testing but, so far, it's working great. It seems like you can use a low strength e.g. 0.3 to create a soft neutral voice or crank it all the way up to 1.0 or above to go for the full hypnosis voice!

    plz12345
    Author
    Jun 14, 2026

    Yeah, I've heard it's very situation-aware, as well. Like if the subject is further away from the camera, the strength should be adjusted. Pretty wild that LTX didn't just support this without this kind of LoRA, but it was a neat process creating it.

    treebenches785Jun 17, 2026
    CivitAI

    Exactly what I was looking for, but unfortunately the lora doesn't work in WanGP.

    plz12345
    Author
    Jun 17, 2026

    Never used it. Is that for WAN? This LoRA is not for WAN.

    leppo1001185Jun 28, 2026

    Works great in WanGP v12.288

    felipe781Jun 17, 2026· 1 reaction
    CivitAI

    Hmm got an error when loading in Wan2gp: "Sequential" object has no attribute weight.

    plz12345
    Author
    Jun 17, 2026· 1 reaction

    Never used wan2gp but this LoRA is for LTX2.3, not WAN.

    felipe781Jun 18, 2026· 1 reaction

    @plz12345 The software is called wan2gp, but inside there are many models, i used 10erosq6, from ltx 2.3.

    plz12345
    Author
    Jun 19, 2026

    @felipe781 Not sure what to tell you. Works fine for others.

    felipe781Jun 19, 2026· 1 reaction

    @plz12345 Yeah, it was an error within Wan2GP, now it's fixed by the dev. Sorry to bother you about it 😭

    sarashinaiJun 18, 2026
    CivitAI

    This is really excellent, it adds a much needed auditory awareness to LTX and a control.

    I wonder how varied your training data was, it may be bad prompting on my part, but I'm getting only a single female and male voice.

    plz12345
    Author
    Jun 18, 2026

    8 voices total, 4 female each doing sets of whispering and soft speaking, same spread for male.

    sarashinaiJun 19, 2026

    @plz12345 Did you tag the different voices in some way (e.g. woman1, woman2, man1, man2)?

    plz12345
    Author
    Jun 20, 2026· 1 reaction

    @sarashinai no, i kept it more generic captioning, like (woman, speaking softly): "Now let's go ahead and...", intimate ASMR

    sarashinaiJun 21, 2026

    @plz12345 Rightio. Perhaps something to consider for any future versions, if you decide to make them.

    plz12345
    Author
    Jun 21, 2026· 1 reaction

    @sarashinai not even really sure how that'd work. each clip and caption is its own thing, so calling one "Susie" or "Steven" has no real value.

    AnistasiataJun 28, 2026
    CivitAI

    As someone who's been listening to whisper audios for relaxation for about 15 years, but now has left Youtube because they demand I expose my computer to ad-delivered viruses to use it? This was practically created for me. Greatly appreciate it.

    So I'm clear: Are "whispering" and "soft-spoken" the exact words you'll use to prompt each type of speech?

    plz12345
    Author
    Jun 28, 2026

    You can vary it, it's not really a hard keyword trigger. Best to experiment with your own prose, as well as the LoRA strength. Start at 1.0, and you can dial up/down if needed.

    AnistasiataJul 5, 2026

    @plz12345 I've been doing so and getting pretty good results on both, still working on how to most reliably get soft-spoken volume. I'll try adjusting lora strength, didn't think of that for some reason! Either way, it works really well, thanks for the lora. :)

    LORA
    LTXV 2.3

    Details

    Downloads
    1,369
    Platform
    CivitAI
    Platform Status
    Available
    Created
    6/13/2026
    Updated
    7/14/2026
    Deleted
    -

    Files

    a_gentle_whisper.safetensors