Everything here is free and stays free — the format spec, the nodes, the workflows, the cartridges, the LoRAs. If it saved you a night of debugging (it contains several hundred of mine), tips keep the 5090 warm:
🔁 Liberapay (recurring)
⚡ Or right here: the Civitai tip button on this page sends Buzz directly.
LTX-2.3 American Accent LoRA — make accent prompts actually work
Makes accent wording in your prompt actually work. LTX-2.3's voice prior ignores accent requests in exactly the regions where you need them most — young female characters default to Australian-leaning voices even when the prompt says "in a casual American accent". This LoRA turns that wording into a reliable control.
Read this first: the 24 fps rule
This LoRA cannot help you at the wrong frame rate. LTX-2.3's joint audio-video prior is 24 fps-native, and render fps is a hidden accent dial: at 25 fps the same prompt and seed render non-rhotic southern British, at 30 fps broad Australian — and off-24 fps overrides accent wording entirely, LoRA or no LoRA. Set your workflow's fps/frame_rate widgets to 24, then use this LoRA. (Dose-response verified by A/B on identical configs.)
What it does — measured
A/B matrix on the hardest region (young pale-skinned woman, casual camcorder monologue), 24 fps, two seeds per cell, blind phonetic review:
spoken-line wordingwithout this LoRAwith it @ 1.0nothing statedAustralianAustralian"saying in a casual American accent"Australian (both seeds)General Americanrich scaffold (voice timbre + accent binding)Australian (both seeds)General American
Read the top row carefully: this LoRA does not force American unconditionally. It makes the model obey the accent you ask for where the base model refuses. Ask for nothing and you still get the base model's lean — so state the accent on every spoken line and let the LoRA do the enforcing.
Known ceiling — read before you try for a regional accent
It enforces General American and flattens regional sub-flavours. "Soft Southern" and "Boston-flavored" both render as General American in testing. If you need a specific regional dialect, this is not that tool yet — that needs dialect-labelled training data, not different prompt wording.
Lip-sync safe by construction
1,152 LoRA tensors, all in the audio branches (audio_attn1 / audio_attn2 / audio_ff). Zero video tensors, zero cross-modal tensors. Faces, video content, and lip motion are mathematically untouched.
Usage
Strength 1.0 in any LTX-2.3 LoRA loader. Verified in long multishot production runs alongside a video-branch LoRA with no interference.
Compatible with the LTX-2.3 22B family: stock distilled 1.1 and the JoyAI-Echo surgical merges.
Render at 24 fps. Say the accent. That is the whole recipe.
Training
ai-toolkit, rank 32 / alpha 32, 3,000 steps, lr 1e-4, qfloat8. American-English read speech from LibriSpeech (CC BY 4.0), muxed over static video so only the audio lane carries signal; captions are verbatim transcripts with the accent deliberately unnamed, so the American prior trains as always-on behavior rather than a trigger phrase. The audio-branch-only module scope keeps the static training video harmless — the video branches never receive a gradient.
Also on HuggingFace: joeygambino/ltx23-accent-american-audio-lora
Description
FAQ
Comments (13)
Hihiii. this is cool. does it work with the word "water" ?
I'm honestly not sure what you mean. Were you having a problem with people saying the word water?
@joeygambino i think its a joke about british and aussie accents saying "water" :p some british say "wo' ah"
woter mate
Very nice! Hope you'll continue working on this concept. Would be great to get a generalised 'Southern' accent.
I am hoping to have some other US-based accents soon and then work my way out internationally.
Nice work, not enough people are working on accent loras, would love to see more for not only the US but other English speaking countries though I appreciate it's probably an assload of work!
I am working on more! A little trick for LTX (although this is also VRAM intensive) - as you raise FPS past 24, other accents emerge. You start getting British/Aussie accents at 25-30 fps - unfortunately it's kind of luck of the draw.
@joeygambino I've had quite good luck with British accents (a generic southern one at least) most of the time, even at 24fps. It seems to restrict itself to the same 1 or 2 American voices when using certain loras, assume that has something to do with how they were trained but am glad someone is thinking about it at least!
would love when this gets specific accents like different country accents in particular.
It's on the list!
Hi.
Any specific tutorial you would suggest for someone trying to train a voice/accent?
I'll see if I can come up with a basic set of instructions and post them in articles. Voice and dialog datasets can be found free (OpenSLR, Mozilla Common Voice), so it's easy to get training data.
Details
Files
ltx23_accent_american_v2_rank32.safetensors
Mirrors
ltx23_accent_american_v2_rank32.safetensors
American Accent - LTX2.3 - saying in a casual American accent,australian accent,general american accent.safetensors
American Accent - LTX2.3 - saying in a casual American accent,australian accent,general american accent.safetensors
American Accent - LTX2.3 - saying in a casual American accent,australian accent,general american accent.safetensors