JoyAI-Echo's performance on LTX-2.5's engine. LTX-2.5 makes picture and sound in one pass, at any length, in one generation. JoyAI-Echo - a fine-tune of LTX-2.3 - is the better actor: natural lip-sync, expressive faces, a voice that stays put. The two transformers are shape-identical, so JoyAI-Echo's video attention / feed-forward delta was transplanted onto the official LTX-2.5 dev transformer and the official LTX-2.5 distilled LoRA baked in at 0.5. Few-step files: same speed, same VRAM, same nodes as LTX-2.5 distilled. Nothing retrained.
v2. The first build put the delta straight onto the distilled transformer and came out over-saturated, with hard contrast and blown highlights - measured +16% saturation and 2-5x the clipping of stock on the same seed, obvious on any sunlit scene. Those files are withdrawn; every file here is the dev-base rebuild. The plain dev merges - for applying your own distilled LoRA at your own strength - are a separate page, Joy-LTX 2.5 DEV Merges, and on Hugging Face (joyai-echo-ltx25-echoVid-dev).
What you get
LTX-2.5 distilled is fast and clean but a flat performer; this merge keeps its speed and adds the thing Echo was trained for - mouths that shape the words, faces that move with the line, a voice that stays the same person across a long take. It runs in the stock LTX-2.5 graph, and in the ComfyUI-JoyLTX25 canvases (Civitai: Joy-LTX 2.5 workflows): a one-prompt Take sized to your card by a planner, and a Multishot that joins shots seamlessly at the VRAM of one shot. A 2-minute four-shot take rendered with it was described blind as "a single continuous take" with six seconds of dead air.
Two doses
070T30 (default) - 0.7 x Echo on video attention / FF, 0.3 x on the modulation tables, distilled LoRA 0.5. Natural grade, cleaner skin. Start here. (The files on this page.)
100T50 (strong) - 1.0 x / 0.5 x, distilled LoRA 0.5. Livelier on loud, comic, animated performances; a touch hotter on contrast. Same file names with
100T50, on Hugging Face.
Pick a version by your card, then a file by your GPU family
Every version holds two files for the same tier - one GGUF, one comfy-native. Same model; the format is the only difference.
RTX 30 / 40 (Ampere, Ada): take the
.gguf. Measured on a 3090, Q5_K_M / Q6_K are 4-8x faster than any 4-bit comfy-native file. Needs ComfyUI-GGUF installed (the JoyLTX canvas loader takes .gguf and .safetensors alike; in a stock graph useUnet Loader (GGUF)).RTX 50 (Blackwell): take the
comfy-*.safetensors. Stock ComfyUI 0.32+, plainLoad Diffusion Model. int8 / w4a8 / w4a4 / nvfp4 / mixed 4-8 all land at roughly 90-120 s per 8 s clip on a 5090; add--enable-triton-backendto the launch line or they run the eager fallback at about half speed.
16 GB cards (and 12 GB) -
Q4_K_S.gguf(12.9 GB) /comfy-w4a8(12.5 GB); 12 GB cards:Q3_K_M.gguf(10.6 GB);comfy-nvfp4for RTX 50.24 GB cards -
Q5_K_M.gguf(15.9 GB, the 24 GB default) /Q6_K.gguf(17.7 GB) /comfy-mix4x8-17.0GB(RTX 50).32 GB cards -
comfy-int8(21.5 GB, fastest on RTX 50) /Q8_0.gguf(22.7 GB).
Everything, both doses, all quants (Q2_K to Q8_0; int8, w4a8, w4a4, nvfp4, mixed 13.8 and 17.0 GB): GGUF repo · comfy-native repo.
You also need (Lightricks/LTX-2.5 on Hugging Face)
ltx-2.5-video-vae-bf16.safetensorsandltx-2.5-audio-vae-bf16.safetensors→models/vae/ltx-2.5-latent-spatial-upscaler-x2-bf16-1.0.safetensors→models/latent_upscale_models/(optional for 24 GB:ltx-2.3-spatial-upscaler-x1.5-1.0.safetensorsfrom LTX-2.3)text encoder
gemma4-12b-with-proj-ltx-2.5-comfy-int8-convrot.safetensors→models/text_encoders/(16 GB cards: the 10.6 GBgemma4-12b-ltx25-comfy-w4a8.safetensorsfrom my LTX-2.5-Quantized page)the workflows: ComfyUI-JoyLTX25 (Manager → Install via Git URL, or the zip on the Joy-LTX 2.5 workflows page - the writer with our prompts is inside it). No writer? paste your own prompt.
Dials worth knowing
video_cfg - 1.0 is the distilled default. The v2 files still sit a notch warmer than stock LTX-2.5 (saturation tracks stock, contrast +7-30% on the same seed); 0.7-0.85 brings them level, at the cost of the negative pass (~1.7x time). Same dial in the stock graph's guider.
Long takes - 30 s single pass at 1280x736 fits 24 GB; the Multishot canvas's identity anchor at 0.6 holds a 4 x 30 s take's texture flat (measured 1.00 0.88 1.11 1.10 vs 1.00 1.23 1.42 1.59 without).
Say the sounds - LTX invents drones and music for anything left unsaid; name the room tone. Say "American voice" if you want one.
Numbers
8 s at 960x544, two passes to 1920x1088: ~2 min on a 5090 (int8), ~7 min on a 3090 (Q5_K_M).
30 s single pass at 1280x736: 5 min (5090) / 17 min (3090). 12 s two-pass x1.5 on a 3090: ~7 min.
Multishot: ~2 min per 8 s shot on a 5090, ~4 min on a 3090; a 4 x 30 s take ~40 min on a 3090.
What is verified, and what is not
Verified: every quant on both repos rendered (5 s and 10 s, high-key exterior and interior prompts, same seed against stock LTX-2.5 distilled); 2-minute takes on a 5090 (int8) and 3090 (Q5_K_M); identity across cut scenes; joins below motion.
Known: the warmer grade above; Q2_K is soft and hazy (it is a 2-bit); Q3_K_M is the floor for faces.
Not claimed: lip-sync scores - reviewers scored the merge and stock within noise of each other on an 8 s talking head; the difference is in the acting you can see, not in a number.
Credits
JoyAI-Echo by JD (jdopensource/JoyAI-Echo). LTX-2.5 by Lightricks. Merge, quantisation, nodes and canvases by me. Licensed under the LTX-2.x Community License (inherited from both parents). Questions: comment here or open a discussion on the Hugging Face page - I answer.
Description
For 16 GB cards (and 12 GB with the Q3_K_M). Pick by GPU family:
RTX 30 / 40 - GGUF (needs ComfyUI-GGUF, Unet Loader (GGUF))
LTX25dist-echoVid-070T30-v2-DiT-Q4_K_S.gguf- 12.9 GB. The 16 GB default. 8 s clip ~8 min on a 3090.LTX25dist-echoVid-070T30-v2-DiT-Q3_K_M.gguf- 10.6 GB. For 12 GB cards; still looks right.
RTX 50 - stock ComfyUI 0.32+, plain Load Diffusion Model
LTX25dist-echoVid-070T30-v2-DiT-comfy-w4a8.safetensors- 12.5 GB. The 16 GB default on Blackwell (~90 s per 8 s clip on a 5090). Also runs on 30/40 series but slowly - use the GGUF there.LTX25dist-echoVid-070T30-v2-DiT-comfy-nvfp4.safetensors- 12.5 GB. Blackwell only.
Text encoder for 16 GB: gemma4-12b-ltx25-comfy-w4a8.safetensors (10.6 GB, on my LTX-2.5-Quantized page). Workflows: ComfyUI-JoyLTX25 (Manager). The 100T50 dose twins are on Hugging Face.
