CivArchive
    Preview 139991509

    Support

    Everything here is free and stays free — the format spec, the nodes, the workflows, the cartridges, the LoRAs. If it saved you a night of debugging (it contains several hundred of mine), tips keep the 5090 warm:

    Type a story. Get one continuous video, with sound. Multi-shot scenes render as a single take - no last-frame chaining, no quality loss from shot to shot. That's the whole pitch.

    Both workflows now read left to right: numbered lanes, and you only ever touch lanes 2-4. Everything below the main row is optional.

    What you need

    • ComfyUI + this node pack (Manager: MiniMax-H3 Multishot, or the zip on this version).

    • A MiniMax-H3 checkpoint (links on this page). 24 GB card? Take a GGUF.

    SEAMLESS CHAIN - multi-shot scenes as one take

    Lane by lane:

    • README - the quick start lives on the canvas itself.

    • 1 - MODELS - pick your H3 checkpoint and text encoder. VAEs are preset; LoRA slots are empty until you fill one.

    • 2 - ANCHORS (optional) - a photo to open shot 1 on (enable its gate), and a short voice clip to lock the speaker's voice.

    • 3 - YOUR PROMPTS - type your idea in the box, or point the switch at a prompt file. The writer expands it into shot prompts. Writing your own? Set the writer to passthrough (raw JSON, skip LLM) and paste shots separated by --- lines.

    • 4 - CONTROLS - size, frames per shot, steps, and take_seconds (total length; 30 is a good first run). The switches stay off unless you installed the pack a switch names.

    • 5 - ENGINE - nothing to change. The remote encoder lives here if you want the text encoder on a second PC: enter its address, flip the encoder switch, free ~15 GB.

    • 6 - OUTPUT - your video and its audio save here.

    Optional panels below the main row: reference images (your character, ref2va checkpoints - folder per character + AUTO REFS on), V2V reference (a clip whose look guides the render), FFLF plates (flf_chain mode only), audio spine (a soundtrack the take follows).

    EXTEND TAKE - one person talking, as long as you want

    Same lanes, different job: one premise becomes ONE continuous speech cut across windows.

    • 1 - MODELS - same as above.

    • 2 - ANCHORS - a photo of your speaker (shot 1 opens on them) and a voice clip. More useful here than anywhere: one person carries the whole take.

    • 3 - YOUR PROMPT - ONE premise, one speaker. The writer writes the whole speech. num_shots 0 = it decides. Passthrough works here too.

    • 4 - CONTROLS - take_seconds is the star: 30 ships, 60 clears TikTok's minute. window stays on auto - it sizes itself to your card.

    • 5 - ENGINE / 6 - OUTPUT - same as above.

    Keep takes to about 4 windows for now - very long takes slowly sharpen.

    Rules of thumb (both workflows)

    • Spoken lines: 8-12 words per shot. Short lines sync; long lines garble.

    • Say the sounds you want ("rain on the roof, a fridge hum") or it invents its own.

    • Keep your character's face in frame - faces carry identity between shots.

    If something breaks

    • Red node? Update the pack in Manager, restart, reload the workflow from disk.

    • Render crawls at low wattage? Lower resolution or frames per shot, or use the remote encoder.

    • Only one of the two workflows shows in your sidebar? Fixed in 2.6.5 - re-download both.

    • Still stuck: comment with your console log. I answer.

    Deep dives: the two articles linked on this page. Every lane also has a short note on the canvas.


    Detailed guide for people that can read good:

    Every setting explained: the Seamless Chain deep manual | Civitai

    Description

    2.6.0 — the extend take: one prompt, one continuous speech, as long as you want

    Type a length. Get a take. Set take_seconds on MASTER CONTROLS (or open H3_Extend_Take, now the main workflow, shipped at 1280x736 with a 30-second take), give the writer one premise, queue. Guides: the 5-minute guide and every setting explained. The panel sizes a window for your card (window = auto - wire the loader's MODEL into the panel for a real weight size) and the number of windows that fills the time. The writer's new extend take join style writes ONE continuous speech and cuts it across the windows at sentence boundaries - no airlock, no settle, no per-shot dialogue budget. The chain continues it under context_pin in H3's own voice. No TTS.

    Verified before shipping: seven writer-driven renders on a 32 GB card at 141/192/243-frame windows plus a 65-second, seven-window take - every join continued the speech, zero repeats, zero clipped words, reviewed blind as one uninterrupted take. Dead air at a join turned out to be a word-fill problem, not a join problem: the writer now aims for the upper half of each window's word budget (measured: 11 words in a 10-second window left a 5-second hole; the same config after the steer, 1.8 seconds).

    Known limit in 2.6.0: the chain's texture ratchet is not fully solved for long takes - we measure about +13% fine texture per join at 736x1280 with every anti-drift setting on. Under about 4 windows (30-40 seconds) it is slight; at 7 windows it is visible sharpening. Keep extend takes to about 4 windows for now; a fix is in progress.

    • Auto refs now reach the sampler on their own. The REFERENCE IMAGES lane had two switches in series - USE AUTO REFS, and a REFERENCE gate downstream that silently threw the auto refs away unless you also flipped it (the log said "3 ref(s)" and the character rendered from prose alone). USE AUTO REFS alone is enough now; the gate is the MANUAL REFS gate and only guards the two LoadImage slots. Proof line in the log: "N reference image(s) ride in every shot".

    • The writer no longer overrides your reference photographs. Same seed, photographs attached: the writer's identity sentence ("a woman in her thirties with dark hair tied back") rendered that person; the same prompt with the sentence replaced by "looks exactly as in the reference photographs" rendered the person in the photographs. The writer has a refs_attached input, wired from USE AUTO REFS in the shipped canvases; when on it points at the photos and describes nothing about face, hair, age or build.

    • Saved 2.5.x canvases keep working: the VRAM / SPEED panel dropped its five dead toggles but keeps all eight output slots in the original order (saved graphs link by slot number); the removed slots always emit false.

    • take_seconds = 0 (the default) is exactly the old behaviour; saved workflows are unaffected.

    • audio_pin_frames on the memory sampler: the audio reference window independent of the picture pin (96 = a 4-second audio memory; measured neutral so far, ships 0).

    • Reserve-planning fixes from the 24 GB test lab: no more phantom x1.6 on CORE-workflow reserves; first runs borrow measurements only from the same quant family (GGUF figures do not transfer to w4a8/int8); streamed masters embed the real graph again.

    • Writer fix: a silent shot whose boilerplate mentioned lip movement, with a voice anchor in the chain, could re-speak an earlier line. Fixed at the source, and the writer now warns by shot number.

    2.5 — the memory release

    Four new memory systems, all measured, two of them fully automatic. If renders have ever randomly crawled, if a long chain has ever killed your machine at the final join, or if the text encoder's 15 GB has ever been the difference between fitting and not — this release is about you.

    New: the random-slowdown fix (automatic)

    The single biggest fix this pack has shipped. High-resolution renders would randomly take anywhere from 27 minutes to 3 hours for identical work — same seed, same settings, a lottery. The cause: when the model weights plus working memory fill the card past roughly 95%, the Windows driver starts quietly demoting GPU memory, and whether your render crawled was luck.

    The pack now detects that zone before sampling and deliberately streams a few GB of weights instead of riding the ceiling — streamed weights are nearly free on modern ComfyUI, the last few percent of VRAM are not. The lottery render became 15 minutes, every time. Nothing to configure; the console prints a "driver headroom" line whenever it saves you.

    New: low_ram_master — long chains without the system-RAM cliff

    Off (the default), every finished shot is held in system RAM until the final join — historically the point where long renders killed whole machines, after all the sampling was already paid for. Switch it ON and each shot streams to lossless disk staging as it finishes; the master is assembled from disk through the exact same levelling math. Verified at 42.8 dB against the in-RAM path, which is codec noise. Peak RAM stays near two shots no matter how long the chain is. Turn it on if you chain long or run under 32 GB of system RAM. The finished file's path comes out of the new master_path output either way.

    New: remote text encoder — the 15 GB your render card never needed

    The text encoder works for a few seconds per shot and occupies 15+ GB the rest of the time. If you have any second PC with ComfyUI: install this pack on it, point the new H3 Remote Text Encoder node at it, turn remote_encoder ON in the VRAM / SPEED SWITCHES panel, and prompts are encoded over there while your render card keeps the memory. Results are identical — verified across two different machines to float precision — and repeated scene text is answered from a local cache with no network call at all. It ships wired into the full workflow; the panel flag defaults OFF, so single-PC setups see zero change. A how-to note sits right next to the node.

    New: H3 TAE Decode — 2-second draft previews

    A 9 MB tiny decoder (Kijai's taeh3, one small download into models/vae_approx/) turns latents into full-resolution draft frames in about 2 seconds, versus roughly a minute per shot through the real VAE. Drafts smear fine texture, but composition, framing and motion read clearly. Use it to audition seeds or triage a batch, then decode the keepers through the real VAE. Never for finals.

    New: the Speed Boosters panel — every option, measured, with the honest verdicts

    Spectrum, TeaCache, block cache and ComfyUI's own EasyCache, each behind a switch on one node. All measured on the same seed, and all eye-tested on the same finished videos:

    • block cache: no effect at 14 steps — 0 cache hits on every run on two cards (the earlier -11% did not reproduce). Ships OFF; only worth it at 30+ steps.

    • Spectrum: -29%, but visible distortion on people. Rooms and environments were fine.

    • TeaCache: -14% to -21% depending on strength, same story — people distort, environments hold.

    • EasyCache (built into ComfyUI, nothing to install): -21%, same family of people distortion.

    • Stacking Spectrum + TeaCache: no faster than Spectrum alone and severely damaged output. Never stack them.

    If your scene has no people in it, the aggressive boosters are real money. If it has people, leave them all off. A booster whose pack is not installed prints an install link and passes the model through unchanged — it no longer breaks the graph.

    Changed: defaults tuned for 16-24 GB cards

    736x1280 at 192 frames and 14 steps, euler/beta57, the curve-Q5_1 checkpoint (~14 GB), and the full anti-drift set on: context_pin, chain_gain_control=flatten, master_normalize=luma+contrast, pin_renorm, memory_frames=0. The resolution floor is deliberate — the base model distorts faces below roughly 1 megapixel, so shrinking the canvas to save VRAM costs you faces first.

    Changed: the writer now knows how long a shot really is

    Chained shots spend about five seconds of every block on the replay, airlock and settle, so the speakable span is shorter than the clip. The prompt writer now receives that real span and sizes every line against it, and overruns or fully-silent scripts are flagged at generation time with the numbers. Measured on renders: lines over budget garble, lines under budget drag, matched budgets came back natural and fully intelligible.

    Changed: writer craft rules learned from failed renders

    • Silent shots with visible people now state what mouths are doing — stops the model inventing mumbling.

    • Things revealed mid-shot are written as already present — stops pop-ins.

    • A shot that opens a mouth without dialogue must name the sound it makes.

    New: clickable title blocks

    Every workflow now carries its version number and live links to the GitHub repo, the HuggingFace page and this page, plus a plain-language pass over every on-canvas note. If you have ever wondered which version of a workflow you had open — that is over.

    Every new control in this release is appended at the end of its node's input list, so workflows you saved against 2.2.x load unchanged.

    FAQ

    Comments (13)

    ProfugoBarbatusAug 17, 2026· 1 reaction
    CivitAI

    Before I download, update and start rebuilding my tweaks, do the 2.5 version memory fixes apply to the keyframes workflow as well? I've been keeping up with the release notes and its not been clear what, if any of the seamless shot fixes improve the keyframes flow as well. Its an incredibly useful flow for specifically framed and scripted action sequences, as you can force rapid movement via tight keyframes in a way you can't with the extended takes flow.

    joeygambino
    Author
    Aug 17, 2026· 2 reactions

    Yes and no. Yes, in that the nodes exist if you've downloaded the files, and you can certainly wire them in yourself, as long as you have the packages. No, in that I haven't actually touched the workflow much since the first releases, only because people seem to like it as is and I was trying to keep a couple of them less cluttered than the main one.

    That said, I will certainly work on adding the memory optimizations I think will work well with that workflow.

    snake88Aug 20, 2026· 1 reaction
    CivitAI

    Thank you for this again - I am on 2.6.1 and noticed the usual newline --- for multi prompts manually doesnt seem to work?

    Even something like

    Test.
    ---
    Test.

    Does not result in two shots (because two prompts). My understanding is the two shots per prompt is to create more looping/variety so the above example should create 4x the shot frames if shots per prompt was set to 2?

    snake88Aug 20, 2026

    cancel that - figured it out, shots per prompt needs to be set to 0 for --- to work... leaving the comment for others.

    joeygambino
    Author
    Aug 20, 2026· 1 reaction

    You are right, and the name is the problem: shot_count is the TOTAL number of shots, not shots per prompt. At 0 the script decides - one shot per --- block, which is what you want when you have written the shots yourself. Set it to a number and that number wins: extra blocks get dropped, and a script with fewer blocks repeats its last one as a continuation. So with two blocks and shot_count at 1, block two is thrown away, which is exactly what you saw. The console says so too ("dropping 1 extra script prompt(s)"), but that is no help if you are not looking at it.

    There is no per-prompt multiplier. If you want four shots from two written blocks, write four blocks - one per shot is the whole grammar here, and each block is its own prompt for that shot.

    Next release rewords the tooltip and the docs so nobody has to work this out twice.

    snake88Aug 22, 2026

    @joeygambino ty sir makes sense!

    egin1992654Aug 20, 2026· 1 reaction
    CivitAI

    I tried cloud model for prompt but waiting it too long. Is it should work like that?

    joeygambino
    Author
    Aug 20, 2026

    It depends which cloud model. I've found the very slowest one so far for me has been minimax-m3. The fastest have been glm5.2, qwen3.5 and kimi-k2.6. I feel like GLM 5.2 has gotten the structure of my prompts right more than most. DeepSeek 4 Pro Cloud has done pretty well too now that I think about it.

    egin1992654Aug 21, 2026
    CivitAI

    I tried auto refs it was hard to understand lol. Couldnt get good work with 5 characters with 3 shorts. 3 characters in first shot then add 2 in second and no use one from first. It shows wrong characters often.

    joeygambino
    Author
    Aug 21, 2026

    Thanks for the exact numbers - 5 characters across 3 shots is precisely where the current version breaks, and you found a real limit, not user error. What's happening on 2.6.1: the automatic scan casts at most 3 characters per run (first mentioned wins). Your other two characters get NO photos - but the writer still points everyone at photographs, and a character pointed at photos that don't exist comes back as a random person. That's the wrong faces you're seeing.

    To make it work on the version you have now:

    1. Type all five folder names into the characters box on the Auto Refs node, comma separated (this skips the 3-character scan): alice,bob,carol,dan,eve

    2. Set max_per_character to 1. The model has 9 reference slots total - at 1 photo each, all five characters fit. At the default 3, the first three characters use all 9 slots and the last two still get nothing.

    3. Use each character's folder name in every shot where they appear, spelled exactly like the folder.

    4. Keep it to 2 or 3 people ON SCREEN in any one shot. All five photo sets stay attached for the whole run, so people can enter in shot 2 or sit out a shot - but five faces in frame at once will blend into each other no matter how good the references are. That one is the video model, not the node.

    Next update does step 1 and 2 for you: the scan takes up to 9 characters and splits the 9 slots automatically (2 characters keep 3 photos each, 4 get 2, five or more get 1), and the console prints the split so you can see who got what. One photo per character holds a face less firmly than three, so fewer characters per run is still the stronger setup.

    egin1992654Aug 21, 2026

    @joeygambino yes I tried it as you say
    1. Type all five folder names into the characters box on the Auto Refs node, comma separated (this skips the 3-character scan): alice,bob,carol,dan,eve

    2. Set max_per_character to 1. The model has 9 reference slots total - at 1 photo each, all five characters fit. At the default 3, the first three characters use all 9 slots and the last two still get nothing.

    3. Use each character's folder name in every shot where they appear, spelled exactly like the folder.

    4. Keep it to 2 or 3 people ON SCREEN in any one shot. All five photo sets stay attached for the whole run, so people can enter in shot 2 or sit out a shot - but five faces in frame at once will blend into each other no matter how good the references are. That one is the video model, not the node.
    but what my case is in first shot I have 3 diff characters, in second and third I have 4 characters, so it yet swaps them or mix features and dialogues.
    the automatic scan casts at most 3 characters per run (first mentioned wins) - I also thought model chose by who shows in scene earlier. it was 2,6,2 ver. I will try 2.6.3

    sebboraketti22295Aug 21, 2026
    CivitAI

    What do you think, would it be worth adding the new Sparse Attention to the latest workflow to speed things up? I’m still trying to find a way to use longer reference videos, but I haven’t had any luck yet :D For now, I’ve just been “playing around” with your workflow :D

    joeygambino
    Author
    Aug 21, 2026· 1 reaction

    Both good questions, and they have very different answers.

    SPARSE ATTENTION - yes, and it already drops in Sol-Attn (the NVlabs training-free sparse attention) has a ComfyUI extension and it was built and tested against MiniMax H3 specifically. It works as a MODEL patch, so it goes between your model loader and the sampler - nothing in the chain workflow needs changing.

    It finds H3's video segment on its own and keeps the conditioning rows (text, audio, reference) exact, which is the part that matters: those rows are what hold identity and voice, and you do not want them approximated.

    Three things worth knowing before you judge it:

    1. First run is SLOWER. The Triton kernels compile on first use. Judge the speed on run two, not run one. On Windows you need a Triton build matched to your torch/CUDA - the plain pip package is usually the wrong one.

    2. Start with start_percent above 0. Early steps set composition and identity; let those run dense and sparsify the later ones. tau is the quality/speed dial.

    3. On CHAINED renders be more conservative than you would be on a single clip. Every shot conditions on the previous shot's output, so any per-shot approximation error compounds down the chain rather than staying put. If shot 6 looks worse than shot 1, back start_percent up. A single take or a retake has no feedback loop and can take a more aggressive setting.

    The extension also ships a block probe that reports sparse-vs-dense error per block. If you want to tune it honestly rather than by eye, that is the tool - it will tell you which blocks tolerate sparsity and which do not.

    It is marked experimental by its author, so I am not going to bake it into the shipped canvas yet. Wire it in yourself and it works today.

    LONGER REFERENCE VIDEOS - this one is a wall, not a missing setting

    You have not found the trick because there is not one. Two separate limits:

    - H3's reference videos were trained at roughly 2-15 seconds. Past that you are outside what the model learned to use, and it does not degrade gracefully - it just stops contributing.

    - Cost. A reference video is subsampled to 2 fps for the text encoder, but the frames are ALSO VAE-encoded into reference rows that are re-read on EVERY sampling step. Double the reference length and you pay that on every step of every shot. That is why the node prints a warning past 64 frames.

    So a long reference does not fail loudly, it just gets expensive and stops helping. What actually works instead:

    - Trim to the 2-4 seconds that contain what you are referencing, with H3 Reference Video (start_seconds + seconds). A short clip of the RIGHT thing beats a long clip containing it. Pick the window where the subject is well lit and facing roughly forward.

    - For identity across a long piece, use reference IMAGES rather than a long video. Two or three stills per character are cheaper and hold a face harder than a clip does.

    - For continuity across a long piece, that is what the chain modes are for - context_pin carries the previous shot's actual latents forward, which is a much stronger link than any reference video, and it costs nothing extra.

    If what you want is "this whole 30-second clip's look", reference a couple of stills from it plus a written description of the grade, rather than the clip itself.

    Workflows
    MiniMax H3

    Details

    Downloads
    688
    Platform
    CivitAI
    Platform Status
    Available
    Created
    8/17/2026
    Updated
    8/25/2026
    Deleted
    -

    Files

    minimaxH3MultishotSeamlessChain_v26ExtendTake.zip

    minimaxH3MultishotSeamlessChain_v26ExtendTake.zip