CivArchive
    MiniMax H3 Continuum – Long-Form Video & Audio for ComfyUI - v3.6
    NSFW
    Preview 140985456

    Development Focus: Stabilization & Optimization

    What’s New in V3.7

    V3.7 focuses on faster, more reliable high-resolution Second Pass refinement. The new Conditioning Adapter rebuilds First/Last images and continuation context at the target resolution, while RefineSchedule provides exact Full, Tail, Partial, and External refinement ranges.

    Using Tail 6 instead of Tail 10 reduced Second Pass sampling from 100.27s to 64.18s in our reference test—about 36% faster—while preserving the original first-pass audio bit-exact.

    Production defaults remain unchanged, and existing V3.6 workflows remain compatible. The new Still Image Guide is included as an Experimental feature because hard anchors may cause abrupt trajectory changes.

    What’s New in V3.6.1

    https://github.com/ukr8b3g-cmyk/ComfyUI-H3-Continuum

    V3.6 introduces a new Masked AV Continuation backend for faster and cleaner long-form generation.

    Instead of adding the previous chunk again as a separate Reference block, V3.6 keeps the finalized Video/Audio latent prefix directly inside the next target and samples only the new region.

    • Masked AV Continuation — new Standard backend

    • Less Attention overhead — FL2VA test reduced packed rows by 8.01%

    • Faster continuation Sampling — median Terminal Group sampling improved by 4.06%

    • Bit-exact Video + Audio prefix preservation

    • T2VA / I2VA / FL2VA supported

    • Reference Image + Reference Audio supported

    • Long Terminal Merge fully supported

    • Run Storage / Resume / Regenerate From supported

    • Compatibility mode keeps the previous V3.5 Reference Context backend

    • Chunk duration expanded to 4–30 seconds — 5–15 seconds remains recommended

    V3.6.1 Hotfix

    • Improved mixed Timeline / --- prompt syntax handling and warnings

    • Prevents unintended repeated chunk prompts

    • Non-Balanced Audio Continuity settings safely fall back to the compatible Reference Context path instead of stopping generation

    Same sampling quality, less redundant continuation work, and full compatibility with existing V3.5 workflows.

    H3 Continuum V3.5.2 — Stabilization & Optimization Update

    V3.5.2 focuses on stability, efficiency, and cleanup rather than adding major new features. Repeated runs with the same Prompt/CLIP conditions now avoid unnecessary re-encoding, and Video Guide preprocessing uses substantially less temporary RAM on longer inputs.

    Existing V3.5.x workflows and generation contracts remain unchanged. Sampling, Terminal Merge, Run Storage, Reference handling, and output behavior were preserved while the updated paths were validated with CPU and GPU A/B testing.

    V3.5.2 is an internal stabilization and optimization update. Existing V3.5.1 workflows can be used as-is, with no changes required to nodes, connections, or workflow structure.

    > v3.5.1 Update

    This update improves the Second Pass workflow, reference handling, and seed behavior. It adds the new Conditioning Bridge V3.5 for external sampling workflows, optional Reference Audio conditioning.

    Video reference inputs are also clarified as Video Guide Frames / Video Guide Size, while existing V3.4/V3.5 workflows and backend connections remain compatible.

    Updating to V3.5 does not automatically replace V3.4 nodes in saved workflows. When updating an older workflow, replace not only the Sampler but also the assembler with H3 Continuum Assemble + Seam V3.5. Add H3 Continuum Hi-Res Fix V3.5 when Hi-Res Fix is required.

    GitHub:
    https://github.com/ukr8b3g-cmyk/ComfyUI-H3-Continuum

    [ V3.5 ]

    V3.5 introduces two major additions:

    - Continuum-aware Second Pass / Hi-Res Fix

    Refine externally processed H3 latents while preserving Continuum physical groups, prompts, seeds, ordering, and first-pass audio. An integrated one-node 2x Hi-Res Fix path is also included as an experimental feature.

    - Low-memory Assemble + Seam V3.5

    Adds Auto, RAM, and Disk-backed video-buffer modes. Disk-backed assembly significantly reduces system RAM/private-memory usage for long or high-resolution outputs while preserving Exact Duration, Seam, Terminal Merge, and audio behavior.

    All V3.4 nodes remain available for saved-workflow compatibility. Existing V3.4 workflows continue to work unchanged.

    The V3.5 release passed 430 automated tests and representative GPU acceptance tests.

    Note: The integrated Hi-Res Fix remains experimental. Long 2x workflows can require substantial GPU VRAM.

    #5 3 chunks x 15 seconds

    #6 6 chunks x 15 seconds

    W576xH576

    [ v3.4 ]

    Long-form MiniMax H3 video and audio generation for ComfyUI with chunked generation, persistent references, restartable runs, and user-controlled audio.

    ### What's new in v3.4

    - Driving Audio: preserves the supplied audio as the final audio while guiding generation across chunks.

    - Video Reference: provides persistent visual reference for identity, motion, framing, and scene appearance.

    - Restartable chunks: reuse completed chunks with Run Storage and regenerate only the required part.

    - Improved Core compatibility: unknown upstream or custom nodes are not rejected merely because they are not recognized by Continuum.

    - Simpler stable interface: obsolete compatibility controls and experimental Timeline inputs are hidden from the V3.4 public workflow.

    - Spectrum interoperability: Spectrum remains optional and can use the official H3 Continuum Interop API.

    ### Direction change from v3.3

    V3.4 focuses on predictable reference workflows rather than experimental Timeline Video and timeline-audio generation.

    Driving Audio preserves the original user-supplied audio. Video Reference provides persistent visual guidance without requiring exact frame-by-frame copying. Existing V3.3 workflows remain available through legacy compatibility paths.

    ### Updating

    For an existing Git installation:

    git pull --ff-only origin main

    V3.4 input connection patterns

    V3.4 separates the visual reference input from the driving-audio input. Choose the connection pattern that matches your source material.

    1. Audio only

    Connect Load Audio to driving_audio. Use this when an existing song, dialogue track, or sound effect should remain the final audio. A Video Reference is not required.

    Driving Audio connection

    2. Video with its own audio

    Connect Load Video (Upload) IMAGE to Video Reference. If the uploaded video contains the audio you want to preserve, connect its AUDIO output to driving_audio as well.

    Video Reference and embedded audio connection

    3. Video and audio from separate sources

    Connect Load Video (Upload) IMAGE to Video Reference, then connect a separate Load Audio node to driving_audio. Use this when the visual reference video and the final audio source are different files.

    Separate Video Reference and Driving Audio connection

    Both inputs are optional. Connect Video Reference when visual guidance is needed, and connect driving_audio when the supplied audio should be preserved in the final output.

    Video Reference frame rate

    Use a 24 fps source for Video Reference. Load Video (Upload) may accept files recorded at 25 fps or another frame rate, but acceptance alone does not guarantee correct temporal alignment with H3. For a non-24 fps source, set force_rate to 24 in Load Video (Upload), or convert the file to 24 fps before loading it. If the source is already 24 fps, leave force_rate at its default and do not resample it.

    Current validation status

    [ v3.3 ]

    V3.3 adds Timeline Video conditioning for long-form MiniMax H3 generation. A reference video can now be processed in chunk-local time slices, allowing motion and scene continuity to be carried across multiple 5-second chunks while keeping the reference resolution independent from the output resolution. The Efficient 0.4 MP mode helps reduce memory usage and processing time.

    Video assembly has also been improved. Auto seam handling analyzes chunk boundaries and applies guarded corrections for transient flicker, micro-flash, exposure, and color differences. This helps produce more natural transitions between generated chunks without changing the original sampling process.

    Existing V3.2.4 workflows remain available as Legacy nodes for compatibility.

    [ v3.24 ]


    Generate longer native MiniMax H3 video and audio sequences in ComfyUI.

    H3 Continuum is a ComfyUI custom node that generates a longer sequence as connected chunks and assembles them into one continuous video.

    ```text
    3 × 5-second chunks → 15-second video
    6 × 5-second chunks → 30-second video

    The previous video and audio latent context is passed into each continuation chunk. This is not a simple video concatenation workflow.

    Main purpose: longer MiniMax H3 generation, not faster generation.

    Easy Installation

    H3 Continuum can be installed directly from ComfyUI Manager.

    1. Open ComfyUI Manager

    2. Search for H3 Continuum or Continuum

    3. Select Install

    4. Restart ComfyUI

    5. Load one of the included sample workflows

    Manual installation and the latest documentation are available on GitHub:

    GitHub:
    https://github.com/ukr8b3g-cmyk/ComfyUI-H3-Continuum

    What It Does

    H3 Continuum divides a longer generation into manageable chunks.

    MiniMax H3 Model
    ↓
    H3 Continuum Sampler
    ↓
    ComfyUI Core Video / Audio VAE Decode
    ↓
    H3 Continuum Assemble
    ↓
    Final video

    Each continuation chunk receives latent context from the preceding chunk. Overlapping context is removed during assembly, and the final frame and audio counts are aligned to the requested duration.

    Main Features

    • Connected long-form MiniMax H3 generation

    • Native video and audio latent continuation

    • Fixed, List, and Timeline prompt formats

    • Automatic prompt-format detection

    • T2VA, I2VA, FL2VA, Last Frame and Reference workflows

    • Up to three Reference Images

    • Reference Audio conditioning

    • First Frame and Last Frame conditioning

    • Configurable continuity context

    • Run Storage and automatic resume

    • Partial regeneration from a selected chunk

    • Optional Spectrum interoperability

    • Standard and Turbo sample workflows

    • ComfyUI Core VAE Decode compatibility

    Included Sample Workflows

    Two example workflows are provided.

    Standard Workflow

    Recommended when output quality and temporal consistency are the priority.

    • Standard MiniMax H3 sampling

    • Spectrum can be enabled

    • Suitable for quality-focused generation

    • Reference Image and Reference Audio supported

    • RTX upscaling can be enabled when required

    Turbo Workflow

    Recommended for faster tests and iteration.

    • LightX2V MiniMax H3 Turbo LoRA

    • 8-step example configuration

    • Spectrum is bypassed by default

    • Faster than the standard workflow in tested configurations

    • Some loss of facial detail or additional artifacts may occur

    Turbo LoRA models:

    https://huggingface.co/lightx2v/Minimax-h3-Turbo/tree/main

    MiniMax H3 models and documentation:

    https://huggingface.co/MiniMaxAI/MiniMax-H3

    Models and LoRAs are not included with this custom node.

    Reference + Continuation

    Reference Images remain available across all generated chunks.

    A typical setup is:

    Picture 1 → face and identity
    Picture 2 → full-body appearance and clothing
    Picture 3 → environment or an additional visual reference
    Audio 1   → vocal, music or audio-performance reference

    Ref2VA is the reference-specialized checkpoint and is generally the first choice for stronger reference fidelity.

    FL2VA with Reference conditioning is also allowed. H3 Continuum does not automatically replace or switch the connected model.

    Spectrum Integration

    Spectrum is optional. H3 Continuum also works without it.

    With a compatible Spectrum release, H3 Continuum sends a continuation signal only when generating later chunks.

    Chunk 1 → normal Spectrum sampling
    Chunk 2+ → Continuum Actual Prefix 2

    This allows Spectrum to coordinate its spectral forecasting with the continuation context instead of treating every chunk as an unrelated generation.

    Benefits include:

    • Automatic identification of continuation chunks

    • Actual Prefix applied only where required

    • No manual prefix switching between chunks

    • Reduced risk of duplicated prefix processing

    • Compatibility with standard ComfyUI workflow execution

    Spectrum remains an approximate accelerator. Motion, anatomy, audio and detail can differ from a non-Spectrum result, so quality comparisons should use the same prompt and seed.

    Spectrum:

    https://github.com/xmarre/ComfyUI-Spectrum-MiniMax-H3

    Run Storage and Resume

    Enable Save + Auto Resume to preserve completed raw video and audio chunks.

    If a generation is interrupted, H3 Continuum can reuse compatible saved chunks and continue from the first missing chunk.

    It can also regenerate from a selected chunk while preserving the compatible prefix.

    Chunk 1–3 completed
    ↓
    Generation interrupted
    ↓
    Queue the workflow again
    ↓
    Chunks 1–3 reused
    ↓
    Generation continues from Chunk 4

    Run Storage verifies the sampling contract, model route, references, resolution and saved chunk files before reuse.

    Prompt Formats

    Fixed

    One prompt is used for every chunk.

    List

    Separate prompts are divided with:

    ---

    Timeline

    [0-5s]
    First scene description
    
    [5-10s]
    Second scene description
    
    [10-15s]
    Third scene description

    Prompt Format = Auto detects the appropriate format automatically.

    Incomplete timeline coverage produces diagnostics and safe fallback behavior rather than unnecessarily stopping every generation. Structurally unusable input is still reported as an error.

    Tested Configuration

    The current Windows implementation has been tested with:

    GPU                 NVIDIA RTX 5060 Ti 16GB
    ComfyUI             MiniMax H3-compatible Core build
    Chunk Duration      5 seconds
    Typical Length      3 or 6 chunks
    Continuity          Balanced 22 frames
    Standard Sampling   RES Multistep
    Spectrum Interop    Actual Prefix 2

    The node is not limited to RTX 50-series GPUs. Actual compatibility, generation speed and usable resolution depend on the MiniMax H3 model, GPU memory, ComfyUI configuration and installed acceleration nodes.

    RTX 4060 and other configurations have not been formally validated by this project.

    Frequently Asked Questions

    Is this only a workflow?

    No. H3 Continuum is a ComfyUI custom node package. The included workflows are ready-to-use examples.

    Does it generate one native 30-second sample?

    No. It generates connected chunks and assembles them into one longer output while carrying video and audio latent context forward.

    Does it make MiniMax H3 faster?

    Speed is not the primary purpose. H3 Continuum is designed for longer generation. Spectrum and Turbo LoRAs can reduce generation time in some configurations.

    Is Spectrum required?

    No. It is an optional acceleration and interoperability path.

    Can I use the Turbo LoRA?

    Yes. A Turbo sample workflow is provided. Spectrum is bypassed by default in that workflow because combining both can change quality or introduce artifacts.

    Which model should I use for Reference Images?

    Ref2VA is the reference-specialized option. FL2VA with Reference conditioning is also allowed, but reference fidelity may differ.

    Are the models included?

    No. MiniMax H3 checkpoints, text encoders, VAEs, Turbo LoRAs and optional acceleration nodes must be installed separately.

    Can an interrupted generation be resumed?

    Yes. Enable Run Storage before generation. Compatible completed chunks can then be reused.

    Can I regenerate only the later part?

    Yes. Run Storage supports regeneration from a selected chunk while retaining a compatible earlier prefix.

    Are chunk boundaries always invisible?

    No generative continuation system can guarantee a completely invisible boundary. H3 Continuum preserves latent context and removes duplicated overlap, but difficult motion, lighting changes and large prompt transitions can still produce flicker or visual changes.

    Does Reference Audio guarantee exact lip synchronization?

    Reference Audio conditions MiniMax H3’s native joint video/audio generation. It can guide vocals, rhythm, expression and mouth movement, but it does not guarantee sample-identical audio reproduction or frame-perfect lip synchronization in every generation.

    Does it support audio continuity?

    Yes. Video and audio latent context are carried together. The assembler also provides an optional Audio Seam mode for boundary-local audio correction.

    Is RTX 5090 required?

    No. Development and runtime validation were performed on an RTX 5060 Ti 16GB. Lower-memory configurations may require reduced resolution, offloading or other ComfyUI memory optimizations.

    What license is used?

    H3 Continuum is released under the MIT License.

    Custom Nodes

    The included Standard and Turbo workflows use the following custom nodes.

    - H3 Continuum

    https://github.com/ukr8b3g-cmyk/ComfyUI-H3-Continuum

    - rgthree-comfy

    https://github.com/rgthree/rgthree-comfy

    - ComfyUI-Easy-Use

    https://github.com/yolain/ComfyUI-Easy-Use

    - ComfyUI-KJNodes

    https://github.com/kijai/ComfyUI-KJNodes

    - ComfyUI-Spectrum-MiniMax-H3

    https://github.com/xmarre/ComfyUI-Spectrum-MiniMax-H3

    - NVIDIA RTX Nodes for ComfyUI

    https://github.com/Comfy-Org/Nvidia_RTX_Nodes_ComfyUI

    Spectrum and RTX upscaling are optional generation paths, but installing all listed custom nodes allows the included workflows to load without missing-node warnings.

    Models

    - MiniMax H3

    https://huggingface.co/MiniMaxAI/MiniMax-H3

    - LightX2V MiniMax H3 Turbo LoRA

    https://huggingface.co/lightx2v/Minimax-h3-Turbo/tree/main

    Models and LoRAs are not included in the workflow ZIP.

    Main Links

    - GitHub and documentation

    https://github.com/ukr8b3g-cmyk/ComfyUI-H3-Continuum

    - Install from ComfyUI Manager

    Search for H3 Continuum

    Description

    FAQ

    Comments (23)

    koalabearrAug 27, 2026
    CivitAI

    Has anyone tried it with 8GB of VRAM?

    hatt2Aug 27, 2026

    I'm on 8GB VRAM. 32 RAM. Works well for me.

    koalabearrAug 27, 2026

    What's your production speed for a 15-second video?@hatt2 

    abeslu425Aug 27, 2026
    CivitAI

    Hi, If I need to modify only the second part keeping the first part of the generation is it possible with this workflow?

    ukr8b3g201
    Author
    Aug 28, 2026

    Yes, it is possible, but Run Storage must be enabled during the original generation.

    For example, if your video is configured as 2 chunks × 5 seconds, Chunk 1 is the first 5 seconds and Chunk 2 is the second 5 seconds.

    Before the first generation:

    Set Run Storage to Save + Auto Resume.

    Give the run a fixed Run Name.

    Keep a fixed Base Seed.

    Generate the complete 10-second video normally.

    Continuum will save the generated latent data for both chunks.

    To create a different second part:

    Keep the same Run Name, Base Seed, model, LoRA, resolution, sampler, steps, references and Continuation settings.

    Set Regenerate From to Chunk 2.

    Queue the workflow again.

    Continuum will load the saved Chunk 1 instead of sampling it again. It will use the end of Chunk 1 as the continuation context and generate a new Chunk 2. The status report should show:

    1 reused, 1 generated

    This means the first 5-second chunk was preserved and only the second 5-second chunk was sampled again. The complete 10-second video is then decoded and assembled again.

    If you want different instructions for each part, prepare separate prompts before the first generation:

    Prompt for the first 5 seconds
    ---
    Prompt for the second 5 seconds

    Changing the model, LoRA, resolution, sampling settings, references or prompt plan after the original run may create a new Run Storage revision. In that case, Continuum may need to regenerate the complete video.

    If the original video was generated with Run Storage = Off, the first chunk was not saved and cannot be reused afterward.

    Important exception: an FL2VA workflow using Long Terminal Merge samples the final two chunks as one unit, so those two chunks must be regenerated together.

    abeslu425Aug 28, 2026

    @ukr8b3g201 Thanks!!

    hatt2Aug 29, 2026· 2 reactions
    CivitAI

    Back with more questions. 🙂
    Is it possible for us to change the Shift?
    - Video shift
    - Audio shift

    ukr8b3g201
    Author
    Aug 29, 2026· 1 reaction

    Yes. Current ComfyUI Core provides a ModelSamplingMiniMaxH3 node with separate controls for:

    Video shift — default 12.0

    Audio shift — default 3.0

    Place it after the MiniMax H3 diffusion-model loader and before the rest of the model chain connected to Continuum. Continuum already preserves this official upstream model patch, so no special Continuum implementation is required.

    Changing either value alters the denoising schedule, so it should be treated as an advanced/experimental setting and regenerated from scratch rather than reusing an older Run Storage revision. I have not established recommended alternative values yet.

    hatt2Aug 29, 2026

    @ukr8b3g201 Ok thank you for confirming!

    themarkedone830Aug 29, 2026
    CivitAI

    thanks for the workflow. 2 things could any share some of the prompts cause i not think i fully understand how to prompt for this and does any one know why im getting random dialog that is not in the prompt

    thanks

    CannBoyoAug 29, 2026

    IF I understood it correctly which seems to work in my prompting...separate chunks (video parts) as [Chunk 1] etc, then in the main thing you set how many chunks to generate, and the chunk length aka seconds/duration of each chunk
    for example
    [Chunk 1]
    bla bla bla character picture 1 bla bla
    she walks in and waves
    [Chunk 2]
    bla bla bla character picture 1 bla bla
    she walks out and closes the door


    random dialogue (if using reference audio) might be due to sampler/scheduler, too few or too many steps, random words outside of desired places, like putting trigger words in the beginning of the prompt (happened to me several times), or you don't shut up the character by saying something like 'she finishes speaking her sentence'
    after the dialogue
    OR you have some other trigger words in the prompt which can have a starting effect, like "says" "said" "voice" etc

    themarkedone830Aug 29, 2026

    ya my prompt look the same and ive put the she finishes speaking her sentence at the end of dialog and tried removing my audio ref still doing. could you tell me what sampler/scheduler your using and any speed lora and how many steps you do please. thank you

    CannBoyoAug 29, 2026· 1 reaction

    @themarkedone830 right now for this workflow Euler - beta57, 4 steps first run and 4 steps 0.4 denoise second run with x1.0 upscale factor... seems aight but I'm still experimenting
    using minimax_h3_fl2v_lightx2v_v0.1_dareties_v4_step600_comfy_fro_no_adaln_proj.safetensors lora on 1 strength, perhaps important factor.. I'm using a hybrid checkpoint called minimax_h3_hybrid_fl2va_ref2va_b25-49-int8.safetensors

    hatt2Aug 29, 2026· 1 reaction

    There's a bunch of different ways to prompt. You can see another method from the skateboarding video.

    https://civitai.red/models/2860061/minimax-h3-continuum-long-form-video-and-audio-for-comfyui?modelVersionId=3230627

    XarfaiAug 30, 2026

    I tend to put complete prompts for every segment, so far gives me the best results, since I mainly do longer segments (15s chunks). Technically shouldnt make a difference, but its what I had the best experience with.
    The checkpoint I use is the minimaxH3Ref2vaXUELUO_v10.safetensors one, with int8 Text encoder and VAE (links on the civitai site of the checkpoint)
    Then is use an 8 Step lora minimax_h3_fl2v_turbo_8step_v1.0_comfyui_bf16.safetensors
    Sampler is er_sde and simple scheduler.
    At 1mp with 2x RTX upscale that usually gives me enough quality that I dont need the 2nd pass, but if I want, ill add the 10 steps with only 1x scale, just to refine the video.

    Also I find that seed hunting is very much helpful, I find myself often just doing a 5s segment of the video to find a good starting scene and then screenshotting it and using it as a reference image to cut down on the variance.

    hatt2Aug 30, 2026

    For now this format is working for me. If I deviate from this I get weird results. Like bad camera work or voices that shouldn't be talking etc


    Line 1 [Chunk 1]

    Line 2 integrated_multimodal_description: [Shot 1] detail scene

    Line 3 overall_soundscape: detail sound stuff

    Line 4 non_diegetic_music: N/A

    Line 5

    Line 6 [Chunk 2]

    Line 7 integrated_multimodal_description: [Shot 1] detail scene

    Line 8 overall_soundscape: detail sound stuff

    Line 9 non_diegetic_music: N/A

    Line 10

    Line 11 [Chunk 2]

    Line 12 integrated_multimodal_description: [Shot 1] detail scene

    Line 13 overall_soundscape: detail sound stuff

    Line 14 non_diegetic_music: N/A

    Line 15

    gigomikolAug 30, 2026
    CivitAI

    This is my favorite workflow, and tonight v37 dropped, however i have a few questions (am novice)
    1. for v37 it says i need the continuumv37 sampler, and after git clone of the repo, and reboot i still dont have this node.
    2. i still dont understand fully the chunks, i can do a 30 second video by selecting 15seconds @ 2 chunks, but my output is the same video twice in 30 seconds, like a loop, not a continuation or a 30 second long single scene.
    3. how do i change the resolution of my image? i see the 4mp and 6mp options, but unsure of how to upscale this to get a fhd 2k or 4k output.

    thank you for your work on this workflow, its one of my favorites, but im sure im barely using it as intended.

    ukr8b3g201
    Author
    Aug 30, 2026

    Thank you — I’m glad you enjoy the workflow! Here are answers to each question:

    1. V3.7 Sampler node is missing

    The display name is *H3 Continuum Sampler V3.7**.

    Please confirm that:

    - The repository is directly under ComfyUI/custom_nodes/ComfyUI-H3-Continuum

    - ComfyUI is version 0.32.0 or newer

    - You completely restarted the ComfyUI backend, then refreshed the browser

    - You do not have both a Manager installation and a second manual clone

    For an existing Git installation:

    ```bash

    cd ComfyUI/custom_nodes/ComfyUI-H3-Continuum

    git pull --ff-only origin main

    ```

    During startup, you should see something similar to:

    ```text

    H3 Continuum 3.7.0 loaded

    ```

    If that line does not appear, please send the startup-console error around ComfyUI-H3-Continuum.

    The V3.6 sampler remains compatible, so you can continue using it while resolving the installation. V3.7 does not change the existing Production defaults.

    2. The 30-second result appears to repeat

    Continuum does not generate one native 30-second H3 sample. With 2 chunks × 15 seconds, it generates two physical sections and passes the end context of the first section into the second.

    If both chunks receive the same Fixed prompt, H3 may repeat the same described action. For more deliberate progression, use Chunk/List prompts:

    ```text

    [Chunk 1]

    The woman enters the hallway. The camera follows her from behind.

    [Chunk 2]

    Without a cut, she continues forward, turns right, and enters the next room.

    The camera keeps tracking her in the same direction.

    ```

    Avoid restarting the action in Chunk 2—for example, do not say “she enters the hallway” again.

    For initial testing, I recommend:

    - Standard continuation

    - Balanced — 22 frames

    - Audio Continuity ON

    - 3 × 10 seconds or 6 × 5 seconds

    - A different prompt section for each chunk

    If the two halves are literally identical rather than merely similar, temporarily set Run Storage = Off, disconnect Video Guide Frames if used, and check the Detailed Report for reused chunks.

    3. Resolution and 2K/4K output

    The Efficient 0.4 MP and Balanced 0.6 MP controls are for Reference/Video Guide preprocessing. They do not set the final output resolution.

    The First Pass output resolution comes from the workflow’s Resolution Selector, which feeds the Sampler width and height. H3 is best used near its native canvas—up to approximately a 768-pixel short edge and 768×1344, aligned to multiples of 32.

    For FHD/2K/4K, the practical approach is:

    ```text

    Generate near native H3 resolution

    → Assemble the complete video

    → Apply a video upscaler afterward

    ```

    The workflow also includes the Experimental H3 Continuum Hi-Res Fix V3.5. A 2× setting can produce a higher-resolution generative Second Pass, but it is much slower and uses substantially more VRAM. A single 576→1152 five-second test passed, while long 2× sequences can exceed 16 GB VRAM.

    Therefore, for a 30-second video I recommend generating at native H3 resolution first and using a post-process video upscaler for 1080p, 2K, or 4K.

    1. The V3.7 Sampler node is missing

    The exact node name is H3 Continuum Sampler V3.7.

    Cloning the repository does not update an older installation that may already exist. If you previously installed Continuum with Git, please run:

    cd ComfyUI/custom_nodes/ComfyUI-H3-Continuum

    git pull --ff-only origin main

    Then completely restart the ComfyUI backend and refresh the browser.

    If you installed it through ComfyUI Manager, use Manager’s Update button instead. Please do not keep both a Manager installation and a second manual Git clone, because duplicate copies can cause ComfyUI to load the wrong version.

    During startup, look for:

    H3 Continuum 3.7.0 loaded

    Then search for the exact node name:

    H3 Continuum Sampler V3.7

    If it is still missing, please share the startup-console lines mentioning ComfyUI-H3-Continuum. The V3.6 sampler remains compatible while the installation issue is being resolved.

    gigomikolAug 30, 2026

    @ukr8b3g201 What a great and detailed reply, Thank you
    1. i was able to resolve the "missing node" i think you were right about cached browser, i chose rightclick reboot from the comfyui window, but nodes did not refresh, so after closing and restarting the red missing node disappeared.
    2. i will use [Chunk 1] , [Chunk 2] etc.. to space out my story, follow up on this, if i describe the details of the character at the top will it follow all the chunks? or do i have to re-add them? some chunks i forsee the character may go "off screen", how do i keep continuity?
    also what are the maximum chunks, can i create a whole story like 100 chunks?
    3. ok generate, then upscale afterwards, do you recommend/create any?

    at this point im hoarding workflows and going cross eyed, but there are a handful (like this one) which are just great!, keep up the awesome work!

    ukr8b3g201
    Author
    Aug 30, 2026

    @gigomikol 最後の追加質問への返答案です。そのままCivitaiへ投稿できます。

    You’re welcome—and yes, those are good follow-up questions. 😊

    1. Shared character description and off-screen continuity

    Text placed before the first [Chunk 1] header is treated as a global preamble. Continuum automatically includes it in every chunk prompt.

    For example:

    GLOBAL CONTINUITY:
    The same young woman appears throughout the story. She has long black hair, green eyes, a red jacket, black jeans, and silver earrings. Her face, clothing, age, and body proportions remain unchanged in every chunk.
    
    The location remains the same modern hotel, with warm evening lighting and dark wooden walls.
    
    [Chunk 1]
    The woman enters the hotel hallway. The camera follows her from behind.
    
    [Chunk 2]
    Without a cut, the same woman continues down the same hallway. She turns right and exits the frame while the camera remains in the corridor.
    
    [Chunk 3]
    The same woman re-enters from the right side, still wearing the same red jacket, black jeans, and silver earrings. She opens the next door.

    You therefore do not have to copy the complete character description into every chunk.

    However, for important characters and long sequences, I still recommend briefly repeating the most important details when the character returns:

    The same woman re-enters from the right, wearing the unchanged red jacket and silver earrings.

    This gives H3 both the global identity description and a local reminder.

    If the character remains off-screen for several chunks, identity drift becomes more likely because there is no visible character information in the recent video context. To reduce this:

    Connect a clear character image through Reference Image

    Keep the shared identity description above [Chunk 1]

    Explicitly say that the character is temporarily off-screen rather than replaced

    Restate her identity, clothing, and entry direction when she returns

    Avoid keeping an important character absent for too many consecutive chunks

    Reference conditioning and Continuum context improve consistency, but they cannot guarantee perfect identity over an unlimited story.

    2. Maximum number of chunks

    The current sampler supports 1–16 logical chunks, so 100 chunks cannot be generated as one Continuum run.

    Although 16 chunks are available, I would not recommend treating one run as a complete movie. Small differences in identity, lighting, camera position, contrast, and motion can accumulate as the sequence becomes longer.

    A more reliable structure is:

    Scene 1: 3–6 chunks
    Scene 2: 3–6 chunks
    Scene 3: 3–6 chunks

    Generate each scene separately, then use an accepted final frame from the previous scene as the next scene’s First Frame. Keep the same Reference Image and global character description, and assemble the completed scenes afterward in a video editor.

    That gives you natural scene boundaries and makes it much easier to reroll or repair only one part of the story.

    3. Post-upscaling recommendations

    For a beginner, I recommend keeping upscaling separate from the Continuum generation workflow:

    Generate near native H3 resolution
    → assemble the complete video with audio
    → upscale the finished video

    The easiest paid option is Topaz Video. It accepts a completed video directly and is simpler than processing frames manually.

    Inside ComfyUI, the lighter option is a conventional ESRGAN-style image upscaler using the built-in UpscaleModelLoader, followed by video reassembly. This is relatively easy and works on modest hardware, but independent frame upscaling may introduce some temporal shimmer.

    A more advanced video-aware option is SeedVR2 Video Upscaler. It can preserve temporal consistency better, but it is significantly heavier. Its documentation warns that even the 3B model may need around 18 GB of VRAM, so it is not my first recommendation for a novice or a 16 GB GPU.

    I may provide a separate optional post-upscale workflow rather than adding more complexity to the main Continuum workflow. For now, generate and approve the complete video first—then upscale only the final accepted result.

    And thank you again. “Hoarding workflows and going cross-eyed” is a very accurate description of learning ComfyUI. 😄

    hatt2Aug 30, 2026· 2 reactions
    CivitAI

    Hello,
    I have more questions. 😅

    I was reading the official minimax H3 documentation today for REF2VA.
    https://huggingface.co/MiniMaxAI/MiniMax-H3/blob/main/docs/VIDEO_PROMPT_WRITING_GUIDE_ref_en.md

    I was wondering if we should use the same format rules for Continuum? With their format, I'm not sure where to put [Chunk 1] [Chunk 2] etc


    subject_definitionsDefines referenced content and its reference labelssummarySummarizes the task type, target video, and main reference relationships

    retention_analysisDescribes how referenced content is preserved, transferred, or reused

    detailed_descriptionDescribes visuals, actions, shots, sound, and dialogue in playback order

    overall_soundscapeSummarizes ambience and physical sounds

    non_diegetic_musicDescribes background music audible only to the audience

    ukr8b3g201
    Author
    Aug 30, 2026· 3 reactions

    技術的ですが、結論は明確です。REF2VAを使う場合は公式の6セクション形式が推奨ですが、[Chunk N]はMiniMax形式ではなくContinuum側の振り分け記号です。

    Civitaiには、次のように返すと分かりやすいです。

    Yes—when you are actually using the Ref2VA model and Reference inputs, following MiniMax’s six-section Ref2VA format is a good idea. However, it is not required for every Continuum workflow. T2VA/I2VA/FL2VA can continue using the normal H3 prompt format.

    The important distinction is:

    subject_definitions, summary, retention_analysis, etc. are instructions for the H3 model.

    [Chunk 1], [Chunk 2], or [0-10s] are routing headers used only by Continuum.

    Continuum removes the Chunk header before sending the resolved prompt to H3.

    For two 10-second chunks, I recommend Prompt Format = Timeline and this structure:

    subject_definitions:
    <Subject 1> is the woman referenced from <Picture 1>, preserving her face, hair, clothing, and overall identity.
    
    [Chunk 1]
    summary:
    [reference generation] The first target segment introduces <Subject 1> and begins the action.
    
    retention_analysis:
    <Subject 1> (appears in [Shot 1]): fully_preserved - her identity, face, hairstyle, and clothing are retained.
    
    detailed_description:
    [Shot 1] A medium close-up introduces <Subject 1>...
    [Shot 2] At 00:05.000, the camera moves...
    
    overall_soundscape:
    The environmental ambience and physical sounds for the first segment...
    
    non_diegetic_music:
    A restrained electronic score begins and continues without resolving.
    
    [Chunk 2]
    summary:
    [reference generation] The second target segment continues the same subject, environment, action, and audiovisual direction.
    
    retention_analysis:
    <Subject 1> (appears in [Shot 1]): fully_preserved - the same identity, clothing, and visual characteristics continue.
    
    detailed_description:
    [Shot 1] The action continues immediately from the inherited ending state, without resetting the pose, camera direction, lighting, or environment...
    [Shot 2] At 00:05.000, the camera moves...
    
    overall_soundscape:
    The ambience and physical sounds continue naturally from the previous segment.
    
    non_diegetic_music:
    The same score continues across the segment boundary without restarting.

    Text before the first [Chunk 1] header is a global preamble. Continuum automatically places it at the beginning of every resolved chunk prompt, so shared subject_definitions are a good fit there.

    Each chunk should still contain its own:

    summary

    retention_analysis

    detailed_description

    overall_soundscape

    non_diegetic_music

    Within each chunk, restart the model-facing shot notation at [Shot 1]. The [Shot N] timestamps are local to that chunk, while [Chunk N] belongs to Continuum.

    A few important points:

    Put [Chunk 1] and [Chunk 2] on separate lines.

    Do not mix [Chunk N] sections with --- List separators.

    Define every configured chunk; otherwise Continuum may reuse the previous prompt as a fallback.

    If one identical full Ref2VA prompt should apply to every chunk, simply use Fixed mode—no Chunk headers are needed.

    Use [video continuation] in the Ref2VA summary only when an actual reference video is being continued. Continuum’s internal latent continuation alone does not automatically make the external Ref2VA task type video continuation.

    Also, if a reference image only defines a person’s identity, define it as a <Subject N> sourced from <Picture N>. Use a standalone <Picture N> entry mainly when that picture is a concrete First Frame, Last Frame, keyframe, or composition anchor.

    The six-section rules come from MiniMax’s official Full-Reference Mode guide. Continuum does not rewrite those sections; it only selects the appropriate complete prompt for each chunk and passes it through to H3.

    hatt2Aug 30, 2026

    ok thank you for the informative reply! I have a lot of studying to still do. I will be back with some test results soon ! 🙂