Getting Started: Chunks, Prompts & Reference Images
3.9.1 — what changed (2026-10-04)
This patch fixes Second Pass conditioning, Take reuse and Review audio:
Second Pass: preserve audio-only/mixed keyframes and inherit the verified conditioning and Reference assignments of the actual Review output, including a single middle chunk.
Take reuse: include sampler closures and MODEL CFG/wrappers in compatibility checks to prevent reuse under different generation settings.
Review Driving Audio: cut the source PCM once at the physical group's natural time, with the same start position for Exact ON/OFF.
Audio resampling: prefer ComfyUI Core's standard API, with the legacy fallback for older Core versions. The Reference Encode Cache limitation for direct VAE weight changes is documented below.
V3.9 Performance Fix — October 1, 2026
Reduced unnecessary preparation time when using Fixed prompts without Reference Images. First Image is still supported. In matched tests, the earlier slowdown compared with V3.8 was no longer observed under these conditions.
Update Continuum from GitHub main, then fully restart ComfyUI. No workflow changes are required for this fix.
What’s New in V3.9
V3.9 introduces per-chunk Reference Images.

Reference Images 1–9 now keep fixed IDs @R1@R9), and the new Reference Images V3.9 node lets you choose which images are active for each chunk. All references are connected to the V3.9 Sampler through one Reference Images input.
For the supplied 2 × 10 s workflow, put each time header on its own line and describe only that chunk's action. For example, after enabling the matching image loaders and checking their chunk assignments:
[0-10s]
@R1 walks forward through a quiet hallway. The camera follows smoothly.
[10-20s]
@R2 enters from the right. The camera turns to follow @R2.
### V3.8X2 Compatibility
V3.8X2 is still supported as-is.
You can drag & drop your existing V3.8X2 workflow and continue using it normally with the V3.8 Sampler.
The new V3.9 Reference system uses different wiring and is not automatically applied to old workflows. To use V3.9 features, either load the included V3.9 workflow or build/rewire your own workflow with the V3.9 Sampler and Reference Images V3.9 node.
https://github.com/ukr8b3g-cmyk/ComfyUI-H3-Continuum/tree/main#v39-reference-images
What’s New in V3.8X2
V3.8X2 is a small update to V3.8X. The existing sampler, Review, Resume, Retry, Takes, and continuation behavior are unchanged.
Up to 9 Reference Images
Reference Images 4–9 can now be added through the new Reference Images helper. Using many or large reference images can use substantial VRAM and system RAM, so resizing source images first is recommended.
Built-in Decode Cache Helper
A Decode Cache Helper node is included for repeated decoding of unchanged latents. It can help with reuse and resume cases, but it does not speed up new Sampling and may not be useful for normal fresh generations.
Use the included V3.8X2 workflow.
https://github.com/ukr8b3g-cmyk/ComfyUI-H3-Continuum
What’s New in V3.8X
V3.8X is the new main line of H3 Continuum, focused on more reliable Review, Resume, Retry, and Continue workflows.
Current-chunk Review preview
Validated prefix reuse
Atomic crash-safe chunk saving
Reliable Resume, Retry, and Takes
No unnecessary regeneration of accepted chunks
Same 7 public nodes and sampler interface
Use the included V3.8X workflow.
V3.8.0 remains available for existing setups.

For a new Git installation:
cd ComfyUI/custom_nodes
git clone https://github.com/ukr8b3g-cmyk/ComfyUI-H3-Continuum.gitFor an existing Git checkout:
cd ComfyUI/custom_nodes/ComfyUI-H3-Continuum
git pull --ff-only origin mainRestart ComfyUI after cloning or pulling. If ComfyUI Manager installed the node, use its Update action instead of
What’s New in V3.8
Please refer to the GitHub repository for usage instructions.
https://github.com/ukr8b3g-cmyk/ComfyUI-H3-Continuum

V3.8 focuses on workflow control, recovery, and ease of use. It does not introduce a new image-quality enhancement; equivalent settings continue to use the established V3.7 Production sampling foundation.
Review Each Chunk — inspect each generated chunk before continuing.
Keep, Retry, or Finish — use the current chunk, generate it again, or finish all remaining chunks.
Resume & Takes — reopen saved progress and choose from previous Takes.
Partial Regeneration — regenerate from a selected chunk without rebuilding the accepted beginning.
Simplified Size Controls — choose First Image sizing or exact Manual dimensions.
Built-in Input Bypass — quickly enable or disable Image, Audio, and Video inputs.
Three Reference Audios — use up to three ordered audio references.
Simplified Public Surface — seven supported V3.8 Continuum nodes.
Spectrum / Turbo Workflow — the supplied Spectrum graph can also be configured for LightX2V Turbo.
Known limitation: Issue #13 remains open. V3.8 does not claim to improve or solve cumulative image-quality drift during long continuation.

What’s New in V3.7

V3.7 focuses on faster, more reliable high-resolution Second Pass refinement. The new Conditioning Adapter rebuilds First/Last images and continuation context at the target resolution, while RefineSchedule provides exact Full, Tail, Partial, and External refinement ranges.
Using Tail 6 instead of Tail 10 reduced Second Pass sampling from 100.27s to 64.18s in our reference test—about 36% faster—while preserving the original first-pass audio bit-exact.
Production defaults remain unchanged, and existing V3.6 workflows remain compatible. The new Still Image Guide is included as an Experimental feature because hard anchors may cause abrupt trajectory changes.
What’s New in V3.6.1
https://github.com/ukr8b3g-cmyk/ComfyUI-H3-Continuum

V3.6 introduces a new Masked AV Continuation backend for faster and cleaner long-form generation.
Instead of adding the previous chunk again as a separate Reference block, V3.6 keeps the finalized Video/Audio latent prefix directly inside the next target and samples only the new region.
Masked AV Continuation — new Standard backend
Less Attention overhead — FL2VA test reduced packed rows by 8.01%
Faster continuation Sampling — median Terminal Group sampling improved by 4.06%
Bit-exact Video + Audio prefix preservation
T2VA / I2VA / FL2VA supported
Reference Image + Reference Audio supported
Long Terminal Merge fully supported
Run Storage / Resume / Regenerate From supported
Compatibility mode keeps the previous V3.5 Reference Context backend
Chunk duration expanded to 4–30 seconds — 5–15 seconds remains recommended
V3.6.1 Hotfix
Improved mixed Timeline /
---prompt syntax handling and warningsPrevents unintended repeated chunk prompts
Non-Balanced Audio Continuity settings safely fall back to the compatible Reference Context path instead of stopping generation
Same sampling quality, less redundant continuation work, and full compatibility with existing V3.5 workflows.
H3 Continuum V3.5.2 — Stabilization & Optimization Update
V3.5.2 focuses on stability, efficiency, and cleanup rather than adding major new features. Repeated runs with the same Prompt/CLIP conditions now avoid unnecessary re-encoding, and Video Guide preprocessing uses substantially less temporary RAM on longer inputs.
Existing V3.5.x workflows and generation contracts remain unchanged. Sampling, Terminal Merge, Run Storage, Reference handling, and output behavior were preserved while the updated paths were validated with CPU and GPU A/B testing.
V3.5.2 is an internal stabilization and optimization update. Existing V3.5.1 workflows can be used as-is, with no changes required to nodes, connections, or workflow structure.

> v3.5.1 Update
This update improves the Second Pass workflow, reference handling, and seed behavior. It adds the new Conditioning Bridge V3.5 for external sampling workflows, optional Reference Audio conditioning.
Video reference inputs are also clarified as Video Guide Frames / Video Guide Size, while existing V3.4/V3.5 workflows and backend connections remain compatible.
Updating to V3.5 does not automatically replace V3.4 nodes in saved workflows. When updating an older workflow, replace not only the Sampler but also the assembler with H3 Continuum Assemble + Seam V3.5. Add H3 Continuum Hi-Res Fix V3.5 when Hi-Res Fix is required.
GitHub:
https://github.com/ukr8b3g-cmyk/ComfyUI-H3-Continuum
[ V3.5 ]
V3.5 introduces two major additions:
- Continuum-aware Second Pass / Hi-Res Fix
Refine externally processed H3 latents while preserving Continuum physical groups, prompts, seeds, ordering, and first-pass audio. An integrated one-node 2x Hi-Res Fix path is also included as an experimental feature.
- Low-memory Assemble + Seam V3.5
Adds Auto, RAM, and Disk-backed video-buffer modes. Disk-backed assembly significantly reduces system RAM/private-memory usage for long or high-resolution outputs while preserving Exact Duration, Seam, Terminal Merge, and audio behavior.
All V3.4 nodes remain available for saved-workflow compatibility. Existing V3.4 workflows continue to work unchanged.
The V3.5 release passed 430 automated tests and representative GPU acceptance tests.
Note: The integrated Hi-Res Fix remains experimental. Long 2x workflows can require substantial GPU VRAM.
#5 3 chunks x 15 seconds
#6 6 chunks x 15 seconds
W576xH576

[ v3.4 ]
Long-form MiniMax H3 video and audio generation for ComfyUI with chunked generation, persistent references, restartable runs, and user-controlled audio.
### What's new in v3.4
- Driving Audio: preserves the supplied audio as the final audio while guiding generation across chunks.
- Video Reference: provides persistent visual reference for identity, motion, framing, and scene appearance.
- Restartable chunks: reuse completed chunks with Run Storage and regenerate only the required part.
- Improved Core compatibility: unknown upstream or custom nodes are not rejected merely because they are not recognized by Continuum.
- Simpler stable interface: obsolete compatibility controls and experimental Timeline inputs are hidden from the V3.4 public workflow.
- Spectrum interoperability: Spectrum remains optional and can use the official H3 Continuum Interop API.
### Direction change from v3.3
V3.4 focuses on predictable reference workflows rather than experimental Timeline Video and timeline-audio generation.
Driving Audio preserves the original user-supplied audio. Video Reference provides persistent visual guidance without requiring exact frame-by-frame copying. Existing V3.3 workflows remain available through legacy compatibility paths.
### Updating
For an existing Git installation:
git pull --ff-only origin mainV3.4 input connection patterns
V3.4 separates the visual reference input from the driving-audio input. Choose the connection pattern that matches your source material.
1. Audio only
Connect Load Audio to driving_audio. Use this when an existing song, dialogue track, or sound effect should remain the final audio. A Video Reference is not required.
2. Video with its own audio
Connect Load Video (Upload) IMAGE to Video Reference. If the uploaded video contains the audio you want to preserve, connect its AUDIO output to driving_audio as well.
3. Video and audio from separate sources
Connect Load Video (Upload) IMAGE to Video Reference, then connect a separate Load Audio node to driving_audio. Use this when the visual reference video and the final audio source are different files.
Both inputs are optional. Connect Video Reference when visual guidance is needed, and connect driving_audio when the supplied audio should be preserved in the final output.
Video Reference frame rate
Use a 24 fps source for Video Reference. Load Video (Upload) may accept files recorded at 25 fps or another frame rate, but acceptance alone does not guarantee correct temporal alignment with H3. For a non-24 fps source, set force_rate to 24 in Load Video (Upload), or convert the file to 24 fps before loading it. If the source is already 24 fps, leave force_rate at its default and do not resample it.
Current validation status
[ v3.3 ]
V3.3 adds Timeline Video conditioning for long-form MiniMax H3 generation. A reference video can now be processed in chunk-local time slices, allowing motion and scene continuity to be carried across multiple 5-second chunks while keeping the reference resolution independent from the output resolution. The Efficient 0.4 MP mode helps reduce memory usage and processing time.
Video assembly has also been improved. Auto seam handling analyzes chunk boundaries and applies guarded corrections for transient flicker, micro-flash, exposure, and color differences. This helps produce more natural transitions between generated chunks without changing the original sampling process.
Existing V3.2.4 workflows remain available as Legacy nodes for compatibility.
[ v3.24 ]
Generate longer native MiniMax H3 video and audio sequences in ComfyUI.
H3 Continuum is a ComfyUI custom node that generates a longer sequence as connected chunks and assembles them into one continuous video.
```text
3 × 5-second chunks → 15-second video
6 × 5-second chunks → 30-second video
The previous video and audio latent context is passed into each continuation chunk. This is not a simple video concatenation workflow.
Main purpose: longer MiniMax H3 generation, not faster generation.

Easy Installation
H3 Continuum can be installed directly from ComfyUI Manager.
Open ComfyUI Manager
Search for H3 Continuum or Continuum
Select Install
Restart ComfyUI
Load one of the included sample workflows

Manual installation and the latest documentation are available on GitHub:
GitHub:
https://github.com/ukr8b3g-cmyk/ComfyUI-H3-Continuum
What It Does
H3 Continuum divides a longer generation into manageable chunks.
MiniMax H3 Model
↓
H3 Continuum Sampler
↓
ComfyUI Core Video / Audio VAE Decode
↓
H3 Continuum Assemble
↓
Final videoEach continuation chunk receives latent context from the preceding chunk. Overlapping context is removed during assembly, and the final frame and audio counts are aligned to the requested duration.
Main Features
Connected long-form MiniMax H3 generation
Native video and audio latent continuation
Fixed, List, and Timeline prompt formats
Automatic prompt-format detection
T2VA, I2VA, FL2VA, Last Frame and Reference workflows
Up to three Reference Images
Reference Audio conditioning
First Frame and Last Frame conditioning
Configurable continuity context
Run Storage and automatic resume
Partial regeneration from a selected chunk
Optional Spectrum interoperability
Standard and Turbo sample workflows
ComfyUI Core VAE Decode compatibility

Included Sample Workflows
Two example workflows are provided.
Standard Workflow
Recommended when output quality and temporal consistency are the priority.
Standard MiniMax H3 sampling
Spectrum can be enabled
Suitable for quality-focused generation
Reference Image and Reference Audio supported
RTX upscaling can be enabled when required

Turbo Workflow
Recommended for faster tests and iteration.
LightX2V MiniMax H3 Turbo LoRA
8-step example configuration
Spectrum is bypassed by default
Faster than the standard workflow in tested configurations
Some loss of facial detail or additional artifacts may occur

Turbo LoRA models:
https://huggingface.co/lightx2v/Minimax-h3-Turbo/tree/main
MiniMax H3 models and documentation:
https://huggingface.co/MiniMaxAI/MiniMax-H3
Models and LoRAs are not included with this custom node.
Reference + Continuation
Reference Images remain available across all generated chunks.
A typical setup is:
Picture 1 → face and identity
Picture 2 → full-body appearance and clothing
Picture 3 → environment or an additional visual reference
Audio 1 → vocal, music or audio-performance referenceRef2VA is the reference-specialized checkpoint and is generally the first choice for stronger reference fidelity.
FL2VA with Reference conditioning is also allowed. H3 Continuum does not automatically replace or switch the connected model.
Spectrum Integration
Spectrum is optional. H3 Continuum also works without it.
With a compatible Spectrum release, H3 Continuum sends a continuation signal only when generating later chunks.
Chunk 1 → normal Spectrum sampling
Chunk 2+ → Continuum Actual Prefix 2This allows Spectrum to coordinate its spectral forecasting with the continuation context instead of treating every chunk as an unrelated generation.
Benefits include:
Automatic identification of continuation chunks
Actual Prefix applied only where required
No manual prefix switching between chunks
Reduced risk of duplicated prefix processing
Compatibility with standard ComfyUI workflow execution
Spectrum remains an approximate accelerator. Motion, anatomy, audio and detail can differ from a non-Spectrum result, so quality comparisons should use the same prompt and seed.
Spectrum:
https://github.com/xmarre/ComfyUI-Spectrum-MiniMax-H3
Run Storage and Resume
Enable Save + Auto Resume to preserve completed raw video and audio chunks.
If a generation is interrupted, H3 Continuum can reuse compatible saved chunks and continue from the first missing chunk.
It can also regenerate from a selected chunk while preserving the compatible prefix.
Chunk 1–3 completed
↓
Generation interrupted
↓
Queue the workflow again
↓
Chunks 1–3 reused
↓
Generation continues from Chunk 4Run Storage verifies the sampling contract, model route, references, resolution and saved chunk files before reuse.

Prompt Formats
Fixed
One prompt is used for every chunk.
List
Separate prompts are divided with:
---Timeline
[0-5s]
First scene description
[5-10s]
Second scene description
[10-15s]
Third scene descriptionPrompt Format = Auto detects the appropriate format automatically.
Incomplete timeline coverage produces diagnostics and safe fallback behavior rather than unnecessarily stopping every generation. Structurally unusable input is still reported as an error.
Tested Configuration
The current Windows implementation has been tested with:
GPU NVIDIA RTX 5060 Ti 16GB
ComfyUI MiniMax H3-compatible Core build
Chunk Duration 5 seconds
Typical Length 3 or 6 chunks
Continuity Balanced 22 frames
Standard Sampling RES Multistep
Spectrum Interop Actual Prefix 2The node is not limited to RTX 50-series GPUs. Actual compatibility, generation speed and usable resolution depend on the MiniMax H3 model, GPU memory, ComfyUI configuration and installed acceleration nodes.
RTX 4060 and other configurations have not been formally validated by this project.
Frequently Asked Questions
Is this only a workflow?
No. H3 Continuum is a ComfyUI custom node package. The included workflows are ready-to-use examples.
Does it generate one native 30-second sample?
No. It generates connected chunks and assembles them into one longer output while carrying video and audio latent context forward.
Does it make MiniMax H3 faster?
Speed is not the primary purpose. H3 Continuum is designed for longer generation. Spectrum and Turbo LoRAs can reduce generation time in some configurations.
Is Spectrum required?
No. It is an optional acceleration and interoperability path.
Can I use the Turbo LoRA?
Yes. A Turbo sample workflow is provided. Spectrum is bypassed by default in that workflow because combining both can change quality or introduce artifacts.
Which model should I use for Reference Images?
Ref2VA is the reference-specialized option. FL2VA with Reference conditioning is also allowed, but reference fidelity may differ.
Are the models included?
No. MiniMax H3 checkpoints, text encoders, VAEs, Turbo LoRAs and optional acceleration nodes must be installed separately.
Can an interrupted generation be resumed?
Yes. Enable Run Storage before generation. Compatible completed chunks can then be reused.
Can I regenerate only the later part?
Yes. Run Storage supports regeneration from a selected chunk while retaining a compatible earlier prefix.
Are chunk boundaries always invisible?
No generative continuation system can guarantee a completely invisible boundary. H3 Continuum preserves latent context and removes duplicated overlap, but difficult motion, lighting changes and large prompt transitions can still produce flicker or visual changes.
Does Reference Audio guarantee exact lip synchronization?
Reference Audio conditions MiniMax H3’s native joint video/audio generation. It can guide vocals, rhythm, expression and mouth movement, but it does not guarantee sample-identical audio reproduction or frame-perfect lip synchronization in every generation.
Does it support audio continuity?
Yes. Video and audio latent context are carried together. The assembler also provides an optional Audio Seam mode for boundary-local audio correction.
Is RTX 5090 required?
No. Development and runtime validation were performed on an RTX 5060 Ti 16GB. Lower-memory configurations may require reduced resolution, offloading or other ComfyUI memory optimizations.
What license is used?
H3 Continuum is released under the MIT License.
Links
Custom Nodes
The included Standard and Turbo workflows use the following custom nodes.
- H3 Continuum
https://github.com/ukr8b3g-cmyk/ComfyUI-H3-Continuum
- rgthree-comfy
https://github.com/rgthree/rgthree-comfy
- ComfyUI-Easy-Use
https://github.com/yolain/ComfyUI-Easy-Use
- ComfyUI-KJNodes
https://github.com/kijai/ComfyUI-KJNodes
- ComfyUI-Spectrum-MiniMax-H3
https://github.com/xmarre/ComfyUI-Spectrum-MiniMax-H3
Spectrum are optional generation paths, but installing all listed custom nodes allows the included workflows to load without missing-node warnings.
Models
- MiniMax H3
https://huggingface.co/MiniMaxAI/MiniMax-H3
- LightX2V MiniMax H3 Turbo LoRA
https://huggingface.co/lightx2v/Minimax-h3-Turbo/tree/main
Models and LoRAs are not included in the workflow ZIP.
Main Links
- GitHub and documentation
https://github.com/ukr8b3g-cmyk/ComfyUI-H3-Continuum
- Install from ComfyUI Manager
Search for H3 Continuum
Description
H3 Continuum V3.5.2 — Stabilization & Optimization Update
V3.5.2 focuses on stability, efficiency, and cleanup rather than adding major new features. Repeated runs with the same Prompt/CLIP conditions now avoid unnecessary re-encoding, and Video Guide preprocessing uses substantially less temporary RAM on longer inputs.
Existing V3.5.x workflows and generation contracts remain unchanged. Sampling, Terminal Merge, Run Storage, Reference handling, and output behavior were preserved while the updated paths were validated with CPU and GPU A/B testing.
V3.5.2 is an internal stabilization and optimization update. Existing V3.5.1 workflows can be used as-is, with no changes required to nodes, connections, or workflow structure.
FAQ
Comments (18)
Thanks for the new 3.5 Version, excited to try it out.
After some quick tests, the performance of the Continuum Sampler seems to have massively improved (close to double speed for 5s videos for me).
This causes some Questions for me regarding the 2nd pass node and execution times, maybe you can help me understand.
For my setup, a 5s video with the 2x upscaler at 10 passes enabled at 0.6mp took 483s
If I switch the upscaler to from 2 to 1 but raise the mp to 1.5, it took 318s.
If I understand correctly, the 2nd option should be the better quality since it is rendered natively at a better quality with all else being the same, correct?
Secondly, if i am not using the upscale function on the 2nd pass node, is there any difference to adding the extra steps to the inital sampling?
As always, thank you for your hard work.
Thanks for testing it — those timings actually make sense with how the V3.5 Hi-Res Fix works.
One important detail is that scale_by is a linear scale. So if your First Pass is around 0.6 MP and you use 2x, the Second Pass is working at roughly 4x the pixel area, around 2.4 MP after aspect-ratio and size rounding.
With your second setup, 1.5 MP + scale_by = 1, there is no spatial upscale. V3.5 detects that the target size is unchanged, skips the Pixel/VAE resize round-trip, and performs a same-resolution Second Pass directly at the First Pass resolution.
So these are actually two rather different workloads:
0.6 MP → 2x: lower-resolution First Pass, then a much larger high-resolution Second Pass.
1.5 MP → 1x: higher-resolution First Pass, followed by same-resolution refinement.
I would not say that the second one is automatically higher quality. A higher-resolution First Pass can give H3 more spatial information from the beginning and may be a very good quality/time tradeoff. On the other hand, the 0.6 MP → 2x path produces a larger final image and gives H3 another sampling stage at that larger resolution, so it can potentially add or reconstruct fine detail differently.
For a proper quality comparison, I would compare the two approaches at approximately the same final resolution, rather than only comparing the number of Second Pass steps.
I should also add an important limitation from my own testing: my test GPU is an RTX 5060 Ti 16 GB. With longer Continuum runs such as 3 × 5 seconds, a 0.6 MP First Pass followed by a 2x Hi-Res Fix can exceed the available VRAM and OOM, especially on the larger Terminal Merge group.
Because of that, my V3.5 Hi-Res Fix testing so far has mainly focused on confirming that the workflow, conditioning, Second Pass, references, and final assembly operate correctly. I do not yet have enough controlled long-form data to make a strong claim about which of these two strategies consistently gives the best visual quality. Your timing and quality comparisons are therefore genuinely useful data for me as well.
For your second question: a 1x Second Pass is not the same as simply adding more steps to the First Pass.
Extra First Pass steps continue the original denoising trajectory. A 1x Second Pass instead starts from the already generated latent and performs a new low-denoise refinement pass with its own seed/SIGMAS while reusing the captured Continuum conditioning context.
So at scale_by = 1, I would think of it as a same-resolution refinement pass, not an upscaler.
If the First Pass is already good and you only want a little more convergence, adding a few First Pass steps may be simpler and faster. If you specifically want a separate refinement stage, want to rework details after the initial generation, or are using an external latent processor/upscaler, the Second Pass becomes more useful.
Also, thanks for reporting the V3.5 speed improvement and the actual execution times. Those comparisons are very helpful, especially since high-resolution long-form testing is VRAM-limited on my current hardware.
@ukr8b3g201 Thank you for the detailed explanation.
I have always produced at a higher inital resolution since it felt like scene consistency regarding movements/extra limbs etc got better then, though might be anecdotal with how much variance there is between seeds.
I also did a test with an idea of slightly combining 3.4 and 3.5.
I took my 3.4 relsolution of 1mp used the 2nd pass for refinement (10 steps at 1x) and then added the RTX Superscaler back in at the end for 2x upscale. I am aware from your explanation that the 2nd pass doesnt run at the same resolution, but overall I saw very little difference in quality to 0.6mp with 2x, maybe slightly better for the 1mp version actually.
Time is where the big difference was, for 2x5s the 0.6 with 2x took about 23.5 minutes, whereas the 1mp version took about 10 minutes.
And a second thing I noticed, with maybe a future improvement idea.
When you start the workflow with a fixed seed, you can run the workflow, see if you like the result and then turn on the 2nd pass and it will start there with the generated video. Super helpful for time savings.
Sadly, if you start with randomize and then switch it to fixed, after it generated, this doesnt work and it generates a new video with the same seed.
I dont know how technically challenging it would be to allow that to work, but that would be awesome.
@Xarfai Thanks — that is a very useful observation, and I agree this is worth fixing.
What is happening is mostly related to ComfyUI's standard seed control rather than the Second Pass itself. When the First Pass is generated with Randomize, ComfyUI changes the seed widget to the next random seed after the queue. So when you then switch it to Fixed, the value shown is already the next seed, not necessarily the seed that produced the video you just accepted.
That makes Continuum see the First Pass input as changed, so it correctly invalidates the cached result and generates it again.
I plan to improve this behavior in a V3.5.x update. The goal would be to preserve or recover the last actually executed First Pass seed, so the workflow can behave like this:
Randomize → generate a result you like → enable Second Pass → reuse that exact First Pass → run only the refinement.
I want to implement it carefully so that Continuum never reuses an old First Pass when the user intentionally changes the seed.
As an immediate workaround, there is also a ComfyUI setting worth trying: change Widget Control Mode from After Generate to Before Generate. With seed randomization happening before the queue instead of after it, the seed displayed after generation should remain the one that was actually used. You can then switch it to Fixed before enabling the Second Pass, which should allow the cached First Pass to be reused.
And your 1 MP → 1x Second Pass → RTX 2x result is interesting as well. Roughly 10 minutes versus 23.5 minutes for 2×5s is a very large difference, especially if the visual quality is comparable or slightly better. That looks like a very practical workflow for people who have enough VRAM for a higher-resolution First Pass but do not want to run the H3 Second Pass at the full 2x resolution.
This is fantastic work, thank you so much for sharing and for all the improvements you've made ! !
thanks for the update. will test ASAP.
regarding to the previous issues, here is a reply from a WF creator that uses the same nodes. maybe it helps or inspires you.
have a great day.
Your other issues were fixed in my other workflows that don't use Continuum.
The crash isn't VRAM and isn't Spectrum: it's the H3-Continuum assembler (a different pack from this one) trying to allocate one 9–11 GB block of system RAM to hold your entire 4×15s chain at once (assembly.py line 132). On a loaded Windows box that single contiguous allocation fails even with plenty of total RAM.
Ways forward, pick one:
Use this pack's chain instead: the H3_Seamless_Chain / Extend_Take workflows here assemble the master by streaming each shot to disk — peak system RAM is one shot, not the whole take (low_ram_master, on by default). Your 4×15s runs in a few GB of RAM.
Stay on Continuum: shorten the chain per assembly, or raise your Windows page file so the big allocation can land — but that's working around a design that holds everything in memory.
Two side notes from your logs, both worth fixing regardless:
Ref2VA requires ffmpeg/ffprobe on PATH; ffprobe=None — your audio path is missing ffprobe. Install a full ffmpeg build and add its folder to PATH, or audio features will misbehave silently.
The minimax_h3_..._pruned_int8_convrot checkpoint is a third-party pruned cut — we've had a report in another thread of degraded audio from pruned files. If you hear mumbling or artifacts, test against the Q5_1 GGUF from this listing before blaming settings.
The numba/NumPy warning is unrelated (WAS suite) and harmless here.
Thanks for sharing this.
That diagnosis is correct for the older Continuum assembler. We had actually anticipated this long-chain RAM issue already, and V3.5 introduced the low-memory Auto / RAM / Disk-backed assembler to avoid requiring one huge system-RAM allocation.
So it’s interesting to see another workflow creator independently arrive at essentially the same bottleneck and solution direction.
The ffprobe and pruned-model notes are useful too — I’ll keep those in mind for audio-related reports.
Thanks again for bringing this over.
@ukr8b3g201 you're most welcome
Dam this is good. Even at low MP it still produces clear images. Thanks
Thanks! Really glad to hear it. H3 seems to hold up surprisingly well even at lower MP settings, which is especially useful for longer Continuum runs.
Youre the best with those insanely quick iterations.
Thanks! I’ve been trying to keep each iteration small and focused so I can test and improve things quickly.
How do we prompt for the reference background? I uploaded the screenshot below. I couldn't get the correct background to appear in the video generation. 🤔
Any suggestion would be much appreciated.
is this correct?
<Picture 3>
<Picture 4>
<Picture 5>
---------------------
<Picture 1> reference for the woman clothing, hair, facial features
<Picture 3> weak_reference for location, environment, props, architecture, lighting, hue, light exposure, contrast and brightness
<Picture 4> weak_reference for location, environment, props, architecture, lighting, hue, light exposure, contrast and brightness
<Picture 5> weak_reference for location, environment, props, architecture, lighting, hue, light exposure, contrast and brightness
[Chunk 1]
Start the scene with a side shot view of the woman's entire horizontal body, flying through the air towards the right side of the camera, her arms are extended forward, her red cape is flapping in the wind. She is moving at blistering fast speeds, <Picture 3> weak_reference, behind her the building and hallways are a blur as she flies by at intense speeds
[Chunk 2]
tilt down, behind her back medium shot view, she is placed on left side of camera in entire clip. She flies into the screen. her red skirt, cape flaps violenting as she flies with her hands extended forwards. The interior bus <Picture 4> around her is a blur as she flies at rapid speeds.
[Chunk 3]
close front view of her face and chest, her arms are reaching forward into the screen balled up as a fist. She has a determined smile on her face. <Picture 5> The abandoned building is a blur as she flies forwards. Near the end of the video she flies past the camera and the outside building comes into focus revealing a large empty dead city, with crumbling buildings.
Sounds of strong wind in the background.
Your Picture numbering is correct if both First Frame and Last Frame are connected: Picture 1 = First Frame, Picture 2 = Last Frame, and Reference Images 1–3 become Pictures 3–5.
One important point: weak_reference is not a special MiniMax H3 or Continuum weighting keyword, so I would not rely on it. Instead, explicitly describe the role of each image.
For example:
"<Picture 3> is the environment reference for this scene. Preserve its architecture, hallway layout, materials, lighting, color palette and atmosphere. Use <Picture 3> for the environment only, not for the woman's appearance."
Then use Picture 4 in Chunk 2 and Picture 5 in Chunk 3 in the same way.
I would also avoid listing <Picture 3> <Picture 4> <Picture 5> by themselves at the top. Keep the woman/identity reference global, but tell each chunk clearly which background picture should control that scene.
Because all connected Reference Images remain available across the Continuum chunks, explicitly assigning one background reference to each chunk usually gives the model a clearer instruction.
@ukr8b3g201 ok thank you for the explanation. I will try to figure this out again !
Ok I'm back with more feedback. 😁
I found another bug.
I noticed this will force [Chunk 2] to loop for the rest of the video. The workflow will ignore Chunk 3, Chunk 4, Chunk 5
[Chunk 1]
Some text
[Chunk 2]
Some text
[Chunk 3]
Some text
[Chunk 4]
Some text
[Chunk 5]
Some text---
Some text
---
Some text
---
Some text
---
Some text
---
Some textCould we get a red error tellling us this is causing problems for the generations? Currently the workflow will silently fail, generate a movie that will be incorrect.
Thanks for reporting this. I found a prompt parser issue that can cause this behavior.
The problem occurs when Timeline syntax such as [Chunk N] is mixed with List separators (---). In some cases, later chunks are not recognized as separate Timeline sections, so the previous chunk prompt can be reused for the remaining chunks.
I’ve now fixed this on the GitHub main branch. The parser will:
detect mixed Timeline/List syntax,
warn with H3C-P105,
ignore standalone --- separators when Timeline mode is active,
clearly report when a previous prompt is reused for a missing chunk.
It will no longer silently include the --- separators in the Timeline prompt.
For now, please use either [Chunk N] sections or --- separators, not both.
The fix is currently on GitHub main and is planned for the V3.6.1 maintenance release. Thanks again for catching this.
@ukr8b3g201 Thanks for the continued support !






