The Infinitizer v2 is done. Due to overwhelming demand from that one person, I've added optional color correction & brightness/contrast control to the loop seam, addressed some bugs from v1 and cleaned up/streamlined the graph- I made subgraphs for self-contained modules with a much easier to understand layout. Copious notes abound, I know I have a tendency to plaster notes everywhere, but here it is warranted, as understanding how each stage works and what can go wrong is critical. This process has idiosyncrasies that are mostly consistent for workflows of the looping kind, but there are key differences from the usual flf setup. So I tried to cover everything. But there is no substitute for the fafo.
I'm not sure how this new scheme with multiple files works... if there is a mandatory dropdown, the zip file is only needed if you want to use autoprompting nodes.
Please use a ref2va model, not an flf2va model. The context node is designed for reference. You can certainly use a hybrid if you like, and flf may work too, but please don't.
Turbo LoRAs can kill this process. I can't emphasize this enough. There is a turbo option in here, and it might work sometimes, but mostly you are just asking for trouble. You only need to do a few seconds here, so running at least 20 steps is not going to kill you. So if you're running with turbo and it's not working at all, turn it off.
There are lots of combinations of post-sampler modules you can play with if there is trouble at the seam, you can re-run post without having to sample again if you keep your seed fixed, which you should always do with this concept anyway. If you have fast motion, for example, a 4 frame seam interpolation can introduces an ugly morph, so you can play with lowering the window or bypassing it entirely. Only takes a few seconds to redo that part.
Once you've familiarized yourself with the flow and know what to expect you can crank out loops pretty easily. But due to the inherent black-magic nature of diffusion, it's always best to watch your previews live. If you read the notes you will know almost immediately if a run is not going to end well and you can kill it after the first few steps and turn some knobs. This makes ditching turbo a lot more palatable, for those of you who are addicted to turbo and have forgotten how much it cramps your style. You don't have to wait for all 20 steps unless things are going swimmingly.
It's not critical to this workflow, but I've added Qwen prompt writing, as I'll be updating my continuation workflow with it ( it's absolutely killing it there...writes all seven prompts, so awesome) so I figured I would throw it in here as well. It uses a special instruction set just for this workflow, I'll paste that at the bottom of the description. You can use this with any LLM that can look at stuff. Or generic nodes, as abliteration is really not required here like it is in the other workflow.
I have created a template workflow for the Auto Prompting.
Please see that page to download the latest patches- I will remove the archive from this post when that posts.
Current custom nodes in the workflow. Some can probably be substituted, some are essential to the core process:
Description stuff prior to v2, anything mentioned above obviates whatever is down here:
Ok, v2 is what I really intended this to be. I was overeager to post something, what for to proselytize for our blessed H3 and its astonishment capabilities, so v1 is kind of blah, doesn't really show off what the model is capable of. v2 takes an input video and makes it infinite. That's it. It can make your three second WAN BJs watchable. Maybe. I don't know why I'm so incredibly neurotic about infinite loops. No on seems to care. But I do. So harken back to the days of yore, the animatediff days, when you could watch a video without vomiting and/or infarct of a random organ.
This is REF model. You can try a hybrid if you like, b25-49 works pretty well. But my default is basic REF, whatever quant you like, doesn't seem to matter. Uses a node from the same pack I used in my Infinite Extension WF, H3 Masked AV Bridge this time. Default is 39 frames of context, from both sides of the input. Sampling makes the bridge, then some minor tweaking to interpolate a few frames (the joint is as good as I've ever seen, but it still needs a nudge at the seam, as this works differently than straight FFLF), then it's stitched. DaSiWa RTX upscaler node and interpolation (whole-output interpolation, not the aforementioned part) are optional before the final combine. Uses the new rife setup. Makes sure everything is up to date. Update to next month's comfy if you can, that might hold you over for a few days.
There are a bunch of optional tweaks that I experimented with, I've left them in for now, best left disabled (group bypassers switch everything). Copious notes in every stage.
VERY little prompting is needed, I'd start with the bare minimum. The video loader can sometimes throw an error if you try to load a clip without audio- I put a reroute next to it in case you do need to manually unplug. There is also an option to silence the output if you want silence. Lots of stuff, but it's really, really simple. You should be able to crank stuff out really fast. I'll post more examples as I make them. Would be nice if you cited me if it works for you, I'm interested to know if it does. I've been working on various incarnations of this concept for years. It can be exceedingly frustrating at times. I've never been able to find workflows that do the things that I'm actually interested in. Not trying to craft absurd AIOs that promise everything and generate nothing. Trying to make tools that do specific things well. Most fail. I'm very pleased with this one though.
At long last. The LF that was promised is here. Praise the overlord. It's actually the frame you give it. No junk, no trash to fix. What it is supposed to be. So there you go. H3 solved the Mystery of the Final Frame.
This does the same thing all my previous FLF workflows attempt to do. Go from point A to point B without exploding. That's it. I am posting an example comp with NO corrections applied to the cuts in post. Nothing. One clip abuts another, the FULL output, nothing trimmed, no duplicated hold frames with dissolves, no morphs, none of the crap I always have to do. If you do this kind of work you will understand why this is a big deal.
Simple modular auto-replace setup so you don't have to constantly retype the required parts (per the official guide). So just describe the first frame only and any actions/movements necessary to get to the second. Below that, the camera movements needed, if any. Then same as always with H3, the two sound prompts.
This version is for HQ image inputs, there are optional interpolation, sharpen and upscale stages but it's mainly fucking around with what the model can do with very little. Experimentage. The latter uses https://github.com/darksidewalker/ComfyUI-DaSiWa-Nodes to run RTX super ultra megascale whatever. It's a nice pack, try it. It does real things, actually useful things.
This model permits users to:
Use the model without crediting the creator
Sell images they generate
Run on services that generate for money
Run on Civitai
Share merges using this model
Sell this model or merges using this model
Have different permissions when sharing merges
This means you can do whatever the fuck you want to do with it and I do not have the right to complain about it, no matter your intent was. Because I checked the slop waiver, acknowledging that this is all perverse, pointless garbage. No matter how smart we may think ourselves.
Useful generic prompt for Infinitization - use MM tag for insertion, or remove.
[Shot 1] Generate only the missing interval between the protected end of the first excerpt and the protected beginning of the second excerpt to form an unbroken continuous transition. Preserve the existing environment, natural lighting, color temperature, optical depth of field, and authentic physical motion throughout the transition. ##MM## Continue every visible subject, fluid dynamic, and background element smoothly along its current trajectory and natural pace. Keep the camera's existing vantage, focal length, and gentle physical drift completely continuous, avoiding any cut, speed ramp, freeze, duplicated element, or sudden composition change, and converge early enough that positions, focus, lighting, and movement settle exactly into the protected second endpoint.
Instructions for auto-prompting context-dependent infinite loop, if you want to use them elsewhere. The instructor only sees start & end frames, not all 78 context frames. I add a few brief sentences in the user prompt to nudge it in the right direction, but it does great when fed only the images:
You are an expert technical director and multimodal prompt engineer for generative video interpolation and seamless loop creation using MiniMax H3.\n\nYou will be provided with two images:\n- Image 1: The starting anchor frame (the end state of the original clip).\n- Image 2: The destination anchor frame (the beginning state of the original clip to complete an unbroken infinite loop).\n\nYour task is to analyze both boundary frames and write an unbroken, highly specific transition prompt adhering strictly to the three-tier MiniMax H3 schema:\n\n1. integrated_multimodal_description:\n- Begin immediately with '[Shot 1] Generate only the missing interval between the protected end of the first excerpt and the protected beginning of the second excerpt to form an unbroken continuous transition.'\n- Identify the shared environment, subjects, lighting conditions, color palette, camera elevation, focal depth, and texture between both frames.\n- Direct the subject trajectory, posture, gait, rotation, and environmental dynamics (e.g., fluid flow, wind, particles, ambient activity) so that elements visible in Image 1 evolve naturally and converge precisely into their corresponding positions, groupings, and states in Image 2.\n- Direct camera mechanics: enforce that the existing camera trajectory, height, drift, or gentle pan continues physically uninterrupted without any sudden composition snaps, speed ramps, freezes, or jump cuts, settling cleanly into the Image 2 vantage.\n- Include negative continuity guidance naturally within the prose: explicitly forbid cuts, speed ramps, duplicated subjects, ghosting, temporal freezes, or sudden lighting changes, ensuring movement converges early enough to lock seamlessly into the destination endpoint.\n\n2. overall_soundscape:\n- Detail the continuous, diegetic ambient audio and location Foley connecting Image 1 directly into Image 2 (e.g., continuous wind, water movement, mechanical hum, rustling, footfalls).\n- Match acoustic level, spatial characteristics, and room acoustics at both protected endpoints without adding speech, sudden volume spikes, or transients.\n\n3. non_diegetic_music:\n- Output 'N/A' unless explicit non-diegetic instrumentation is required.\n\nOUTPUT FORMAT RULES:\n- Output ONLY the formatted prompt beginning immediately with 'integrated_multimodal_description:'.\n- Do NOT output markdown code blocks, backticks, preambles, introductory greetings, or conversational commentary.\n- Use the exact section headers with no extra spaces.
Description
Makes a source video infinite.
REF model. Uses this context node. Same pack I use in my continuation workflow.