Basic Refmod creator for Minimax H3
Combine and encode your REF2VA audio/image/video references into a single .safetensor for use with MMH3 Director Workflows
Nodes: search for "Minimax H3 Refmod" in your manager/extensions
Description
FAQ
Comments (9)
no better creating page to creating minimaxh3 for free?
not sure I follow your question
@FirstPrinciples mean is not try creating page to generate video ai for minimaxh3
I’ve seen many people on Reddit recommending this LoRA type that bundles various reference materials. Does this mean that when creating such a LoRA, the source files must all be of the same type—such as bundling only images, only videos, or only audio (even if the specific formats differ, like mixing AVI and MP4 or PNG and JPG)—rather than bundling a folder containing a mix of different media types (like audio, video, and images all together) into a single LoRA?
Well refmods aren't really Lora's, although if that helps conceptually. Normally when you do REF2VA, you would add your image and audio references directly to the timeline, then the workflow encodes them as it goes. What the refmod does is combine all those images and audio into a single package that can be used for a reference in the timeline, and since they are pre-encoded not only does it save file size, but it saves processing time too since they dont need to be encoded at runtime.
Think of them like little character packages where (yeah, conceptually similar to a LoRA) you would have like a front, back, quarter, clothed, nude, expressions, voice sample etc packaged into a single .safetensors that you can then drop in your REF2VA timeline - that way instead of something like <image 1> <image 2> <image 3>, etc, you can just tie the <Subject> to <Refmod 1> which contains everything. The source files don't need to be the same type or even size, you can mix the characters voice samples with different image types - anything that you would normally drop into a REF2VA timeline you can drop into the refmod instead.
Hope that helps.
@FirstPrinciples Thanks for the reply. Yesterday, I tried packaging an image of a person and some video footage (two 7-second MP4 clips) into a single RMOD, and I noticed it was automatically named <VIDEO_1>. How do I specify in the prompt that H3 should use this RMOD? Should I simply write something like <video_1> walking towards the camera?
Do you also recommend using one dedicated RMOD per character or object? If I were to mix different subjects (e.g., a soccer star and a basketball star) into a single "sports star" RMOD, how would I phrase the prompt so that H3 can distinguish between the different assets contained within that mixed RMOD?
@wyxzddsjj919 Anytime you package in a video it will add everything to the video section of the timeline, but the images are still there for you to reference - in your example with a sports star with multiple outfits you could do something like:
<Subject 1> is a (describe the person like tall althetic male with x hair and eyes whatever) fully referenced in <Refmod 1>
<Subject 2> is a white basketball outfit with the number 7 on the jersey fully referenced in <Refmod 1>
<Subject 3> is a red and blue "Richmond" soccer jersey fully referenced in <Refmod 1>
Then the prompt detail would say something like:
At 00:00.000, <Subject 1> is fully clothed in <Subject 2>, he turns to the camera and says <d>[ENGLISH] But I love soccer too! </d>. At 00:04.000, <Subject 1> snaps his fingers and in a poof of white smoke, <Subject 2> is replaced by <Subject 3>
Obv your prompt would be better but I hope you get the idea - you can reference anything in the refmod or use different <Subjects> tied to different parts of the refmod to build your prompt
Is this limited to just 6 images? can you do it from a folder with a bunch of images?
The standard REF2VA limits apply:
Images - Max 9
Videos - Max 3, 2-15 seconds per, 15 seconds combined total
Audio - Max 3, 2-15 seconds per, 15 seconds combined total
Max total references 12
So you could do 9 images, 3 video or 6 images 3 video 3 audio... etc etc.
