A MiniMax H3 studio workflow for text-to-video, image-to-video, and reference-to-video. One fused turbo model, 8 steps, and one video file per run.
Pick a mode and queue.
- T2V: prompt only
- I2V: unmute First Frame. Turn on Use Last Frame if you also want an end frame
- REF2V: add pictures and tag them <Picture 1> and <Picture 2>. For a clip, unmute Reference Video and Unpack Reference Video together and tag <Video 1>
Grey nodes are optional and stay muted until you have a file for them. Duration is 5–15 seconds. Stay at or under 0.98 MP (1344×768). 0.8 MP is the faster everyday size.
SeedVR2 and X2 Detail are toggles. Leave both off for the raw generation. Turning either one on still saves a single file.
Model links and the custom nodes used here are listed inside the workflow. Tested and built for a 16 GB card.
If you're going to ask why 0.98 MP and not 1.0 MP, honestly i have no idea... but if for some reason you are interested in finding out, here's the link: https://platform.minimax.io/docs/guides/local-deploy-h3
Description
Initial release.
FAQ
Comments (5)
Note: if you want one single model that supports every format this workflow uses, use the one I selected. The link is inside the workflow.
I built this because most MiniMax H3 workflows are overcomplicated. This one is meant to stay compact enough that a beginner can follow it.... and most importantly, it's 3 workflows into one!
Each video I posted has a comment with the resolution I used (in MP) and credit for the source image.
Hoping you may be able to help with pointing me in the right direction. If I wanted to add in Faceswap nodes using Reactor, where would that need to go? I've tried a couple of places, and while the log window looked like it was performing the swap, the final video didn't have it. I'll also keep trying to figure it out.
Is Reactor still the nodes to use for faceswaps or is there now something else recommended?
Cheers.
EDIT: Sorted it with a little help from AI
Glad you got it sorted. I do not use Reactor in this workflow. The usual miss is that the swapped frames are not the ones feeding Video Combine, so the log shows a swap and the MP4 does not. H3 will also overwrite a swap if Reactor sits before the sampler. For a face change here, put the clip in Reference Video, the identity still in Reference 1, and prompt it as an edit: face identity from Picture 1, motion, camera, scene, and audio fully preserved from Video 1. Reactor or FaceFusion still work as a pass on the finished file.
And also, i personally don't see a reason to use that custom node since the minimax h3 model can already do face swaps, i do it with a prompt like this one:
- source clip in Reference Video, the face I want in Reference 1... here's an detailed prompt example:
```
subject_definitions:
<Subject 1> is the replacement person whose facial identity comes only from <Picture 1>: [face shape, eyes, brows, nose, lips, skin tone, age range, hair color, hairstyle, facial hair or makeup]. Body, pose, gesture timing, expression timing, and performance come from the original performer in <Video 1>. Wardrobe, hands, and body proportions stay those of <Video 1> unless they are also visible and intentional in <Picture 1>.
<Subject 2> is the original scene in <Video 1>: background, props, lighting, shadows, framing, and occlusion.
<Video 1> is the source video for the target video edit, including its camera path, cuts, pacing, and motion.
<Audio 1> is the synchronized audio track of <Video 1> and is the complete final audio of the target video.
summary:
[video editing + reference generation + audio reuse] The target video is an edited version of <Video 1>. The original performer’s face is replaced by <Subject 1> from <Picture 1> for the entire duration. Motion, pose, timing, camera, cuts, environment, lighting, and <Audio 1> stay those of the source. No new dialogue, music, or effects are generated.
retention_analysis:
<Subject 1> (appears in [Shot 1]): attribute_transfer - facial identity, face shape, eyes, skin, and hairstyle from <Picture 1> replace the original face in every frame; the transferred face follows the source head pose and expression timing.
<Subject 2> (appears in [Shot 1]): fully_preserved - background, props, lighting, shadows, and occlusion stay unchanged.
<Video 1> (source video editing): fully_preserved - camera path, framing, cuts, motion blur, and action timing are unchanged; only the face identity is edited.
<Audio 1>: fully_copy - <Audio 1> is reused 1:1 as the target video’s complete final audio track, including voice, room tone, effects, and music. No new audio is added.
detailed_description:
The target video is a live-action edit of <Video 1> in the same photographic style, color, and lighting as the source. [Shot 1] The shot begins on the source framing of <Video 1>. <Subject 2> remains exactly as filmed: same background, props, light direction, shadows, and depth of field. The person performs the same action, body path, head turns, blinks, and mouth shapes as the original performer, at the same speed. From the first frame, the visible face is only <Subject 1> from <Picture 1>, with that face shape, eyes, skin, and hair held stable as the head rotates. The jaw and lips follow the source performance so they stay locked to <Audio 1>. Hands, clothing, and body stay the source performer’s unless covered by <Picture 1>. When hair, hands, or objects pass in front of the face, occlusion matches <Video 1>. The camera repeats the source move and holds a static shot wherever <Video 1> is static. No cut is added. No identity drift, morphing, extra people, or background change occurs. The audible track is only <Audio 1>, unchanged.
overall_soundscape:
The complete soundscape is the fully copied <Audio 1> from <Video 1>: the original room tone, movement sounds, and any non-speech human sounds, with nothing added or replaced.
non_diegetic_music:
Any music is only the music already inside <Audio 1>. No new score is added. If <Audio 1> has no music, N/A.
```
@LowQualityEnjoyer Thanks for the info. So if I'm correct your example prompt is only for doing afaceswap on an existing video rather than do the swap as it goes for T2V or similar? Or for T2V/I2V would you instead either not reference the video or reference the image being used as a guide? I'll look into it further :-)
