【MiniMax H3_SparseRef15_Hybrid】
MiniMax H3_SparseRef15_Hybrid is an experimental of Fl2VA-based hybrid models family that incorporates selected strengths of Ref2VA without any LoRA merging. It is designed to improve reference consistency while preserving natural motion, scene flexibility, and long-form stability.
Note: This is not a native Ref2V model. It is based on Fl2VA with an added "Ref" effect, so please use LoRAs designed for Fl2VA rather than Ref2V LoRAs.
The core concept is a sparse Ref2VA influence applied through AdaLN across 15 main transformer blocks, rather than using Ref2VA uniformly throughout the model.
This sparse structure is intended to retain more of FL2VA's motion freedom and scene behavior while strengthening character identity and reference consistency.
In testing, the model has shown strong resistance to character drift across long multi-clip generations, including cases where the character changes direction, temporarily leaves a clear frontal view, or continues through many consecutive clips.
Different versions may vary in model size, precision, quantization, speed, memory usage, visual quality, and reference strength.
For reference, Pruned_INT8 is a standard, lightweight pruned model in which the full model's "time-conditioning" has been compressed to 8 dimensions and the large-scale Attention/MLP matrices have been converted to the INT8 ConvRot format.
In contrast, the Pruned_Partial-INT8 series also converts large-scale Attention/MLP matrices to the INT8 ConvRot format but reconstructs and retains critical components—such as AdaLN and time-conditioning—at FP32 precision.
<Hybrid_Pruned_INT8>
A lightweight Pruned Partial-INT8 version of SparseRef15 Hybrid.
Technical characteristics:
- 15 sparsely distributed Ref2VA AdaLN blocks within the 50-block main transformer
- Pruned H3 architecture
- Partial INT8 quantization
- Designed for lower memory usage and faster inference
- Strong emphasis on character identity consistency and long-form stability
In extended multi-clip testing, this version maintained character identity with very little visible drift, even over long sequences.
On suitable hardware, it is considerably faster and lighter than the heavier non-INT8 variant while retaining good motion quality and overall visual coherence.
<Hybrid_Pruned_Partial-INT8 Ver.1.0>
A newer Pruned Partial-INT8 Hybrid variant focused on stronger reference consistency and long-form character stability.
Technical characteristics:
- Pruned H3 architecture
- Partial INT8 quantization for reduced model footprint and faster inference
- Hybrid FL2VA / Ref2VA structure
- Reference behavior tuned for stronger character identity persistence
- Designed to remain stable across long multi-clip generations
Compared with "Hybrid_pruned_int8", this version shows noticeably stronger reference persistence.
In long-form testing, character identity remained highly stable across a 21-clip sequence with very little visible drift, even through repeated changes in pose, direction, and clip transitions.
The stronger reference behavior can sometimes reduce scene freedom or resist situations where the character is expected to remain fully hidden for an extended period. In return, this version is particularly well suited to long-form generations where character consistency is the highest priority.
The Partial-INT8 structure also gives this version a much smaller model footprint and significantly faster inference than the heavier non-INT8 Pruned variant on supported hardware.
<Hybrid_Pruned_BF16>
This is a hybrid derived from BF16 that has undergone only minimal pruning.
The hybrid method is the same as Pruned_INT8.
Description
Pruned Partial-INT8 Hybrid
A lightweight H3 hybrid focused on strong reference consistency and long-form stability.
Partial INT8 quantization reduces memory usage and improves inference speed, while character identity remains highly consistent even across extended multi-clip sequences.
FAQ
Comments (32)
Can this be used with any turbo loras?
@kunde2
Please use the Fl2VA LoRA designed for pruned for pruned models. Some full-version LoRAs may not function correctly.
Is the BF16 file no longer available?
@Mr_Fei
It is a large file, so I plan to release it once I successfully upload it.
@Aki7777777 ok
@Mr_Fei
It runs heavily and hasn't been fully tested, but I have released the Prude version.
I would love to hear your thoughts if you'd like to share them.
New version not given any improve.
Not tested 20 steps since its unrealistic case for 99% users.
With any turbo loras its same bad as dasiwa. Distorted sound and blurry as hell even on 8 steps.
With any combines of acceleration nodes, with just full resolution or with lowres+upscale - Eros beats this model in all tests.
@velanteg
That is a shame.
Personally, I’ve had good results with "minimax_h3_fl2v_lightx2v_turbo_4step_v0.1_comfy." Maybe give it a try?
@Aki7777777 I agree, you should give it a try. when i used https://civitai.red/models/2898443/comfyui-fastvideo-fasth3-4-step-lora-unofficial
it surpasses eros. i also compare those models.
after using the accelerator in the link, this one here became my default model. 4-5 steps only, and improved audio
@Adaptalab0r Its fl2v lora. I am talking about using ref2v case where everything breaks.
@velanteg oh, i see. gotta test it a but more myself then. do you have a recommendation which ref turbo lora to use in general??
@Adaptalab0r best result i got with 6-steps turbo lora, at least better video. Audio is unstable.
@velanteg @Adaptalab0r
I’ve added a note to clarify this. This is not a native Ref2V model. It is based on FL2VA with an added "Ref" effect, so please use LoRAs designed for FL2VA rather than Ref2V LoRAs.
@Aki7777777 Thanks!
@Aki7777777 Well, i just tested Sparceref15_prunedpartial and new 4-steps Dasiwa with this Fast 4 steps lora in 6-steps workflow: 4 steps 0.3mp + 2 steps 0.7mp upscale. For now its best pipeline to get quality without blur and still without generation time spike over 5 minutes (takes 260 seconds for 12 seconds video). Dasiwa won, Sparceref15 made it blurry.
5-steps (1 step upscale) also was fine with image (but unstable sound) on Dasiwa and Eros but dont work with Sparceref15.
Plain generation in full resolution i stoped testing - its takes too long and pointless since 4+2 is already perfect and faster.
@velanteg
Please try using only the default ComfyUI nodes and the "low-step LoRA" I recommended.
I cannot offer advice on any non-standard usage beyond that.
You can view examples of "appropriate" output further down the model page for reference.
the Pruned_Partial-INT8 version seems to be able to recover even when faces or accessories are off-screen for long periods of time. its easy to see the Ref effect is good.
@Motenashi
I agree. However, during testing, we occasionally found it to be "too strong," so please try using it selectively depending on your specific needs.
i need workflow plz
@IGLXX47
You can find the official ComfyUI MiniMax H3 workflows here:
https://docs.comfy.org/tutorials/video/minimax/minimax-h3
I recommend start with an FL2VA workflow. If needed, you may also try a Ref2VA workflow depending on your purpose. Use FL2VA-compatible LoRAs.
@Aki7777777 THX
@Aki7777777 BROTHER WHEN I USE UR MODEL THERES SOME DISTORTION IN THE VIDEO
@IGLXX47
What kind of workflow and which LoRA are you using?
Are you using a basic, minimal workflow with standard ComfyUI nodes, and a LoRA designed for FL2VA?
The original Ref2va model seems to have a foreign speech / phoneme / stutter bug / limitation on 90% of the generations, especially when used with turbo 4 or 8 step. I did try the newer turbo versions like 8 step (1.0) 768 but that did not fix it. So I tried a couple of alternate models (Eros10 etc.) to no avail. Then I came across this model right here. This model is the only one that gives awesome speech even in foreign language. What would you attribute this success to ? Because whatever the fix is, I'm going to be looking for more / other models that have this bug fixed.
@randombrowser1234
Is the model in question Pruned_INT8 or Pruned_Partial-INT8 Ver.1.0?
Which one are you referring to?
@Aki7777777 Both deliver very good foreign language speech and very good identity retention from reference images. All other models fail at it. Except minimax_h3_hybrid_fl2va_ref2va_b25-49-int8.safetensors this one also delivers. Is there a connection between your models and the b25 ?
What I understand is that the H3 audio issue happens only with the ref2v architecture but it can be mitigated by hybridizing it with fl2v whichs audio is ok, the drawback being less good Ref2v compared to pure Ref2v.
Just did an additional test with a fixed seed, I am getting better pronunciation with the b25 hybrid model, but yours come very close 2nd. Adding to the confusion, with a different fixed seed the results flip, but nevertheless those are the best two / three models anyway for this use case.
@randombrowser1234 mind posting your examples here? been trying to get different languages to work as well
After more testing I am getting even better audio results with the PARTIAL.
@hatt2 All I can say is in my personal experience as a end-user, I got much better results using the 2 models I referenced in this post. 1) minimax_h3_hybrid_fl2va_ref2va_b25-49-int8.safetensors and 2) the partial version on this page. Good luck.
@randombrowser1234 @hatt2
Thank you for the detailed verification. While I do not have a direct connection to the b25 model, my model approach is differs from simply layering LoRAs; this suggests that—despite differences in the balance regarding audio and reference integration—there may be similarities in behavior. It is also very interesting that the "PARTIAL" version yielded better audio results. This may suggest that some of the audio-related issues stem from specific elements associated with the Ref rather than the overall concept, given that the process involved pruning from "Full Fl2VA," partially adding the Ref effect, and applying selective compression.