【MiniMax H3_SparseRef15_Hybrid】
MiniMax H3_SparseRef15_Hybrid is an experimental of Fl2VA-based hybrid models family that incorporates selected strengths of Ref2VA without any LoRA merging. It is designed to improve reference consistency while preserving natural motion, scene flexibility, and long-form stability.
Note: This is not a native Ref2V model. It is based on Fl2VA with an added "Ref" effect, so please use LoRAs designed for Fl2VA rather than Ref2V LoRAs.
The core concept is a sparse Ref2VA influence applied through AdaLN across 15 main transformer blocks, rather than using Ref2VA uniformly throughout the model.
This sparse structure is intended to retain more of FL2VA's motion freedom and scene behavior while strengthening character identity and reference consistency.
In testing, the model has shown strong resistance to character drift across long multi-clip generations, including cases where the character changes direction, temporarily leaves a clear frontal view, or continues through many consecutive clips.
Different versions may vary in model size, precision, quantization, speed, memory usage, visual quality, and reference strength.
For reference, Pruned_INT8 is a standard, lightweight pruned model in which the full model's "time-conditioning" has been compressed to 8 dimensions and the large-scale Attention/MLP matrices have been converted to the INT8 ConvRot format.
In contrast, the Pruned_Partial-INT8 series also converts large-scale Attention/MLP matrices to the INT8 ConvRot format but reconstructs and retains critical components—such as AdaLN and time-conditioning—at FP32 precision.
<Hybrid_Pruned_INT8>
A lightweight Pruned Partial-INT8 version of SparseRef15 Hybrid.
Technical characteristics:
- 15 sparsely distributed Ref2VA AdaLN blocks within the 50-block main transformer
- Pruned H3 architecture
- Partial INT8 quantization
- Designed for lower memory usage and faster inference
- Strong emphasis on character identity consistency and long-form stability
In extended multi-clip testing, this version maintained character identity with very little visible drift, even over long sequences.
On suitable hardware, it is considerably faster and lighter than the heavier non-INT8 variant while retaining good motion quality and overall visual coherence.
<Hybrid_Pruned_Partial-INT8 Ver.1.0>
A newer Pruned Partial-INT8 Hybrid variant focused on stronger reference consistency and long-form character stability.
Technical characteristics:
- Pruned H3 architecture
- Partial INT8 quantization for reduced model footprint and faster inference
- Hybrid FL2VA / Ref2VA structure
- Reference behavior tuned for stronger character identity persistence
- Designed to remain stable across long multi-clip generations
Compared with "Hybrid_pruned_int8", this version shows noticeably stronger reference persistence.
In long-form testing, character identity remained highly stable across a 21-clip sequence with very little visible drift, even through repeated changes in pose, direction, and clip transitions.
The stronger reference behavior can sometimes reduce scene freedom or resist situations where the character is expected to remain fully hidden for an extended period. In return, this version is particularly well suited to long-form generations where character consistency is the highest priority.
The Partial-INT8 structure also gives this version a much smaller model footprint and significantly faster inference than the heavier non-INT8 Pruned variant on supported hardware.
<Hybrid_Pruned_BF16>
This is a hybrid derived from BF16 that has undergone only minimal pruning.
The hybrid method is the same as Pruned_INT8.
Description
This is for experimental purposes. I cannot accept responsibility for any consequences arising from its use.
FAQ
Comments (5)
Does this model include a Turbo LoRA?
If not, which Turbo LoRA and step count do you recommend?
@kkmw15
To be honest, I’m quite torn on this myself.
I used "minimax_h3_fl2v_turbo_4step_v1.1_768p_comfyui_bf16" for this video, but I also quite like the "minimax_h3_fl2v_lightx2v_turbo_4step_v0.1_comfy" model as a safe bet.
However, since many existing low-step LoRAs tend to alter the image quality or over-sharpen the details, I’m currently exploring ways to create a more natural-looking LoRA.
Does not follow prompts
@Dayer69
Could you provide a bit more detail and, if possible, an example prompt? I’d like to know whether the issue is with motion, camera direction, scene changes, subject behavior, or something else.
@Aki7777777 litterly any promtps, 'The green anthropomorphic female cat with pink eyes and white chest fur sits atop an erect penis in a cowgirl position on a grassy field under a cloudy sky. The green anthropomorphic female cat straddles the man's lap as she moves her hips rhythmly up and down while gripping his waist for stability; her tail flicks behind her and thick foliage sways in the breeze as wispy clouds drift across the blue sky above. She grinds her pelvis downward against him before lifting back up to repeat the cycle of thrusting; her breathing becomes heavy-chested as she accelerates into quicker vertical movements. Her voice is breathy and high-pitched when she moans, "Mmm... so deep," followed by more intense gasps from the woman with pink eyes as they reach peak speed. Ambient forest sounds like rustling leaves provide a low hum beneath sharp foley hits of skin slapping together during each rhythmic pelvic impact.' Just made her shout in a a wierd language
