lightx2v_4step_turbo_v0.1+ rank21 version (300mb) uploaded; works with pruned/non-pruned
Current best settings: 5-8 steps er_sde/res_multistep simple. Strength 0.75 seems to be the best (0.5-0.8)
New larryvrh v4 600_ema uploaded.
This LoRA is an early version of a training run currently underway that larryvrh is doing. Degrades audio quality at 4 steps.
Best used at 6-8 steps res_multistep + simple
Using 850ckpt for anime and 500ckpt for realism works best
Examples are generated with 1.0
I did not train, distill or create the original Turbo LoRA weights. This model page only provides modified compatibility versions intended to allow the remaining compatible LoRA adapters to load with ComfyUI's built-in LoRA loader when using the pruned/curve-form model.
Full credit for the original h3_turbo_4step_pruned belongs to larryvrh. And drbaph for the pruned comfy compatible version
Full credit for the original lightx2v_4step_turbo_v0.1 belongs to lightx2v team and kijai for the comfy compatible version
I'll be uploading different Minimax H3 Turbo LoRAs here and crediting the original creators
Description
FAQ
Comments (61)
is there a recommended strenght ?
Example was generated with 1.0, seen people get good results with 1.0-3.0
Noob question: do i load it normally with [load lora] node after the main model loader?
Yea, workflow is embeded in the example video, download and drag into comfy
@tsolful thank you so much, love your huggingface btw.
Hey does this work with the Ref2V model also? (I know it works for the FL2V model
Have not tried with Ref2V sorry, don't use ref2v model as much as t2v/i2v so i cant speak on it myself, I've seen a generation using it that turned out well
@tsolful Ill try now! Thanks
Also testing Ref2V now. Currently testing 4, 6 and 8 steps at lora strenght 1.0. 8 Steps still running at 4 and 6 steps the video looks kind of "burned out" for the lack of a better term. will try to play with the strength later. takes some time.
Any other results or experiences so far?
I found the 8 Steps and a strength of 0.3 works well for Ref2Vid.
@meeatsmeat Thanks for that! Yeah I ran mine using 4 steps and it was really bad, I think in the same way you described. I heard you have to increase the Lora strength if doing 4 steps.
I am essentially running this as a draft mode generation so it shows the general prompt adherance until I get a good prompt, then I run it full resolution + upscale
So i tried your worfklow with the lora and without it on 8 steps and there is ZERO difference in the quality (both bad) and the generation time
The lora is still in training, the new 500ckpt is improved on audio quality
@tsolful i know, it was just an honest feedback. I mainly rated image quality. It's like there is no impact at all using the lora or not. It's just faster and lower quality bc of less steps (4 or 8). As i used your workflow, i don't think i did something wrong. Can you maybe post a comparison video from your results?
Yeah, I took it as honest feedback. https://civitai.red/images/138905130
thanks!. you are hero.
You came so quickly, thanks
it's not perfect but it also not bad. i am using a 40gb ref2va_pruned_bf16 with that lora. audio is distorted but i don't care. im also using default res_multistep and simple - and 4 stesp is out of the way but 8 steps is usable
Yeah new uploaded 500ckpt version has slight improvement to audio at 8 steps
I got overfitting result trying sampling from 4 step to 10 step, the audio quality is acceptable now but frames seems still a bit off. It's a good start, so I have to wait for a bit.
Cool, thanks. With 8 steps and a Lora strength of 0.6, the results are quite good.
Do these work with the convrot ref2v? Comfy keeps throwing unable to load lora block errors at me which usually means there is a mismatch with the model
Works for me for ref2v with the pruned fp8 model. I found that strenght 0.3 and 8 steps work quite well so far. Can't help you with the lora block errors though.
I think they fixed this in the latest version of comfy
On 3060 12gb 6-8 steps 0.3mp res_multistep + simple best result.
what weight did you put the lora at? I'm using a 3060 and with the lora i get horrible results
@mrweaz I'm also using 3060 12gb, I put the lora at 0.6, 8 steps, 10 seconds length and 0.4 megapixels. I'm getting bad results when there's a lot of motion or any action scene.
actually with ckpt500 version visuals are great. but audio is still bad.
edit: the original creator of this turbo lora also created this sampler node to use with. audio is better with it. still not good as no-turbo though. https://github.com/Larryvrh/ComfyUI-MiniMax-H3-Turbo
The audio issues are waiting to be merged to comfy as the wizard kijai already has a fix pr'd https://github.com/Comfy-Org/ComfyUI/pull/15243
Man you are a savor!. what used to use 200s now takes 90s for 120 frames 300x600px. that chkpt500 is excellent quality at 6 steps euler+beta, In fact if you update comfyui to the very latest version you no longer need any special nodes, just use regular lora and load diffusion models to load them. The adherence to prompt, and adherence to the details of reference images is astonishing, practically making wan and ltx fit for trash only. The audio needs more work though
All credit goes to larryvrh for training the lora and drbaph for pruning. The audio issues are waiting to be merged to comfy as the wizard kijai already has a fix pr'd https://github.com/Comfy-Org/ComfyUI/pull/15243
@tsolful Thanks to the three of you.
@mddw could you kindly share your workflow? i just need to check something cuz my generation times are not improving much.
@necroryona I am using the workflow provided by @tsolful, though i connected the lora using typical power lora loader . The speed issue sometimes has to do with how fast your cpu-ram is: Because the model is so big that offloading it essential to all gpus below 32GB, thus the bottleneck is not gpu speed, but rather how fast your cpu-ram sends and coordinates the shuffling of models. Thats why a 5080 gpu 16GB with a ddr4 ram will generally generate slower than a 3XXX gpu with ddr5 ram!
.6 lora strngth 8 steps , You are legendary pokemon , thanks alot ... from 13m to directly 5m
i tried your setting but all i got is pixel soup :'(
euler is kinda bad if u are using ref2video.
doing a bunch of tests with other ones, they grab more the reference and prompt adherence.
Which did you prefer? Normal Res multi?
my pp has no mouth but it must scream
Thanks for the reminder, I need to finish that game...never read the book though.
Sorry to be that guy...
Is this working with gguf versions?
I tried the Q3 Ref and FFLF, activating and deactivating the lora, but I got the same output (enabled/disabled)
I will download the 21Gb version, but I´m not sure if it will fit in my gpu
Running on Q5, works like a dream
@Unisol58930 where is Q5?, I have the Q3 from RealRebelAi, do you mind sharing the link?
The joeyGambino ones? What gpu are you using? the Q5 is 25Gb, I only have 24Gb of Vram and I need to consider Lora, encoders and vaes too
@Unisol58930 how is the Q5 in terms of speed and quality compare to the pruned verisons? Like the minimax_h3_fl2va_pruned_fp8_scaled and minimax_h3_fl2va_pruned_int8_convrot models?
@bhopping pruned/curved Q5 is nice for 12bg. Same speed as Q4 or Q3
If you use an rtx 5000 family with 16gb VRAM you can run the int8 convrot quant.
I'm on a 12GB card and I run the int8 convrot model without issues. It's way faster and better quality than GGUF.
Yeah, gguf models are on their way out with Kijai's new int4 quantization (w4a8) waiting to be merged. I'm able to run the int8 model with RTX 3060 12 GB + 32 GB RAM, which is still faster than a gguf quant that fits in VRAM + the quality of the int8 convrot is better compared to q8 gguf
for some reason it turned everything slow motion?
@thumperunit441 Yeah i've noticed this aswell for I2VA, bit of a seed hunt for a good output. T2VA is running fine for me
The video is great and very fast but the audio is a complete disaster. Keep up the good work I look forward to seeing how it goes.
Using the authors sampler node helps the audio alot actually but without that its no good.
hi, authors sampler node? audio is destroyed in mmy outputs too
@Baka_Oppai @Sinusplusminus Update to the latest comfyui commit as the audio should be fixed, Which is what ive been running with the examples posted https://github.com/Comfy-Org/ComfyUI/commit/bdcb886a4705a03cf40f4a7226de9fc7c059fc90 If you are still running into bad audio here is the authors custom node https://github.com/Larryvrh/ComfyUI-MiniMax-H3-Turbo
@tsolful thanks
i tried raising the lora strength to 1.8 and my god, the animation was flawless and the prompt adherence was perfect, but the video itself was so mushy and distorted so i had to lower it down to the 1.0, i kinda feel like i had a small feel for bf16 or something, shit is scary good.
i am not sure exactly what this lora is doing, but it's a hit or miss so far, can't wait for the full release and more optimized WFs and loras.
will experiment more and see how it goes and post my findings here, PLEASE keep updating the workflows as you post the new loras and always post the fastest workflows you find (whoever you are lol).
I do most of my testing on a 5090 Runpod, but locally I generate with a 3060 12 GB + 32 GB RAM. I'll upload the simple workflows I use with the 5090 here, and when I get around to fully optimizing my local rig (Currently using Minimax Chunk feed forward, Minimax low vram attention and sage attention from KJNodes), I'll decide if I post a separate workflow or just add to optional files. Also Full credit for the original h3_turbo_4step_pruned belongs to larryvrh. And drbaph for the pruned comfy compatible version
Full credit for the original lightx2v_4step_turbo_v0.1 belongs to lightx2v team and kijai for the comfy compatible version 💚
太棒了,期待后续。手动三联
Seems like using 850ckpt at 1.0 strength will generate overbaked video.
Yeah ive found using 850ckpt for anime and 500ckpt for realism works best
I found that too, but easy fix is just lowering the strength. I didn't extensively test to see how low the sweet spot is, but got down to 0.7 and it seems to work fine.
i use the 800_ema at 0.5 at 10 steps. the 500_non-ema worked at 0.3 for 15 steps. You are right that it seems to overbake at 1.0, but I also didn't try too many settings