CivArchive
    V3 Final! DeJanked Speed Hack Hunyuan T2V Final Boss - v3
    NSFW
    Preview 53565304

    Major Overhaul. True refiner speed hack.

    DeJanked Speed Hack Hunyuan T2V Final Boss:
    Are you tired of your AI video workflow crawling slower than a grandma playing Frogger? Do you crave blistering speed without sacrificing jaw-dropping quality? Buckle up, because this isn’t just a workflow—it’s the Final Boss of Hunyuan T2V optimization. Through the fiery trials of placebo hacks, sanity-testing, and daisy-chain wizardry, this setup slashes render times, keeps your GPU breathing, and still pumps out e-girl-quality frames that’ll have you questioning reality. Dare to try it? It’s fast, it’s smooth, and it might just blow your mind (but not your GPU).

    TL/DR:
    Standard: 180 seconds (Great quality)
    After Speed Hack: 100 seconds (great quality)

    No descaling/rescaling
    No wavespeed and minimal TeaCache.

    No XL bs placebo (I feel scammed! See testing below)

    This is something else (please post your results)


    The testing:

    Wavespeed: I will not test it due to it being dependent on Triton, which lots of people have trouble installing on Windows (and can screw up a lot of things with a windows install...those using WSL and Wavespeed can probably figure out how to shove in wavespeed for their own use. this is for maximum availability to the widest amount of users

    Hardware I am using: 3090 TI 24g VRAM. 64g Ram, WSL.

    100 frames, 10 steps, 3 step refiner 512x512 (no up/downscaling) for base.

    Goal:

    - find speed without losing quality.

    - no use of upscaling or downscaling.

    - no tricks, just using standard nodes to their max potential.

    Method: 3 generations for each phase.

    Vanilla baseline:

    Round 1: (basic Vid Gen, no tweaks, just gen and refine)

    Flow: normal (no glitchy movements of note)

    Quality: high

    180 seconds

    Pass

    Round 2:

    Teacache sampler at 1.6 (fast) for both main and refiner

    Flow: normal

    Quality: high

    172-175 seconds

    Pass

    Round 3:

    First Teacache sampler at 4.4 (shapeless). Refiner at Normal

    Quality: Average-poor, refined was better but losing fidelity, even when kicked up to 4 steps.

    Flow: okay, slightly glitchy (possible exaggerated normal glitchyness)

    154 seconds

    Fail

    Round 4:

    TeaCache Samplers on fast, introduction to TeaCache Thresh node at 0.15.

    Quality: Good

    Flow: good

    180 seconds (???)

    Fail (pointless to possibly clashing with the sampler)

    Result: having samplers in both main and refiner on fast seems to be the happy medium. possible further testing for perhaps faster settings, but will call this enough (a few extra seconds either way isn't the gains I am going for)

    On to XL Workaround to kick and see whats what once and for all.

    XL hack:

    Gone, placebo, nerfed! 185 average. Remove and toss into fires of..etc (I feel I've been scammed!)

    next up, Daisy Chain Refiner Speed Hack

    results:

    1 main and 2 refiner steps without encode/decode:

    100 frames, 9 steps

    Quality: high

    Flow: good

    100.34 seconds

    Quality can be altered by raising or lowering a step at the beginning if desired, but 5,2,2 is producing stellar results. Recommend starting here

    Why does this work?

    I don't know, but I assume the first step gives shape. 2nd fleshes out, and 3rd refines. each one working off the less for less overhead and less need to start from scratch, building off the last. Alternatively, simulation universe and pixie dust...obviously.

    There you have it. poke holes in it if you can.

    Try it, its free, and for me, it is working blazing fast.

    Some odd unique occurances come with some renders moving a bit fast (re-render same seed but drop fps down a bit)

    From here, see if you can improve it...but before you run this, run your normal non weird workflow without wavespeed for your own vanilla testing to sanity test. make sure to run it 3 times though (need time for cache to warm up. by the 3rd run, you're hitting optimal speeds)

    I used the hunyuan 8b 720 (fast) model and the only LoRA I had active was the fastvideo lora (found on civitai) at -0.30 (positive for big model, negative for fast models). Egirl lora added for main video model just for fun but not part of the test.

    WARNING: nothing is downscaled. monitor your GPU. best to maybe go smaller. 512x512 to start, then work up (or down) from there. This speeds up your rendering time, but it doesn't lower the GPU overhead.

    Description

    Total rebuild:
    XL out
    Daisy Chain enabled
    Removed all steps from start until final result double refined.

    FAQ

    Comments (85)

    aliabougazia85159Jan 17, 2025· 1 reaction
    CivitAI

    Well done!. The teacache vid gen node has to have a set threshold value, though.

    saturngfx
    Author
    Jan 18, 2025

    should be on 0.15....hmm, did I set it down?
    Also, what types of times are you seeing (before/after)?

    aliabougazia85159Jan 19, 2025

    To be honest it's not much faster to the ones I already see on my own setup. About 14 min for a 65 frames 720x512p with 20 steps using the original model.

    saturngfx
    Author
    Jan 19, 2025· 1 reaction

    @aliabougazia85159 strange. I am getting on a 1 to 1 comparison a 33% decrease in time generally speaking. wait...14 minutes? you running on CPU or something? Whats your specs? also, why 20 steps? teacache and the fastlora brings that down to 11, Something is odd with your workflow. let me test (working on it now actually.)

    result. aiming for an outcome of 720x512 before upscale on my 3090 with 11 steps (because more is pointless) in full model. gonna do 70 steps (multiplier x 7): yeah, 131.71 seconds
    Even double the steps to 22 wouldn't come nearly that close. you sure you weren't running like 700 steps? keep in mind, if you want 70 steps, you only put 7 in the multiples section (10x7) . I've done that...more times than I care to admit. I'll post the video. anyhow, next version coming out today. V2 is a bit more streamlined, better notes, better refiner, etc.

    aliabougazia85159Jan 19, 2025

    @saturngfx I'm using a 3060 with 12gb vram. I don't know what others are getting, but the fastest I got was 7 minutes using the teacache workflow, but I don't like its outcome especially when I use loras.

    What vram do you have?

    saturngfx
    Author
    Jan 19, 2025· 1 reaction

    @aliabougazia85159 3090 TI has 24. still...something is odd. I...don't know, maybe someone else with 12g card can help. I can only guess, which is what you are doing also.
    Do some reddit digging..see if someone can troubleshoot your setup with hunyuan...it kinda sounds to me that something might be off...like, possibly windows related, like your swap file size may be tiny or your harddrive is all but full, something on those lines. Find out what other people are getting for your card setup using the same settings (just use a simple workflow for now so others can figure it out)

    saturngfx
    Author
    Jan 20, 2025

    @aliabougazia85159 a number of issues happened for some reason. redownload it. I fixed all things, from the teacache, to the secret defaulting of 512x512 for what supposed to be tiny quick filler pics, slowing things down considerably. sorry. not sure why it got nerfed. That alone could cause a big slowdown. 2.01 is now back to where it should be.

    TheAIDoctorJan 18, 2025· 1 reaction
    CivitAI

    Testing out your workflow.

    "come on folks, less people and more styles...if only to help stablize the videos for a better end result. Anyhow, I default using edge of reality because its...sort of a style. your milage may vary"

    Glad you like my Lora ;)

    TheAIDoctorJan 18, 2025

    Any chance you can provide a workflow example or screenshots on the setup using the BF16 full hunyuan model? I'm struggling to get it working, not sure if I just have the wrong clip model or what.

    Requested to load HunyuanVideoClipModel_

    loaded completely 9.5367431640625e+25 7894.8529052734375 True

    CLIP/text encoder model load device: cuda:0, offload device: cpu, current: cuda:0, dtype: torch.float16

    clip missing: ['text_projection.weight']

    model weight dtype torch.bfloat16, manual cast: None

    model_type FLOW

    Token indices sequence length is longer than the specified maximum sequence length for this model (94 > 77). Running this sequence through the model will result in indexing errors

    Requested to load HunyuanVideoClipModel_

    loaded completely 9.5367431640625e+25 7894.8529052734375 True

    Requested to load HunyuanVideo

    saturngfx
    Author
    Jan 18, 2025

    First, your lora really ups the game. great job.

    Now, about the issue:
    ahh, hmm...that issue. Lets see...I want to say its the weight_dtype maybe? if its on fast, put it no regular and reload. I tend to find a quick switch of things like that and reload fixes most issues. check your normal workflows to match 1 for 1. in the hunyuan section (for now) its pretty much a straight shot of however you would normally set up. I am starting to fall out of love with the fast model and becoming quite the fan of the big boy (working on a big upscale tweak now as we speak). If you're pulling your hair out, maybe hit reddit or send me the workflow (is there an option to even send files direct here) and I can poke it.

    766788Jan 19, 2025
    CivitAI

    Very interesting work. Could you explain what the benefit is to using a txt2img for noise?

    saturngfx
    Author
    Jan 19, 2025· 2 reactions

    Although I am not sure, my hypothesis is this. when using hunyuan video, say you want a 512x512 vid at 50 frames. it will make 50 blank frames at that size. it then needs to prepare all those files. This method only has 1 at the size you want. the rest are tiny little images that it can quickly decode, but it then alters them after decoding to the proper size. the decoding process is the slow bit, so by giving them tiny little images for all but 1, it rapidly increases its processing weight at the start.
    Or it could simply be voodoo. I would like someone smarter than me to explain why this works.

    766788Jan 19, 2025· 1 reaction

    @saturngfx Very cool. Thank you!

    marviskealan254Jan 19, 2025
    CivitAI

    Is there any benefit to manually toying with the depthflow before it hits Hunyuan? I know other workflows seem to just automate it. I imagine you could do the same sorta thing by using the seed to influence some random logic for that if you wanted. But so far from my limited testing I am not too sure if manually trying to get the depthflow to somewhat resemble the motion you are describing with your prompt is actually helping or not. What are your experiences?

    saturngfx
    Author
    Jan 19, 2025· 1 reaction

    Oh hell yeah, depthflow tinkering makes all kinds of magic go on. be it slow steady panning, rapid quick movements, deep sinking, etc...

    saturngfx
    Author
    Jan 19, 2025

    btw, I think this was meant for my other workflow (the image2vid one).

    marviskealan254Jan 20, 2025

    @saturngfx It was. Ahaha. My bad. But good to know, thanks.

    saturngfx
    Author
    Jan 20, 2025· 1 reaction

    @marviskealan254 its fine. I am confusing myself also. gotta have the webpage up to know which I am talking about. need a more clear naming convention. heh

    dominic1336756Jan 19, 2025
    CivitAI

    do we need to connect nodes? because apart from the initial rendering, nothing happens, no upscaling or refining.

    saturngfx
    Author
    Jan 19, 2025

    nothing happens? what does your console say? are the switches turned on? did you make sure to swap out the various areas with your model locations? any errors?

    dominic1336756Jan 19, 2025

    @saturngfx No errors, Prompt executed in 43.37 seconds. On the other hand, V3 with Flux works very well.

    saturngfx
    Author
    Jan 19, 2025

    @dominic1336756 @dominic1336756 looks at you stupified So, you got flux working, the thing that has caused me endless grief...but now this pretty much out of the box thing...isnt?
    I...don't know. update maybe? maybe something is going on with something behind the hunyuan refiner node? move the boxes aside and ensure teacache isn't doing something weird, like giving a nan value verses 0.15 or the like. poke around and figure where the issue is basically.I have had a couple others say there are issues with teacache...possibly update issue or something.

    Wait, what do you mean V3? which workflow are you looking at?

    dominic1336756Jan 19, 2025

    @saturngfx sorry my english is bad. The one where you had trouble with Flux works fine. The latest Speed on the other hand only renders the basic image, without upscaling or refining. I hope to be more explicit

    saturngfx
    Author
    Jan 19, 2025· 1 reaction

    @dominic1336756 Got it, yeah, the turbo t2v. Well, nice one.
    Anyhow, others are reporting this problem with the teacache. redownload the model (just updated it a few minutes ago) all it does is separate out the teacache in the refiner area...change it as the note says. that might fix the problem.

    and I might need to revisit the flux on my turbo to see what I broke.

    dominic1336756Jan 20, 2025

    ok my friend

    Output will be ignored

    Failed to validate prompt for output 368:

    Output will be ignored

    Failed to validate prompt for output 367:

    Output will be ignored

    Failed to validate prompt for output 361:

    Output will be ignored

    Using pytorch attention in VAE

    Using pytorch attention in VAE

    VAE load device: cuda:0, offload device: cpu, dtype: torch.bfloat16

    Requested to load AutoencoderKL

    0 models unloaded.

    loaded completely 9.5367431640625e+25 470.1210079193115 True

    CLIP/text encoder model load device: cuda:0, offload device: cpu, current: cpu, dtype: torch.float16

    clip missing: ['text_projection.weight']

    model weight dtype torch.float8_e4m3fn, manual cast: torch.bfloat16

    model_type FLOW

    Requested to load HunyuanVideoClipModel_

    loaded completely 9.5367431640625e+25 7894.8529052734375 True

    Requested to load HunyuanVideo

    loaded partially 8377.999984741211 8377.100646972656 330

    100%|██████████████████████████████████████████████████████████████████████████████████| 11/11 [00:33<00:00, 3.09s/it]

    Requested to load AutoencoderKL

    0 models unloaded.

    loaded completely 9.5367431640625e+25 470.1210079193115 True

    Prompt executed in 306.57 seconds

    saturngfx
    Author
    Jan 20, 2025

    @dominic1336756 ahh, upscaler. what do you have the math set to? does it work without the upscaler turned on? What size image are you doing? also, 300 seconds...is that your first run at launch? (takes awhile for first launches, I tend to do just a starter picture of like 10 frames, 256x256 just to get the initial load out of the way. second run also takes a bit of time, but 3rd time its hitting its stride, this is for all video workflows though, not just this)

    dominic1336756Jan 20, 2025

    @saturngfx rendering is good, good in poor quality, since there's no refiner or upscale. the size is the default, I have the impression that the get_nicerefine node can't find the image.

    saturngfx
    Author
    Jan 20, 2025

    @dominic1336756 confirm version you're using just to make sure we are on the same version.

    saturngfx
    Author
    Jan 20, 2025

    @dominic1336756 Do you have the refiner turned off? right now, nicerefine is set from the results of the refiner. if it is on, then something is janky (go figure).  Lets troubleshoot from the beginning. so, it goes through the hunyuan video process fine, then it shoots over to the refiner (assuming its enabled.) move those panels aside and see if there are weird things, red circles. expand and see where the hangup is.

    dominic1336756Jan 20, 2025

    @saturngfx I've expanded all the nodes, I do have two red nodes, the two links attached to get_nicerefine (upscale image with model and get image size), the refiner is ON, in fact everything is ON

    saturngfx
    Author
    Jan 20, 2025

    @dominic1336756 and confirm the set nicerefine is connected and doing fine at the end of the sharpener. it is in the refiner column and says (experiment, too sharp? lower). pull that aside, follow the noodle to what is coming out...that is where the set_nicerefine should be located. If its not...well, something wonky happened.

    syntecJan 19, 2025
    CivitAI

    The initial output works no issues. The Teacache for vidgen node originally had "True" and "NaN" as the values, which wouldn't work. I tried changing them to "Hunyuan_video" and "0.15" but just get this error:
    Failed to validate prompt for output 368:

    * TeaCacheForVidGen 215:

    - Failed to convert an input value to a FLOAT value: rel_l1_thresh, hunyuan_video, could not convert string to float: 'hunyuan_video'

    - Value not in list: model_type: 'True' not in ['hunyuan_video', 'ltxv']

    What values should be in that node?
    To clarify - the intial gen works regardless. The error messages trigger if I enable upscaling and/or refine. It'll generate the low-res version, then go idle.

    saturngfx
    Author
    Jan 19, 2025

    weird. I don't know why that happened.
    Anyhow, on the teacache node hidden behind the denoise, just the typical on, and the rel I have at 0.15. 4 steps.
    Null? thats an odd one. I just redownloaded it and threw it in, its fine on my end, so maybe an update is needed or something? if that still gives you grief, just unplug it and have the model node directly link into the denoiser. But yeah, check to ensure you are updated across the board. comfyui and extensions (gonna check myself also come to think of it)

    syntecJan 19, 2025

    @saturngfx lol it is probably that, will check for updates and report back if so with my dunce cap clearly in place.

    AicushJan 19, 2025· 1 reaction

    @syntec Just spent the best part of 20 minutes trying to sus this one out X-D if you look it is asking for 215, this one is hidden neatly away at the bottom of HUNRefine (worth it) may have to move some stuff around was wondering why I wasn't getting the tasty upscales

    saturngfx
    Author
    Jan 19, 2025

    @Aicush Why? ugg...why is this null? does it not transfer? ugg...let me reupload a quick update if only to put that node out in the open in case there is weirdness going on. how frustrating.

    AicushJan 19, 2025· 1 reaction

    @saturngfx I don't think it does you know, I have tried a few workflows and always deal with this myself, but yeah that second box was a sneaky one, as a novice I updated the first one thinking I had sussed it until I realised there was a sneaky second one....please keep the comical notes, I had a joy reading them X-D

    saturngfx
    Author
    Jan 19, 2025· 1 reaction

    @Aicush Thanks for the heads up. I just reuploaded the zip, new version now has that bugger pulled out, noted, etc.

    syntecJan 19, 2025· 1 reaction

    @saturngfx Ah, so it wasn't an update issue it seems, but your tip about looking behind the denoiser helped, that appears to be the node that was causing issues, it seemed to be inheriting the "true/null" stuff, but recreating it straightened it out. I was pretty confused when I disconnected the teacache node up in the magic box completely and was still getting an error with teacache :P
    Oh, just saw Aicush's response as I was typing this one - cheers!
    Will redownload the workflow. Thanks for your help, look forward to trying it out properly!

    AicushJan 19, 2025

    @saturngfx Brilliant title to it X-D, yeah both tea cache for vid gen are still defaulting to True with value of NaN on the latest version still. Hopefully this resolves itself at some point, but your big note should reduce some of the queries :)

    saturngfx
    Author
    Jan 20, 2025

    @Aicush Broke more than that. a number of issues happened for some reason. redownload it. I fixed all things, from the teacache, to the secret defaulting of 512x512 for what supposed to be tiny quick filler pics, slowing things down considerably. sorry. not sure why it got nerfed

    the__RealistJan 22, 2025
    CivitAI

    Question from someone who's not used Comfy UI before.

    - Does this work with Forge?

    - Does this work with image to video?, or is it only prompt based.


    Thanks in advance : )

    saturngfx
    Author
    Jan 22, 2025· 2 reactions

    Forge: No. need to up your comfyui game here. Honestly, it would be best not just to learn comfyui, but if you're serious about the AI hobby, learn WSL...talking a huge leap in rendering time for video and llms.

    Currently, officially, it only works with prompts or video 2 video. Check out my other workflow for a sort of hacky workaround for image 2 video, (not their official yet, just a sort of hold over until we get it proper)

    the__RealistJan 23, 2025

    @saturngfx Thanks for the quick and detailed reply!
    comfyui is something i had considered using but never did as forge seemed more 'straightforward' to use. I shall have a look now.

    Melty1989Jan 22, 2025
    CivitAI

    Thanks for this. I seem to be getting a few more blurry artifacts when refine is enabled. Is this expected? I'm mostly using 320x512 resolution

    saturngfx
    Author
    Jan 24, 2025· 1 reaction

    download latest. its a full rethinking. no more blur, true speed.

    Melty1989Jan 24, 2025· 1 reaction

    @saturngfx In the description you say "I used the hunyuan 8b 720 (not fast) model" , but in the workflow it defaults to the fast model. So, which one is it? I used the non-fast fp8 model and the output was always blurry, but with the Fastmodel its fine.

    saturngfx
    Author
    Jan 24, 2025

    @Melty1989 EXCELLENT CATCH! I totally was using the fast model (reloading must have swapped it before testing), but it should hold the same regardless of models, just if using the bigger models, adjust the steps maybe...1 or 2 more, but the speed increase should be universal. Thanks for spotting that. I need more coffee clearly :)

    Melty1989Jan 24, 2025

    @saturngfx Thanks for confirming. Also, moving the FastLora strength to positive helped deblur the image when using the normal (non fast) fp8 model. I set the strength to 1 , but is there a recommended value for this?

    saturngfx
    Author
    Jan 24, 2025

    @Melty1989 yeah, negatives only for the fast one. I seriously don't know what I am doing btw, just a lot of throwing things in a blender and dancing around when things work. the fastlora thing confuses me still in general. its like...magnets man..how do they even work! Anyhow, I just set it to 1 when using it by itself, but when adding loras, I start decreasing it as other loras tend to pick up the slack.
    Anyhow, noticing a speed increase?

    Melty1989Jan 24, 2025

    @saturngfx At 512x512 61 frames, I'm getting 103 seconds (non-fast model). Running a 3080 (10GB). Pretty good , much lower than the previous workflow! Just used the FastLora and Secret Sauce.

    I wonder if I can get this lower by using City96's Q4_K_M Hunyuan Model, and IbnAbdeen's quantized llama model (IbnAbdeen/llava-llama-3-8b-text-encoder-tokenizer-Q4_K_M-GGUF at main)

    Melty1989Jan 24, 2025· 1 reaction

    @saturngfx So, it doesn't play well with GGUF models. It's an additional 40-50 seconds for me with the same settings.

    saturngfx
    Author
    Jan 24, 2025

    @Melty1989 heh, well, remember to run it 3 times. need for the cache to warm up with major changes or startup, but yeah, all about experimentation. Thanks for posting, others may read and come up with new ideas, or at least heed warnings. Glad you're gaining speed. my other workflow I think cheated..damn lying figures (probably had a downscaler hidden somewhere in there skewing the times at testing). This is mostly just a trick of not decoding/encoding..huge waste of time if you're going straight for refined.

    Melty1989Jan 24, 2025

    @saturngfx Yeah i did run it with GGUF thrice, still a bit slower. Anyway, I'm seeing some weird movements with other scenes. (non fast fp8 model, 512x512 61 frames, fastlora at 0.25, secret sauce at 0.8). Like the animation is looping and artefacts: Imgur: The magic of the Internet

    Edit : So running the same prompt on the old workflow results in much more natural transitions. Here it just seems to be looping. Did multiple runs and can confirm this @saturngfx

    Prompt :
    Realistic video of A confident woman resembling Taylor Swift leans against a Harley Davidson motorcycle outside a retro diner. She is dressed in a sleek black leather biker jacket, fitted jeans, and rugged boots. Her wavy blonde hair frames her face with soft, striking features. The diner behind her has vibrant neon signs and large windows, capturing a classic 1950s Americana vibe. The scene is set in the morning, with the medium shot framing both the woman and the motorcycle prominently in the foreground.

    saturngfx
    Author
    Jan 24, 2025

    @Melty1989 yeah I have gotten some with weird seemingly sped up motion, I did kinda solve it by doing an interpolation at the end doubling the frames.

    Melty1989Jan 24, 2025

    @saturngfx How do I do interpolation? Did you run it through Film VFI , AMT VFI etc?

    saturngfx
    Author
    Jan 24, 2025

    @Melty1989 Its a simple node. I was considering tossing it right before it goes through the final upscaler. just double the frames (will need to also alter your fps, else it will be running slow). Just anytime after it goes through the refinery process when it has to unload the model anyhow. I wouldn't do it early on...otherwise you're gonna be waiting far longer. yeah, either before or after the upscaler step. not sure which would be faster...I would think before. ...I'll test

    Melty1989Jan 24, 2025

    @saturngfx Ok , so what's the node called?

    saturngfx
    Author
    Jan 24, 2025

    @Melty1989 Sorry my dude, I haven't had coffee yet (brewing as we speak). it is Rife vfi. gonna go caffeinate before commenting again. :)

    Melty1989Jan 24, 2025

    @saturngfx No worries, thanks for responding quickly. Take care!

    AicushJan 24, 2025· 1 reaction
    CivitAI

    Good Morning - EnhancedLoadDiffusionModel is missing what node group is this coming from? As my missing nodes list appears empty

    saturngfx
    Author
    Jan 24, 2025

    what number does it say? is it the lora group? (aka, you missing Loras? a triple stack right above the clip text prompt)? If so, thats RGThree, but you can probably swap it out with your favorite lora loader)

    Melty1989Jan 24, 2025· 1 reaction

    @Aicush that's from Wavespeed. You might have to manually clone that in from the repo to the ComfyUI custom_nodes folder.

    AicushJan 24, 2025· 1 reaction

    @Melty1989 That was the ticket many thanks :-)

    AicushJan 24, 2025

    @saturngfx Ill test this out on my peasant hardware and see how I get on :-)

    saturngfx
    Author
    Jan 24, 2025

    seen someone comment, but can't see itnow. possible wavespeed? (I don't have wavespeed in this workflow). also, not sure if I am double commenting btw, the page is being super wonky for me at the moment. But yeah, if you can either tell me the number of the node, or its location on the workflow, I can help.

    AicushJan 24, 2025· 1 reaction

    @saturngfx It was the very first one for loading the diffusion model -very top left

    saturngfx
    Author
    Jan 24, 2025· 1 reaction

    @Aicush ahh, just the model loader. swap it out for the other hunyuan video loader if you don't have it. But yeah, that one is part of the Wavespeed stuff, however its only there because after uploading, I reintroduced wavespeed. you can swap it for the standard model loader. should have done that before uploading actually.

    AicushJan 24, 2025

    @saturngfx No worries got the Wavespeed pack as suggested by @Melty1989 seems to be doing the job well, will log my results on a separate post, Once I have finished :-)

    AicushJan 24, 2025· 4 reactions
    CivitAI

    My Config - 4080 mobile 12gb VRAM - 32GB RAM

    Out of the box experience - Using all the same models/loras etc

    I think if you stay around 480x480 - 512x512 the amount of time it takes is worth the quality, if you drop below 480x480 it starts to get a bit shoddy.

    The 512x640 jump is not worth the extra time to run in my opinion.

    Overall quality workflow!

    Below timings on how long each resolution took.

    Default Setup -

    (512 x 640, 101 length) 784.79 seconds First Run

    (512 x 640, 101 length) 555.50 seconds Second Run

    (512 x 640, 101 length) 577.36 seconds Third Run

    Default Setup - Reduced Resolution 512 x 512

    (512 x 512, 101 length) 337.76 seconds First Run

    (512 x 512, 101 length) 274.98 seconds Second Run

    (512 x 512, 101 length) 250.95 seconds Third Run

    Default Setup - Reduced Resolution 480 x 480

    (480 x 480, 101 length) 197.73 seconds First Run

    (480 x 480, 101 length) 196.85 seconds Second Run

    (480 x 480, 101 length) 196.12 seconds Thirds Run

    Default Setup - Reduced Resolution 416 x 416

    (416 x 416, 101 length) 135.02 seconds First Run

    (416 x 416, 101 length) 135.92 seconds Second Run

    (416 x 416, 101 length) 134.86 seconds Thirds Run

    saturngfx
    Author
    Jan 24, 2025

    You are sure you're not running in CPU mode? (looking at the times...thats insane.). It should work better at any resolution. the only thing going on here is skipping the step of decoding/encoding, which takes forever...

    Hmm, upon considering. 16g memory card. naa, still seems far too long. if my 512x512 took 100 seconds, yours should be like maybe 1/3rd longer if things scale like that. You should be getting 130ish times. Something is off.

    saturngfx
    Author
    Jan 24, 2025· 1 reaction

    Thought a bit more (trying to snort coffee to wake brain up.) alright, so I suspect the workflow you might be comparing to might do a descaling process before it runs through the pipeline. This of course will make things go blazing fast. Something I considered doing and simply tell people to use larger sizes, then upscale it post refiner steps. I chose not to on this workflow because people can do that themselves (lots of people simply hate downscale/upscale. I don't mind, but best to leave out then pop in.)
    So, when you go in raw and large, your GPU might get used up and hit a bottleneck. once your GPU hits that like 97% used mark, it tends to get locked and takes forever. this might be whats going on with you.
    What I would do is downscale step before refiner. like, right out of the gate, then add a step after refiner to upscale back. You will lose some quality, but you will have videos before next christmas.

    AicushJan 24, 2025

    @saturngfx Yeah running in GPU mode, This was not a comparison take, this was just ran against yours, I was more than happy with the speed, I have tried pretty much most workflows that has come out, this I think gives me quality results for the time taken, I dont mind the waiting part as i usually run these whilst working X-D.

    I will look into your recommendations as I am trying to improve my ability on this! But I will be upgrading and getting a new pc soon so I wont have to worry about VRAM etc.

    --To confirm my gpu and ram are maxing out with the setting and that is likely the reason why its dragging on hence the massive speed increase when I drop the resolution--

    saturngfx
    Author
    Jan 24, 2025· 1 reaction

    @Aicush Thanks for the update. btw, image 2 video version 4 just dropped using this...if you use that then that should increase your speed there also. just...be warned. Flux is a beast.

    AicushJan 24, 2025

    @saturngfx Legend thanks so much for letting me know :-) Ill give this a go tomorrow morning and report my findings, seriously thanks for the awesome workflows!!

    Lil_WayneJan 25, 2025
    CivitAI

    I'm getting strange checkerboard artifacts on my videos, any idea why? it's like squares with blurred edges, thought it was the VA Decode (Tiled) but when previewing the output before the VAE decodes they still have the same issue

    saturngfx
    Author
    Jan 25, 2025

    Hmm, straight out of the box or can you remember the things you altered?

    Lil_WayneJan 25, 2025

    @saturngfx Out of the box I believe, I'll give it another whirl, would you mind sharing models you used specifically?

    saturngfx
    Author
    Jan 26, 2025

    @Lebofly fast model and the fastlora (at negatives). really thats it. you can add other loras of course and you can use any of the hunyuanvideo models.

    OFIAJan 31, 2025
    CivitAI

    How is the LORA, with diffusion pipe or other? formed, what are the necessary PC capacities?

    saturngfx
    Author
    Jan 31, 2025

    Lora? what LoRA? the fastvideo LoRA?
    When using the Fastvideo lora, if using the fast model, ensure it is a negative value. -0.3 or -0.4. otherwise keep it positive. if you use other loras, you can turn it down, but it helps speed up steps and gives more detail in general...resolves the weird blotchy blocky mess.
    As far as PC capacities? I assume you mean can it run on XYZ type question. I dont know...if you can run regular hunyuan video, you can do it here also. the only secret sauce is no steps inbetween the video process and the refinement process. no decoding just to encode again...otherwise, there really isn't anything special, just bypassing some unnecessary very long steps of unloading and reloading models. Also because it hits a refiner twice, you need less initial steps because the refiners work on the previous picture to fix issues, so...yeah, less steps, faster because no decode/encode...thats it really. everything else is vanilla hunyuan video.

    OFIAJan 31, 2025

    @saturngfx I mean how he trained the "Final V3! DeJanked Speed ​​​​​​Hack Hunyuan T2V Final Boss", with diffussion-pipe?

    OFIAJan 31, 2025

    @saturngfx I mean, how do you train your LORA, Whit diffusion-pipe?

    saturngfx
    Author
    Feb 1, 2025

    @OFIA I think you might be commenting in the wrong area. this isn't a LoRA, its just a workflow. I didn't train any loras, just slapped nodes together.

    yesitsfineFeb 4, 2025· 1 reaction
    CivitAI

    Amazing comfy kung fu! Thank you!

    Workflows
    Hunyuan Video

    Details

    Downloads
    713
    Platform
    CivitAI
    Platform Status
    Available
    Created
    1/24/2025
    Updated
    8/14/2026
    Deleted
    -

    Files

    v3FinalDejankedSpeedHack_v3.zip

    Mirrors

    HuggingFace (1 mirrors)