Workflows are able to +4K Text-to-Image and also Multi-Reference Image-to-Image.
Unless there is a major update all changes will be listed in their own sections below.
Workflows v6.0:

v6.0 Basic:
- For DaSiWa V2
- No RES4LYF custom nodes needed.
- Has a Beta Sampling Scheduler node using the Beta57 Alpha and Beta's. When switching on Beta57, Denoise and Scheduler are disabled.
Node Requirements:
- ComfyUI-MM-1Frame
- ComfyUI-Easy-Use
- ComfyUI-KJNodes
- rgthree-comfy
v6 Updates:
- Moved sampler step selection into the prompt group.
- Added notes for suggested model V2 sampler settings.

v6.0 Standard:
- For the pulled DasiwaMinimaxH3_dasiwaREF2VAHybridV1 model or V2.
- Two samplers: Sampler Custom, and Clown Shark Sampler from RES4LYF, swap between them with the Fast Groups Muter.
- ClownShark group has a chain sampler for the second half of total step count, added a bit more sharpness.
- Several RES4LYF option nodes: Detail Boost, Implicit Steps, Momentum and Sigma Scaling. Notes from the official RES4LYF example workflow for explanations of what they do are included, along with some quick notes on my testing of them.
Node Requirements:
- ComfyUI-MM-1Frame
- ComfyUI-Easy-Use
- ComfyUI-KJNodes
- RES4LYF
- rgthree-comfy
v6 Updates:
- Added a chain sampler to ClownShark group, first half of steps to S1, second half to S2. Added a bit more sharpness without having to change any settings.
- Prompt, resolution selection and sampler step count now grouped.
- Added notes for suggested model V2 sampler settings.The DasiwaMinimaxH3_dasiwaREF2VAHybridV1 model Darksidewalker pulled is a must, every setting in the standard workflow is tuned for this exact model, if you use anything else you're going to have to find your own settings. There is none better.
I tried most of the samplers and landed on what feels like the sweet spot for Clown Shark: diag_implicit/pareschi_russo_2s with Beta57. It’s a little slower, but totally worth it IMO.
This workflow is pretty much a set-and-forget workflow, I found settings that give a nice balanced image so nothing "should" need adjusted, still tweak to your hearts content as there is plenty to adjust if you want.
The Big Three:
The three biggest contributors to image quality is ETA in Sampler Group and Implicit Steps, Lying Strength with Lying Start Step in the Options Group
Adjusting the two down in Lying will produce a more grainy but sharper image, up will make them smooth and blurrier. ETA in Sampler Group does the same thing.
Implicit Steps can have multiple effects, sometimes 2 looks better sometimes 1 will, varies from prompt to prompt and seed to seed. Switch to 2 when you get a image you like to save time. 2 is also good for busy scenery.
Quick notes:
The Beta57 scheduler is a must for image quality.
For extra detail, run a SeedVR detail pass with the same resolution as the image (no upscaling).
I’d keep the LoRA strength light as these were made for video.
Pretty much everything you can tweak is sitting on the subgraphs. Fair warning: cranking the resolution too high will probably throw errors.
Reference Image input size has a great deal to do with rendering time, the bigger the image the slower the renders. If you have a slower card set this lower.
About lighting, it is very hard to tame as MM always wants to shine a spotlight on the characters. Want a dark dimly lit atmosphere, forget about it, at least I haven't found a way yet after many prompts of lighting techniques with very few exceptions. The closest I can get to good lighting is using this phrase at the beginning and adapt the second half per scene: "Ray tracing. Make no studio, set, self emissive character, spot, vanity or hero key lights. Soft, diffused, bounced, natural daylight window key lighting from all sides." If anyone finds a way to get pinpointed lighting that you want let me know, I will be grateful for it.
Some things I learned on this journey and a ongoing blog. MM in this format loves, loves and loves some more the details. If you don't explain exactly what you want in a detailed way you will get bland output. For example just saying a woman is moaning won't get you much, putting things in like the state of her face such as open mouth, furrowed eyebrows and facial tension will get you much further, this applies to everything, even more when it comes to lighting.
Nodes and models you will need to have installed, the list below is the bare minimum.
custom_nodes:
ComfyUI-MM-1Frame (Basic & Standard)
ComfyUI-Easy-Use (Basic & Standard)
ComfyUI-KJNodes (Basic & Standard)
RES4LYF (Standard)
rgthree-comfy (Basic & Standard)
## Model Links
diffusion_models:
text_encoders:
vae:
vae_approx (TAE):
loras/h3:
## Highly Recommended
text_encoders:
