Ready-made ComfyUI workflows for Verboa Image 1.0 (ERNIE-Image fine-tune, 8 steps, photoreal, adults only). Drag a .json onto ComfyUI and press Run.
Included
verboa-image-1.0-hires.json: fp8 safetensors (RTX 40/50) with a built-in 2.3 MP hires pass (upscale, then a second 8-step sampler at denoise 0.42).
verboa-image-1.0-gguf-hires.json: the same hires graph for the GGUF files (needs the ComfyUI-GGUF node).
verboa-image-1.0.json and verboa-image-1.0-gguf.json: the plain 1 MP originals.
Settings baked in
8 steps, CFG 2.0, euler + simple, ModelSamplingSD3 shift 4.0, empty negative prompt, FLUX.2 VAE, Ministral 3B text encoder (CLIP type flux2).
Hires pass: ImageScaleToTotalPixels 2.3 MP (steps of 16), VAE encode, 8 steps at denoise 0.42 with the same prompt. About 60 seconds per image on a laptop RTX 5090.
You also need
The model: Verboa Image (fp8 / bf16 or GGUF Q8_0) from the Verboa Image model page; text encoder ministral-3-3b and VAE flux2-vae from Comfy-Org/ERNIE-Image.
For GGUF: the ComfyUI-GGUF custom node (Unet Loader (GGUF)).
18+ only. Fictional adults only: no real people or look-alikes.
Description
Verboa Video: the video stage of our pipeline as ComfyUI nodes. One node turns a picture into a video with sound (LTX-2.5, two stages), Verboa Act Prompt writes the first-frame prompt and the motion prompt of a clip from a seed, and Verboa Frame Sheet makes a strip of frames to check before you post. Unzip into ComfyUI/custom_nodes/, restart ComfyUI, put the five LTX-2.5 files in their models folders (the list is in the README inside the zip), and load a workflow: verboa-video-from-text.json (Act Prompt, Verboa Image, Verboa Video) or verboa-video-i2v.json (any first frame). Tested on ComfyUI 0.37.4; a 6 s 768x1152 clip takes about a minute on a 24 GB card. The clips on our profile are made this way. LTX-2.5 has its own license: label what you post as AI-generated.

