YuE2-3B
Frontier music generation with editable scores
Developed by Multimodal Art Projection, YuE2 combines lyrics and musical style instructions to synthesize songs with singing and instrumental accompaniment.
- Score control: compose melody and harmony, melody alone, or generate without symbolic planning. An ABC score can be supplied directly.
- Revisions and covers: use an existing score with revised lyrics or style, or edit its melody and harmony before rendering again.
- Local generation: the upstream setup produces 48 kHz stereo audio on a 24 GB GPU without quantization.
- Pipeline access: separate interfaces expose planning and synthesis, with text guidance controls.
ComfyUI files
Packaged for ComfyUI by Comfy-Org, from the original M.A.P models.
- BF16 checkpoint:
yue2_3b_bf16.safetensors(7.80 GB), the bfloat16 build. - INT8 checkpoint:
yue2_3b_int8_convrot.safetensors(3.96 GB), the smaller quantized build. - SheetSage2 BF16 audio encoder:
sheetsage2_bf16.safetensors(1.39 GB), an optional audio-to-score model for transcribing an existing recording before creating a cover. It is not required to generate a song from text and lyrics.
Local setup
Choose one checkpoint and place it in ComfyUI/models/checkpoints/. For audio-to-score workflows, place SheetSage2 in ComfyUI/models/audio_encoders/. See the ComfyUI package and SheetSage2 release for setup details.
Weights license: CC BY-NC 4.0.
Original release · Usage examples · Demos · GitHub · Discord
Description
Comments (25)
YuE2-3B is to music what Minimax H3 is to video.
Can get cutting edge, SAAS quality for free on local hardware and it'll run on a potato.
If Civitai and going to host this music model they really need to allow uploads of .MP3 and/or videos longer than 245 seconds on this page!
This free site is good and can quickly add an image to an .mp3 for free so you can upload it to Civitai: https://www.neuralframes.com/tools/audio-to-video
@J1B Thanks, just what I was looking for. 👍
That would help for my Music Focus collection https://civitai.com/articles/35308/the-music-focus-collection
Although mp3? 🧐
Don't forget FLAC, Opus, m4a (AAC).
@bno84 This site allows you to convert audio formats to FLAC and M4A. For instance, NeuralFlame supports M4A audio generation. The site also includes basic editing features, such as adding fade-ins and fade-outs, which is highly useful since AI-generated music files often end abruptly. 😂Forgot the link:https://mp3cut.net/
Ofcourse downloading Audacity or som other apps is the best way. But that site is easy and quick to use, for fast work.
@J0hn_D0e The conversion isn't an issue (LLMs and ffmpeg exists). But allowing it to be uploaded to Civitai. MP3 is imo completely obsolete so if audio uploads get supported, I think they should consider:
- Opus as the gold standard for lossy
- M4A/AAC for broad compatibility lossy
- FLAC for lossless
@bno84 Sorry I misunderstood your initial post. I agree, lossless formats should be a option. But the cost of hosting them would probably need to be covered somehow.
@J0hn_D0e Yes the size of lossless is gnarly. But the site also has video. My 15 second video clips tend to be about the same as a 2½min flac track ~30MB or so. But with Opus it's 10x smaller and indistinguishable. So maybe they should put a length-cap on the lossless formats if they want to allow but keep costs down. Say 30s for lossless?
Can't load the model without a specific node that git-pulls. HELL. NO.
I don't want the BF16 version and I can't load the int8_convrot with any normal Node?
Easy >>> Shift + Delete
Have you never installed a custom node for Github before? or are you saying there is a security issue with that one?
I have over 200 custom node packs and never really had a problem, although it is not unheard of/without any risk.
i agree, but generally will just go check out the actual node on github or whatever and then manually install it. that's all its doing anyway
@J1B I have absolutely no trust in that blackbox of a node that forces downloads.
But please, tell me the nodes you're using to load the int8_convrot model and the audio encoder separately...
@makiaeveli same question.
@GlowingGuardianGirl Who created the models don't have to obey on what you desire but on their. If they created an AIO model there will be a rerason and maybe if you read into their repository you'll understand why. INT8 convorot is natively supported by ComfyUI. If you don't like it noone force you. Just a little respect on other work especially if they can let you use it for free.
@GlowingGuardianGirl I'm just using the old school checkpoint loader and it can load the int8 and bf16 models, not sure what I did for the audio encoder, I think that was a custom node.
@J1BIf you use the template from Comfyui u'll understand. No need to use a custom node for the audio encoder. Audio encoder is for remixing audio, not to generate a new one.
One question, can this music model understand music theoy, the last music model like Lyrica 3 was really bad at translating my sheet music accuractly no matter how much prompt enginnering went though. Can this model understand music theoy?
Yes, kind of. Check out their website.
You actually get scores as an artifact, you can have an LLM edit the scores for you (they have a skill) and re-render the audio from it. Or have the model complete a short transcription.
You can feed your own music into the Sheetsage2 node, and it will generate an ABC music score. Alternatively, you can simply use any existing ABC score or generate one using an LLM.
Also, when generating, leave the "Style" field blank if you need to adhere strictly to the ABC score, and specify your vocals in the "Vocal" field.
@BibaBobacivitai There are gotchas, they require a simplified ABC dialect.
2 tracks + chord progression, monophonic. Your style prompt needs to influence whether it's likely that monophonic melody becomes power chords, single notes, etc.
This model has issue with gen without vocals. The best that I could do was 50/50.
If anyone has good prompts / workflows to get i.e. techno / electro / synth without lyrics please do share, thanks in advance.
PS. same goes for mula and mmMusic.
There is a instrumental only lora here: https://huggingface.co/Mothersuperior/YuE2-instrumental-cot-full-loras
