This workflow uses the MOSS-TTS-v1.5 model (OpenMOSS-Team) to generate
high-quality Persian (Farsi) speech, patched with 4-bit quantization
(via bitsandbytes) so it can run on 8GB VRAM GPUs (e.g. RTX 3060 Ti).
The unquantized model (8B params, BF16) normally requires ~16GB VRAM.
Built on:
- ComfyUI node: comfyui-moss-tts by richservo
(github.com/richservo/comfyui-moss-tts)
- Model: OpenMOSS-Team/MOSS-TTS-v1.5 (Apache 2.0 license)
My contribution: added an optional 4-bit (NF4) quantization option to
the Model Loader node, which the original implementation did not
support. This patch has also been submitted upstream as a Pull
Request to the original repository.
⚠️ Disclaimer and Terms of Use:
This workflow is shared strictly for educational, research, and
entertainment purposes. Any use of it — particularly the voice
cloning capability — must be done with the consent of the voice's
owner and in compliance with the laws of the user's country. Please
do not use this tool for impersonation, fraud, or any unethical or
illegal purpose.
The creator of this post assumes no responsibility or liability for
any misuse, unethical use, or illegal use of this workflow. Full
responsibility for how it is used rests entirely with the user.
Description
Initial release: MOSS-TTS-v1.5 workflow patched with optional 4-bit
(NF4) quantization to run on 8GB VRAM GPUs.
Comments (2)
Среднечок есть уже давно получше. Это выдаёт роботизацию уже на старте. Для канала youtube я бы не использовал, в фильтр улететь легко можно.
Спасибо за отзыв! Согласен — по естественности звучания (особенно
с 4-битной квантизацией) уступает некоторым платным решениям, это
осознанный компромисс ради работы на 8GB GPU. Но пока это один из
немногих бесплатных open-source TTS с реальной поддержкой персидского
языка. Для YouTube соглашусь, лучше что-то понадёжнее.
