Skip to content

Latest commit

 

History

History
27 lines (24 loc) · 9.23 KB

File metadata and controls

27 lines (24 loc) · 9.23 KB

TTS Engines Reference Tables

Engine Comparison

Engine Isolation Models Size TTS SRT VC ASR Sound Effects Training License Special Features Languages
F5-TTS Main Base, v1, E2TTS + 8 lang models ~1.2GB each CC-BY-NC-4.0 Targeted Word/Speech Editing, Speed control 10
ChatterBox Main EN, DE×3, IT, FR, RU, HY, KA, JA, KO, NO ~4.3GB MIT Expressiveness slider 10
ChatterBox 23L Main v1, v2, v3, Vietnamese (Viterbox), Egyptian Arabic (oddadmix) ~4.3GB MIT V1, V2, and V3 official checkpoints, Emotion tokens (v2; currently ineffective), V3 skips the legacy alignment analyzer and trims the final token artifact 25
VibeVoice Shared 1.5B, 7B, KugelAudio-0 (7B), kugel-2 (7B), Hindi-1.5B/7B 5.4GB / 18GB MIT (research-only per model card) 90-min long-form, Native 4-speaker (Base models), Multilingual (KugelAudio variants), 4-bit quantization 27
Higgs Audio 2 Shared 3B ~9GB Boson Higgs Audio 2 Community License 3 multi-speaker, CUDA graphs (55+ tokens/sec) 5
Higgs Audio v3 Main 4B ~8GB Boson Higgs Audio v3 Research and Non-Commercial License Native inline emotion/style/prosody/SFX tags, Zero-shot voice cloning, 100+ language support 100+
IndexTTS 2 / 2.5 Main IndexTTS-2, IndexTTS-2.5 ~4.7GB / ~5.49GB bilibili Model Use License Emotion Control: 8 vectors, Text as reference, Audio as reference, IndexTTS-2.5 official internal feature-duration scaling (not prosody planning), IndexTTS-2.5 pronunciation annotations 5
CosyVoice3 Main 0.5B, 0.5B-RL ~5.4GB Apache-2.0 Paralinguistic tags 4
Qwen3-TTS Shared 0.6B, 1.7B (CustomVoice/VoiceDesign/Base) ~3-6GB Apache-2.0 Voice design, ASR (Automatic Speech Recognition) 10
Granite ASR Main granite-4.0-1b-speech, granite-speech-4.1-2b, granite-speech-4.1-2b-plus ~4.6GB Apache-2.0 Native speaker attribution / diarization (plus model variant), Native word-level timestamps (plus model variant), ASR (Automatic Speech Recognition), Custom timestamps/SRT via reused Qwen forced aligner, Speech translation (experimental), Optional forced aligner auto-routed through shared legacy T4 runtime 6
Step Audio EditX Main 3B LLM + CosyVoice ~7GB Apache-2.0 (verify before commercial use) Second Pass Speech Editing Node: 14 emotions, 32 speaking styles, Paralinguistic effects, Selectable main, shared, or dedicated Python runtime (shared Transformers 4 runtime recommended) 4
Echo-TTS Main echo-tts-base + fish-s1-dac-min ~5.3GB + ~1.8GB CC-BY-NC-SA-4.0 Diffusion-based (~30s best), Force Speaker KV (speaker drift control) 1
Fish Audio S2 Pro Main S2 Pro 4B / FP8 ~10.3GB / ~8.0GB Fish Audio Research License Free-form sub-word emotion/prosody tags, Native multi-speaker and multi-turn dialogue with dynamic speaker references, Zero-shot voice cloning, Optional per-segment custom character switching, Configurable 4K-32K native context with reduced KV-cache VRAM, Optional community FP8 weight-only checkpoint with BF16 activations, Optional on-the-fly BitsAndBytes INT8/NF4 for the official checkpoint 80+ languages
Dots TTS Main dots.tts-base, dots.tts-soar, dots.tts-mf ~6GB Apache-2.0 Official auto language detect / language control, SOAR and MeanFlow distilled variants 19
DramaBox Main DramaBox 3.3B ~16.4GB LTX-2 Community License Expressive scene prompting and stage directions, Native and SRT-aware duration targeting, Official duration-aware long-form chunking with scene-prefix preservation, Optional 10-second zero-shot voice reference, CFG negative prompt with per-segment switching, Explicit generation/reference durations, CFG rescale control, and optional Perth watermark, Experimental staged and sequential strategies for lowering peak VRAM; official LTX FP8-cast storage, Optional official torch.compile path, Official audio-branch IC-LoRA training workflow 1
OmniVoice Main OmniVoice ~3.7GB Apache-2.0 Inline non-verbal tags and pronunciation overrides, Reference-free voice design, 600+ language support, Upstream long-form chunk orchestration 600+
MOSS-TTS Main Local 1.7B, Delay 8B v1.5/1.0, LAION Voice Acting 8B community fine-tune, VoiceGenerator 1.7B, SoundEffect 8B v1, TTSD 8B ~8.5GB tokenizer + ~6.1GB/17GB/18GB model Apache-2.0 Reference-free voice design with MOSS-VoiceGenerator, Native 1-5 speaker TTSD dialogue, 31-language generation with MOSS-TTS-v1.5, Optional LAION community 8B voice-acting fine-tune, Config-based discovery of compatible local MOSS full checkpoints, Prompt-only sound-effect generation with MOSS-SoundEffect v1, Long-form generation (TTSD/Delay), Duration token hint, Local/Delay/TTSD variants, Initial integrated LoRA training workflow (Delay 8B) 24
MOSS-SoundEffect v2 Main MOSS-SoundEffect-v2.0 ~11.2GB Apache-2.0 Durations up to 30 seconds, Native negative prompting, CFG, flow shift, and diffusion-step controls, Prompt-only text-to-sound generation, 48 kHz mono output, Seeded generation 2
RVC Main Community .pth 100-300MB MIT (framework); community models vary Real-time VC, Integrated training workflow, Pitch shift (±14), 6 HuBERT models, Language-independent Any

Isolation column: Main runs in the main ComfyUI environment. Shared uses a shared secondary runtime reused by multiple engines. Dedicated uses an engine-specific secondary runtime.