VibeVoice

A Microsoft speech synthesis model that generates long dialogues, up to 90 minutes and four distinct voices in a single pass, built for the podcast format. Microsoft pulled the generation code from its repository in September 2025 after uses that went against the stated intent; the weights are still on Hugging Face and a community fork keeps the code alive. Trained on English and Chinese only.

Strengths

Limitations

Best for

Official site

View on Coeurdar