VibeVoice
A Microsoft speech synthesis model that generates long dialogues, up to 90 minutes and four distinct voices in a single pass, built for the podcast format. Microsoft pulled the generation code from its repository in September 2025 after uses that went against the stated intent; the weights are still on Hugging Face and a community fork keeps the code alive. Trained on English and Chinese only.
Strengths
- Up to 90 minutes and four distinct voices in a single generation, built for dialogue
- MIT licence and weights still downloadable from Hugging Face, no subscription
- The Realtime 0.5B variant stays in the Microsoft repository, with first audio in about 300 ms
Limitations
- English and Chinese only: the official card warns other languages may come out unintelligible
- Generation code pulled by Microsoft: you depend on a community fork with no accountable publisher
- Presented as a research model, with commercial use discouraged without further testing
- Hardware needs not quantified by the official pages, to be measured on your machine
Best for
- Testing a multi-voice English podcast format before booking a studio
- Judging the generated-dialogue demos going around on the evidence, before promising the format to a client