Voicebox
Open source voice studio running entirely on your machine: your voices, samples and captures never leave it. It dictates anywhere in the system through a hotkey (local Whisper transcription), clones a voice from a short reference sample, and makes it speak through 7 speech engines, including Qwen3-TTS, Chatterbox and Kokoro, across 23 languages. A built-in MCP server plugs it into Claude Code, Cursor or Windsurf. Filler words get cleaned by a local Qwen3. Over 53,000 GitHub stars as of 13 September 2026, and a free alternative to ElevenLabs.
Strengths
- 100% local: your voice and transcripts never leave the machine
- Native MCP for Claude Code, Cursor, Windsurf, Cline, VS Code
- Integrated LLM cleanup (local Qwen3): "ums huhs" cleaned before paste
- Voice cloning from a short sample, 7 speech engines, 23 languages
Limitations
- Runs without a graphics card, only slower: speed comes from Apple Silicon or a GPU (CUDA, ROCm, DirectML, Intel Arc)
- More steps than an online service: models to download, and on Linux a build from source, since only macOS and Windows get an installer
- No cloud sync: if you change machine, you start over
Best for
- Devs and creators dictating to Claude Code wanting privacy and zero ongoing cost
- Podcasters and content creators wanting voice cloning without cloud dependency