Voicebox
Open-source local-first voice studio: system dictation with local Whisper (MLX Apple Silicon, PyTorch CUDA/ROCm/CPU), 3-second voice cloning, 5 TTS engines, 23 languages. Native MCP for Claude Code, Cursor, Cline. Cleanup of 'ums huhs' via local Qwen3. 28K+ GitHub stars, free ElevenLabs alternative.
Strengths
- 100% local: your voice and transcripts never leave the machine
- Native MCP for Claude Code, Cursor, Windsurf, Cline, VS Code
- Integrated LLM cleanup (local Qwen3): "ums huhs" cleaned before paste
- 3-second voice cloning plus 5 TTS engines plus 23 languages
Limitations
- Requires Apple Silicon or CUDA / ROCm GPU for decent perf
- Initial setup more complex than a SaaS (download models, Qwen3, MLX)
- No cloud sync: if you change machine, you start over
Best for
- Devs and creators dictating to Claude Code wanting privacy and zero ongoing cost
- Podcasters and content creators wanting voice cloning without cloud dependency