VoxCPM
A speech synthesis model that clones a voice from a few seconds of audio, across several languages, and accepts a written description to shape the result (deeper, calmer, stronger accent). Everything runs on your machine, no API key and no subscription, but you need an NVIDIA card with roughly 8 GB of video memory. Output varies between runs: two or three attempts are often needed to land the voice you want.
Strengths
- No cost per generated minute, where services bill by audio volume
- Source audio never leaves your machine, which matters when it is a client voice
- Tuning through a written description, more approachable than a set of technical sliders
Limitations
- An NVIDIA card with roughly 8 GB of video memory is required in practice
- Irregular output: often two or three generations are needed for the same sentence
- Consent from the cloned person stays your problem, the model licence does not address it
Best for
- Narrating your own content without re-recording at every script fix
- Prototyping a voice interface without opening an account or exposing audio to a third party