MiniMax Speech
MiniMax's text-to-speech model (version 2.8, released 23 January 2026), used from the minimax.io web studio or through the API. You type text, it returns a voiceover; give it 10 seconds to 5 minutes of your voice and it clones it. Two tiers: HD for quality, Turbo for cost. Around forty languages, plus interjection tags (laugh, sigh, breath) to break the read-aloud tone.
Strengths
- Published, readable pay-as-you-go pricing, with a cheaper Turbo tier for volume
- Cloning from 10 seconds of audio, straight from the web studio
- Interjection tags (laugh, sigh, breath) that break the recited tone of a voiceover
Limitations
- No documented consent check for cloning
- Everything goes through the cloud: your voice and your text leave your machine
- Output to judge on your own script: the hype around it is not a measurement
Best for
- Content creators producing regular voiceovers who want to cut the cost
- Product teams prototyping narration or a voice agent before picking a provider