MiniMax has shipped Speech 2.8, a voice model whose thesis is that human speech is imperfect — and that the imperfections are the point.
Teaching AI to Hesitate
Speech 2.8 introduces native sound tags: modeling colloquial fillers like “um,” “uh,” and “ah,” along with breaths, chuckles, and pauses. The result preserves the rhythm and emphasis signals that make speech feel alive rather than flattened — MiniMax’s demo has the model casually admitting mid-sentence that it isn’t human, punctuated by a (chuckle).
Cloning in 10 Seconds
The cloning pipeline was re-optimized for fidelity: a 10-second sample captures a speaker’s texture, breathiness, and pace — aimed at narration that sounds like a person, not an announcer.
Purity and Languages
The processing engine was re-engineered to eliminate background noise and synthetic distortion, and cross-lingual “accent bleed” — unnatural tones when a voice speaks a second language — is fixed starting with Mandarin-Japanese, with more languages promised.
Speech 2.8 is live on the MiniMax platform and minimax.io/audio.