MiniMax Speech 2.8: AI Voices That Hesitate and Breathe

MiniMax ships Speech 2.8 with native sound tags for natural fillers, 10-second high-fidelity voice cloning, and studio-grade noise-free output.

Friday January 23, 2026 Source: MiniMax
TL;DR — Quick Answer

MiniMax released Speech 2.8, a text-to-speech model that models the imperfections of real speech — 'um,' 'uh,' breaths, and pauses — through native sound tags. It clones a voice from a 10-second sample with high fidelity, eliminates background noise and digital artifacts, and fixes cross-lingual accent bleed starting with Mandarin-Japanese.

Key Takeaways

MiniMax Speech 2.8: AI Voices That Hesitate and Breathe — AI news article illustration

MiniMax has shipped Speech 2.8, a voice model whose thesis is that human speech is imperfect — and that the imperfections are the point.

Teaching AI to Hesitate

Speech 2.8 introduces native sound tags: modeling colloquial fillers like “um,” “uh,” and “ah,” along with breaths, chuckles, and pauses. The result preserves the rhythm and emphasis signals that make speech feel alive rather than flattened — MiniMax’s demo has the model casually admitting mid-sentence that it isn’t human, punctuated by a (chuckle).

Cloning in 10 Seconds

The cloning pipeline was re-optimized for fidelity: a 10-second sample captures a speaker’s texture, breathiness, and pace — aimed at narration that sounds like a person, not an announcer.

Purity and Languages

The processing engine was re-engineered to eliminate background noise and synthetic distortion, and cross-lingual “accent bleed” — unnatural tones when a voice speaks a second language — is fixed starting with Mandarin-Japanese, with more languages promised.

Speech 2.8 is live on the MiniMax platform and minimax.io/audio.


Frequently Asked Questions

What is MiniMax Speech 2.8?

Speech 2.8 is MiniMax's text-to-speech model released January 2026, focused on vocal authenticity — natural hesitations, breaths, and fillers — plus high-fidelity 10-second voice cloning and studio-grade clarity.

What are sound tags in Speech 2.8?

Sound tags are inline markers like (chuckle), (breath), or (laughs) that Speech 2.8 natively models, preserving the natural rhythm, pitch, and pauses of human dialogue instead of producing flattened, 'too perfect' speech.

How good is Speech 2.8 voice cloning?

Speech 2.8 clones a voice from a 10-second sample, capturing the speaker's texture, breathiness, and speaking pace — MiniMax positions the result as capturing a person's 'vocal fingerprint.'

This article is based on the official announcement from MiniMax . Read the original for full technical details.

Related Articles

Back to all news