Ultra-realistic text-to-speech. Powered by Sonus TTS.
Voice AI has rapidly evolved, yet many enterprise systems still sound unmistakably robotic. Today, we are releasing Sonus, a massive multilingual zero-shot text-to-speech (TTS) model that scales natively to over 45 languages, delivering state-of-the-art intelligibility, emotional prosody, and instantaneous voice cloning.
Highlights.
- Zero-Shot Voice Cloning: Synthesize high-fidelity voice profiles instantly from just a three-second reference clip without any additional fine-tuning.
- Zero-Shot Multilingual Capability: Scales seamlessly to over 45 languages using a massive 581,000-hour training corpus.
- Single-Stage Pipeline: Eliminates the traditional text-to-semantic-to-acoustic bottleneck by directly generating multi-codebook acoustic tokens, drastically reducing inference latency.
- Superior Intelligibility: Achieves industry-leading Word Error Rate (WER) and Similarity (SIM-o) metrics, significantly outperforming commercial models like ElevenLabs and MiniMax.
Model Capabilities.
Traditional TTS architectures are complex and brittle. They typically rely on a cascading two-stage pipeline: first mapping text to semantic tokens, and then converting those semantic tokens into an acoustic waveform. This disjointed process inherently bottlenecks speed, increases Time-To-First-Byte (TTFB) latency, and severely degrades audio quality across language boundaries.
Sonus completely rewrites this paradigm. At its core, Sonus leverages a highly novel diffusion language model-style discrete non-autoregressive (NAR) architecture. By bypassing the semantic layer entirely, Sonus directly maps input text straight to multi-codebook acoustic tokens in a single, fluid generative pass. The result is a model that is profoundly faster (capable of TTFB under 150ms), significantly more natural, and highly resistant to the mispronunciations commonly found in older TTS models.
This dramatically simplified approach is facilitated by a full-codebook random masking strategy for highly efficient training, and robust initialization from a pre-trained LLM. We trained Sonus on an unprecedented 581,000-hour multilingual dataset, curated entirely from open-source data. This extensive data diet allows Sonus to natively understand and output 45 unique languages. It can gracefully handle code-switching—seamlessly transitioning between English, Spanish, Mandarin, and more within the same sentence—while perfectly preserving the speaker's original voice timbre, breath patterns, and accent.
In rigorous human evaluations, Sonus achieved a 3.80 SMOS score—frequently surpassing the original human ground-truth audio in listener preference tests due to its lack of microphone artifacts and perfectly regulated pacing.
| Model | WER ↓ | SIM-o ↑ |
|---|---|---|
| Sonus TTS | 2.850 | 0.830 |
| MiniMax | 3.774 | 0.766 |
| ElevenLabs | 10.950 | 0.655 |
*Note: Lower Word Error Rate (WER) is superior. Higher Speaker Similarity (SIM-o) is superior. Subjective evaluation measured 3.80 SMOS across Chinese and English.
Use Cases.
- Real-Time Conversational Avatars: Thanks to its incredibly low Time-To-First-Byte latency and single-pass architecture, Sonus is the ideal backend for real-time customer support agents, virtual companions, and interactive gaming NPCs.
- Automated Video Localization: Combined with translation models, Sonus can take a video recorded in English, clone the speaker's voice instantly, and dub the entire video into 45 different languages with perfect emotional continuity and accent control.
- Dynamic Accessibility Tools: Integrate Sonus into mobile applications and web browsers to provide high-quality, non-robotic screen reading and audio-book generation on the fly, drastically improving the experience for visually impaired users.
Get started.
Sonus TTS is available today in SAGEA Studio and APIs, and powers remote coding agents and Work mode on the Pro, Team, and Enterprise plans.
It is available for prototyping and production deployment, hosted on SAGEA-accelerated endpoints on platform.sagea.space.
Build the future of agentic systems with us.
We're hiring across research, engineering, and product to push agentic systems further. See our open roles.

