You give it text and it speaks the line back in a natural, expressive voice. It ships with 20 preset voices, and it can pick up a new voice from a short reference: give it about ten seconds of someone speaking and it reads your text in that voice, in any of its nine languages: English, French, Spanish, German, Italian, Portuguese, Dutch, Arabic, and Hindi.
The practical difference is that it is built for live conversation, not just narration. It streams audio as it generates, with the first sound arriving in a fraction of a second, which is what lets a voice agent answer a caller without an awkward pause.
Under the hood it is a 4 billion parameter model built on Mistral's small Ministral language model, and the weights run on a single graphics card with 16 GB of memory. Output is clean 24 kHz audio.