Breeze TTS 2 turns text into natural speech in English and Chinese, and it gives you three ways to shape the voice. You can clone a speaker from a short reference clip and its transcript, invent a brand-new voice by describing it in plain words ("a calm older narrator with a slight rasp"), or clone a voice and then steer its tone, pace, and emotion with a written instruction. You can even drop expressive events straight into the text, like (laugh) or (sigh), and the model performs them in place.
It currently ranks first among open-weight models on the Artificial Analysis TTS leaderboard, ahead of several proprietary systems. The model is also built for real-time use: on server hardware it reaches first audio in under 40 milliseconds and generates about three times faster than playback.