You hand Audex a piece of audio and a question, and it answers like a chat model that happens to have ears. Ask for a transcript and you get one. Ask what a foreign-language speaker said and it translates into English. Ask what is happening in a field recording and it describes the scene: footsteps, rain, a door closing. Flip the direction and it generates audio instead: type a script and it reads it aloud, or produces a sound effect from a description.
The difference in practice is that these were separate tools until now. A transcriber could not answer follow-up questions about the meeting it just transcribed; a text-to-speech model could not hear. Because Audex is one model, you can stay in a single conversation: transcribe an interview, ask which speaker sounded uncertain, then have it read a summary out loud.
It comes from NVIDIA's Nemotron lab as a compact 2B-parameter model, the small sibling of a much larger 30B version, and it keeps the ordinary chat and reasoning skills of its text backbone, including a step-by-step thinking mode.