You paste your lyrics, then describe the music the way you would to a producer: genre, tempo, mood, the kind of voice, the instruments. The model composes and performs the whole thing and hands back a stereo audio file. If you want control over the shape, you can label your lines with tags like [Verse] and [Chorus], which tell it which part of the song each block of lyrics belongs to.
The difference from earlier song generators is stamina. Where most models drift into mush or repeat themselves after a minute, Music 3 was built for long form: verses stay verses, the chorus returns as a chorus, and a bridge actually sounds like a detour. Under the hood two models split the job, a larger one planning the song's structure while a smaller one renders the sound itself in stereo.