TEST RECORD · ← all field reports

FIELD REPORTS/FIELD NOTE/ISSUE #89

Turns lyrics and a style prompt into a full five minute song

MiniMax keeps shipping audio models at a startling pace, and this one goes after the hardest problem in music generation: keeping a song coherent past the first minute. Most generators hand you a thirty second loop; this one holds an arrangement together from intro to outro.

MODELMiniMax-Music3
PUBLISHEDAugust 13, 2026
READ TIME3 min
TESTED BYNeural Expedition
CATEGORYFIELD NOTE

Field notes

01What it does

You paste your lyrics, then describe the music the way you would to a producer: genre, tempo, mood, the kind of voice, the instruments. The model composes and performs the whole thing and hands back a stereo audio file. If you want control over the shape, you can label your lines with tags like [Verse] and [Chorus], which tell it which part of the song each block of lyrics belongs to.

The difference from earlier song generators is stamina. Where most models drift into mush or repeat themselves after a minute, Music 3 was built for long form: verses stay verses, the chorus returns as a chorus, and a bridge actually sounds like a detour. Under the hood two models split the job, a larger one planning the song's structure while a smaller one renders the sound itself in stereo.

02How to try it

Open the demo Space, paste a verse and a chorus, and describe the track in one line, something like "slow soul ballad, warm piano, female voice, 70 BPM". The test worth running: listen to whether the second chorus comes back subtly changed instead of copy-pasted. That is the long-form trick almost every other generator misses. If you want it on your own machine, the weights are on Hugging Face; full precision wants a serious GPU with 24 GB of memory, though it can squeeze into 8 GB with offloading if you accept slower runs.

03Caveat

The section tags steer rather than command: it can gloss over a [Bridge] you asked for, so regenerate if the structure matters. Vocal quality follows the description; vague briefs get generic voices. And the license is a Creative Commons variant rather than a standard permissive one, so read it before putting a generated track in commercial work.

04What you can do with it

  • Demo a song idea with full vocals before spending money on studio time.
  • Score a podcast intro or video outro with an original track instead of stock music.
  • Hear your poem or lyric sheet performed, to find a melody direction worth keeping.
  • Prototype a jingle against a precise brief: tempo, mood, and instrumentation.
  • Generate a reference track that shows a collaborator the vibe you are after.

Try the demo

View model page

Read this issue on neuralexpedition.com →