News
MiniMax Music 3: Five-Minute Songs, Open Weights, and What It Means for Studio Soundtracks
ยท RenderBob team
MiniMax Music 3 generates full songs up to five minutes from lyrics and a description, with open weights and 32kHz stereo output. Here is what that means for studio soundtracks.

ComfyUI already has a MiniMax Music Production Toolkit, a free node pack that turns the graph into a music production studio. This week's news is about what sits underneath it: MiniMax Music 3, an open-weight music generation model producing full songs up to five minutes long, driven by lyrics and a music description, with expressive vocals, evolving arrangements, and 32kHz stereo output.
The toolkit and the model solve different problems. The toolkit is the production layer, LLM-written prompts, genre presets, upscaling, cover art, that makes the model practical to direct without deep music-production expertise. MiniMax Music 3 is the underlying generative model that produces the audio. The open-weight release means a studio can now, if it chooses, run the model itself rather than only accessing it through a hosted node.
Full songs up to five minutes is a meaningful duration ceiling, long enough for most complete pieces of commercial music rather than a short loop or sting. Evolving arrangements suggests the model handles song structure, intro, build, chorus, bridge, rather than generating a static repeating pattern. Lyrics-plus-description as the input means a studio can direct both the words and the musical direction in one generation. 32kHz stereo is respectable for most delivery contexts, though FlashSR upscaling to 48kHz, part of the toolkit, remains the step that gets output to broadcast-acceptable quality for anything more demanding.
Open weights change the calculus: rather than relying solely on a hosted node and its per-call or platform terms, a studio with the hardware can run Music 3 on owned infrastructure, fine-tune it toward a house sound, or route it through the same local-iterate, cloud-burst pattern used for every other heavy generative job. "Open weights" describes access to the model, not automatically a blanket commercial-use clearance, so the specific licence terms are still the thing to check before shipping generated music in paid client work. Between the open model and the production toolkit built on top of it, ComfyUI now has a complete, ownable music pipeline sitting inside the same graph as everything else.
More from the blog
- Visual Dubbing Goes Mainstream: Prime Video Changes the Mouth, Not the Voice
On 9 September 2026, Prime Video launched AI lip-sync for the English dub of Maxton Hall. Human actors record the dialogue. The actors' mouths are regenerated to match.
- Suggestive Editing Arrives: Story-Aware Rough Cuts Move Into the Mainstream
From Eddie AI's story-structured assemblies at NAB 2026 to Premiere's AI Assistant and Resolve 21 search, editing tools are proposing the cut, not only cleaning the footage.