
MiniMax Music 3.0 is an open-weight AI music generator from MiniMax for creators who want complete, vocal-backed songs rather than short loops. Released on 13 August 2026 with weights published on Hugging Face, it generates up to five minutes of arranged music from a text prompt and optional lyrics in a single pass.
Core Features
- A hybrid dual-model architecture: an 8B global LLM for song structure plus a 0.6B local LLM for vocal and timbre quality, with a 2.4B flow module and VAE.
- Full-song generation in one inference: intro, verse, chorus, bridge, instrumental breaks, and outro stay coherent for the full five minutes.
- Structured control through caption tags and section labels like [Verse] and [Chorus], plus an instrumental-only mode and a music-cover mode.
- Studio-grade 32 kHz stereo WAV output with a rebuilt vocal engine that removes the digital artifacts common in earlier AI singers.
- Local deployment via diffusers and an official ComfyUI node, with streaming output for low-VRAM machines around 8 GB.
Use Cases
- Producing royalty-free theme music for YouTube videos, games, and podcasts from a short creative brief.
- Exploring arrangements and styles before booking studio time.
- Building music features into apps via the MiniMax API, which offers a free tier for light use.
Pricing
MiniMax Music 3.0 is free to self-host under the MiniMax Music 3.0 Community License. Through the MiniMax API, a free tier (music-3.0-free) allows about three requests per minute, while paid generation costs roughly $0.15 per five-minute song. Enterprise and higher-volume plans are available on request.
Our Take
Best for indie creators and developers who want a production-grade, locally runnable music model instead of a closed SaaS. The trade-off is that the 8 GB streaming mode lowers fidelity, and you need real GPU memory for the full-quality weights. See our AI audio tools for comparable options.










