
Meta Muse Voice Transcribe is a real-time speech model from Meta's Superintelligence Labs that streams transcripts with speaker separation for up to 20 people at once, built for meeting notes, live captions, and contact-center automation where many voices overlap. Meta released it on September 1, 2026, and the same engine now powers voice dictation inside Meta AI and Muse Code.
Core Features
- Streaming ASR in 80-millisecond chunks with an adaptive per-word delay that balances speed and accuracy on the fly.
- Speaker separation for 20 or more speakers handled inside one model, tagging each passage A through Z with no extra system.
- Joint endpoint and sentence detection trained together with recognition, so turns and boundaries are marked automatically.
- Seventy-plus languages with code-switching and context, keyword, and language bias to sharpen proper nouns.
- Handles recordings longer than one hour end to end with no separate post-processing step.
- Top of the Artificial Analysis streaming speech-to-text leaderboard, posting a 3.1 percent English word-error rate.
Use Cases
- Multi-speaker meeting notes and live subtitles for webinars and conferences.
- Call-center and contact-center transcription at high volume.
- Journalists transcribing interviews where several people talk over each other.
Pricing
Meta prices the API at $0.18 per hour, or $3 per 1,000 audio minutes, which undercuts Cartesia, ElevenLabs, and Deepgram. It powers dictation inside Meta AI and is available through the Meta Model API; the weights are not released.
Our Take
Best for high-speaker-count, cost-sensitive transcription where price matters more than owning the model; the trade-off is that it is Meta-hosted only, with no self-host or open-weight option. Browse more AI audio tools.










