// 9TO5MAC — MOBILE & WEB
Meta launches Muse Voice Transcribe for real-time voice dictation on Mac
Meta is launching Muse Voice Transcribe, its first real-time audio perception model. It brings multilingual, streaming transcription to Meta AI for Mac, Muse Code, and developers through the Meta Model API.
Meta Superintelligence Labs says Muse Voice Transcribe combines streaming automatic speech recognition with speaker diarization and endpointing.
In practical terms, it can transcribe speech as it happens, separate speakers across recordings with 20-plus voices, and determine when someone has finished talking, all without a separate post-processing step.
The model was trained across more than 70 languages, with 25 validated at launch. It supports audio longer than an hour, plus native code-switching within or between sentences. Language, keyword, and context biasing can further improve recognition.
Rather than using one fixed tradeoff between speed and accuracy, Muse Voice Transcribe decides how long to listen before committing each word. Meta calls this “adaptive delay.” The system can move quickly through easier speech while using more audio context for difficult words.
Meta says the model ranks first on the Artificial Analysis streaming speech-to-text leaderboard as of September 1.
Muse Voice Transcribe is available today through the Meta Model API for $3 per 1,000 audio-minutes, equivalent to $0.18 per hour.
It is also already powering dictation in Meta AI for Mac and Muse Code. On Mac, users can hold the Fn key to dictate into any application.
Introducing Muse Voice Transcribe! 🗣️This is our first real-time audio perception model from @AIatMeta. It's trained on 70+ languages and delivers diarization with 20+ speakers. A truly fantastic model.Available now in the Meta AI macOS app! pic.twitter.com/HMAj0i2o5q
You can learn more about the new Muse Voice Transcribe technology here.