Frontier Flux Original
On 1 September 2026 Meta Superintelligence Labs launched Muse Voice Transcribe, its first real-time audio perception model, with streaming ASR, 20-plus-speaker diarization, and endpointing. It is live on the Meta Model API, Meta AI for Mac, and Muse Code. Rankings Meta cites are as of that date.

Abstract concept art for Muse Voice Transcribe streaming speech. AI-generated.
Frontier Flux / AI-generated
Meta Superintelligence Labs launched Muse Voice Transcribe on 1 September 2026. The research post calls it the lab's first real-time audio perception model: real-time streaming ASR, diarization with 20-plus speakers, and endpointing, plus multilingual code-switching and language, keyword, and context biasing.
AIatMeta on X posted the same day: "Introducing Muse Voice Transcribe, the first real-time audio perception model from Meta Superintelligence Labs."
A follow-up from AIatMeta says: "Muse Voice Transcribe is an autoregressive multimodal LLM from the Muse Spark family." The research post calls it an autoregressive multimodal model in that family. Audio is processed in 80ms chunks (12.5 Hz). Each chunk becomes one soft token. At each chunk the model continues listening or emits a text token.
Key facts: - Meta Superintelligence Labs says this is its first real-time audio perception model. - Training includes RL with combined word error rate and delay rewards, per the same follow-up post. - Trained with 70-plus languages; 25 are extensively verified at launch. Meta recommends those 25 first. - The research post says the model natively supports long audio exceeding one hour and up to 20-plus speakers, with no required post-processing. - Meta says it is available today through the Meta Model API, Meta AI for Mac, and Muse Code. On Mac, Meta says users hold the Fn key to dictate into any application. - Meta says it ranks first on Artificial Analysis streaming speech-to-text and on public diarization benchmarks as of 1 September 2026.
How to read this: the architecture, language counts, hour-long context, Mac Fn dictation, and leaderboard ranks are Meta's launch claims. The X follow-up uses "LLM"; the research post says "multimodal model." Rankings can move after 1 September.
Frontier Flux is independent coverage. We did not train or host Muse Voice Transcribe.
Free
Free daily email of model releases, benchmarks, and lab notes from the last 24 hours across OpenAI, Anthropic, DeepMind, Meta, NVIDIA, and more. Capability-first, not company drama.
Opens Buttondown to finish signup. Confirm by email if asked. No ads. Unsubscribe anytime.