Meta releases Muse Spark 1.3 with a new max reasoning tier on Muse Code and the Meta Model API.
Meta Superintelligence Labs Muse Voice Transcribe launch
Meta Superintelligence Labs Launches Muse Voice Transcribe Real-Time ASR
Handling speaker diarization, endpointing, and turn detection natively in a single streaming model differs from conventional speech-to-text APIs, which typically require separate components for each function. Meta reports benchmark-leading results across both streaming speech-to-text and diarization.
The full picture
Muse Voice Transcribe, released September 1, 2026, is Meta Superintelligence Labs' first real-time audio perception model. It delivers streaming ASR with a 3.1% word error rate, speaker diarization for more than 20 speakers, and turn-endpointing natively within a single model. The model processes audio in 80ms chunks, dynamically deciding after each chunk whether to emit text or continue listening. It supports multilingual speech and seamless code-switching, with accuracy improvements available through language, keyword, and context biasing. Meta reports it ranks first on Artificial Analysis streaming speech-to-text and public diarization benchmarks. It is available through the Meta Model API, Meta AI for Mac, and Muse Code. On September 4, Meta also released Muse Spark 1.3 with a max reasoning tier on Muse Code and the Meta Model API, reporting stronger coding and agentic performance than the previously available high and xhigh tiers.
How it developed
Meta Superintelligence Labs released Muse Voice Transcribe on September 1, its first real-time audio perception model, handling streaming speech-to-text, speaker diarization for more than 20 speakers, and turn-endpointing within a single model.
The model processes audio in 80ms chunks and Meta reports a 3.1% final-transcription word error rate and first-place rankings on Artificial Analysis streaming speech-to-text and public diarization benchmarks. It is available through the Meta Model API, Meta AI for Mac, and Muse Code.
Meta Superintelligence Labs releases Muse Voice Transcribe, its first real-time audio perception model, with streaming ASR, diarization, and endpointing in a single model.
Sources
Want this in your inbox?
I send a short email each morning with the stories that moved. If you would rather just read here, that works too.
Subscribe free