Meta releases Muse Voice Transcribe with 3.1% WER, undercutting rivals at $3 per 1,000 minutes
Meta has launched Muse Voice Transcribe, a real-time speech-to-text model achieving a 3.1% word error rate at 0.16-second latency, outperforming competitors like ElevenLabs Scribe v2 Realtime and Cartesia Ink-2. Priced at $0.18 per hour ($3 per 1,000 minutes), it undercuts rivals by over 50%. The model supports 70+ languages, processes audio in 80-millisecond chunks, and features adaptive delay and native speaker diarization. It is available via the Meta Model API, Meta AI for Mac, and Muse Code.
Editorial responsibility
- No named human review is recorded for this page.
- Reports are grouped by semantic similarity and deterministic rules. Language models may assist titles, summaries, translation and cross-source analysis; the page itself is projected from evidence records.
- Current automated evidence projection
Cross-source coverage
Wire timeline
Meta releases Muse Voice Transcribe, its first real-time speech recognition model
Meta AI released Muse Voice Transcribe early this morning, marking its first official entry into real-time speech recognition. The model, developed by Meta's MSL division, claims state-of-the-art performance in streaming speech-to-text, integrating previously separate modules into a single system. It supports over 70 languages and 20 speakers, with native speaker diarization and endpoint detection. A key feature is Adaptive Delay, which dynamically adjusts latency based on speech complexity—waiting longer for difficult words and outputting faster for simple content. This aims to balance real-time performance and accuracy, addressing a common issue in voice agents where cutting in too early interrupts users and waiting too long feels unresponsive. The release is part of Meta's broader Muse series rollout across general models, images, video, and coding agents, potentially becoming the underlying model family for Meta's personal AI system across WhatsApp, Instagram, Meta AI, and Ray-Ban glasses.
Muse Voice Transcribe launches at $3 per 1,000 minutes, undercutting Cartesia and ElevenLabs
Muse Voice Transcribe, a new AI-powered speech-to-text service, has been released with a pricing model of $0.18 per hour, equivalent to $3 per 1,000 minutes of audio. This price point positions it below Cartesia Ink-2, which charges $4 per 1,000 minutes, and significantly undercuts ElevenLabs Scribe v2 Realtime and Deepgram Flux, both of which are priced at $6.50 per 1,000 minutes. The announcement highlights a competitive pricing strategy in the AI transcription market, offering a cost reduction of more than 50% compared to the higher-priced alternatives. The service is available immediately, as indicated by the linked announcement. This pricing move could pressure other providers to adjust their rates or differentiate their offerings in the rapidly evolving AI voice services sector.
Meta releases Muse Voice Transcribe, achieving 3.1% WER at 0.16s latency for streaming speech-to-text
Meta has released Muse Voice Transcribe, a streaming speech-to-text model developed by Meta Superintelligence Labs. The model achieves a 3.1% Word Error Rate (WER) at 0.16 seconds after the end of speech for final transcripts, claiming the top spot on the AA-WER Streaming benchmark. It outperforms competitors like Cartesia Ink-2 and ElevenLabs Scribe v2 Realtime in accuracy, with competitive latency. For first partial transcripts, it achieves 3.6% WER at 0.13 seconds. The model was trained on over 70 languages, with 25 extensively verified, and supports audio inputs exceeding one hour without post-processing, processing audio in 80ms chunks. Pricing is set at $0.18 per hour ($3 per 1,000 minutes), significantly undercutting rivals such as Cartesia Ink-2 ($4 per hour) and ElevenLabs Scribe v2 Realtime ($6.50 per hour). Muse Voice Transcribe is available through the Meta Model API, Meta AI for Mac, and Muse Code.
Show 2 older updatesHide older updates
Muse Voice Transcribe achieves 3.6% WER at 0.13s, beating ElevenLabs Scribe v2 Realtime
Muse Voice Transcribe has achieved a 3.6% Word Error Rate (WER) at 0.13 seconds after the end of speech on the First Partial Transcript benchmark, placing it just ahead of ElevenLabs Scribe v2 Realtime in both accuracy and latency. The model also outperforms Cartesia Ink-2 with external endpoints, which recorded a 4.0% WER at 0.07s, trading some speed for higher accuracy. Additionally, Muse is both more accurate and faster than AssemblyAI Universal-3.5 Pro Realtime, whose Max Accuracy and Min Latency modes achieved 4.0% WER at 0.18s and 0.17s, respectively. These results position Muse as a leading contender in real-time voice transcription, offering a competitive balance of low error rates and rapid processing.
Meta launches Muse Voice Transcribe with 3.1% word error rate for real-time dictation
Meta has released Muse Voice Transcribe, a new real-time voice dictation model that achieves a 3.1% final-transcription word error rate with adaptive delay, significantly outperforming competing models. Unlike conventional speech-to-text APIs, Muse processes audio in 80-millisecond chunks and dynamically decides when to emit text, when to wait, and when a speaker change or turn ends, functioning as a real-time perception layer for voice agents. The model is available through the Meta Model API, Meta AI for Mac, and Muse Code. This release represents a major advancement in streaming voice transcription technology, offering lower latency and higher accuracy for developers building voice-enabled applications.