Skip to main content

ITmatterss

Meta Unveils Muse Voice Transcribe: Supports Hindi, Tamil, Telugu, Kannada and Malayalam

Vertical Share Bar
Meta

News in Short

  • Meta has launched Muse Voice Transcribe, its first real-time audio perception model from Meta Superintelligence Labs.
  • The model supports more than 70 languages, including five Indian languages.
  • Indian language support includes Hindi, Tamil, Telugu, Kannada and Malayalam.
  • It can distinguish more than 20 speakers in a recording.

Meta has launched Muse Voice Transcribe, a new AI model designed to handle real-time speech transcription. The model is the first real-time audio perception model developed by Meta Superintelligence Labs.

One of its biggest highlights is its language support. Muse Voice Transcribe supports more than 70 languages globally, including five major Indian languages — Hindi, Tamil, Telugu, Kannada and Malayalam.

Unlike traditional transcription systems that process an entire recording after it ends, Muse Voice Transcribe can convert speech into text as a person speaks. It also combines transcription, speaker separation and endpointing within a single system.

As a result, developers can use it for applications that need to understand conversations in real time.

Muse Voice Transcribe Can Identify Multiple Speakers

Muse Voice Transcribe is designed for more than simple voice-to-text conversion. Meta says the model can distinguish more than 20 speakers in a recording.

It can also detect when a speaker starts and stops talking. This allows the system to separate conversations into individual speakers while processing the audio.

Another notable capability is its support for long recordings. Meta says Muse Voice Transcribe can handle audio that runs for more than an hour.

That could make the model useful for applications involving meetings, interviews, discussions and other extended conversations.

Supports Code-Switching Between Languages

For Indian users, one particularly useful feature could be Muse Voice Transcribe’s ability to handle code-switching.

People often switch between languages while speaking, sometimes even within the same sentence. Muse Voice Transcribe can process these conversations without requiring users to manually change the selected language every time the speaker switches.

The model also supports language, keyword and context biasing. These capabilities can help it recognise words based on the surrounding conversation and the context of the audio.

How Meta Muse Voice Transcribe Works

Meta says Muse Voice Transcribe processes audio in 80-millisecond chunks. It then decides when it has enough information to transcribe each word.

The system can commit simpler words sooner while waiting longer when a word is more difficult to recognise. Meta says reinforcement learning helps the model make these timing decisions while keeping transcription errors low.

Meta trained Muse Voice Transcribe across more than 70 languages. For its initial release, the company extensively validated 25 languages.

Meta Muse Voice Transcribe Availability and Pricing

Meta has made Muse Voice Transcribe available through the Meta Model API. Developers can therefore integrate the speech transcription model into their own applications.

The model is already being used by Meta AI for Mac and Muse Code for dictation.

For developers, the API costs $3 per 1,000 audio minutes, which works out to approximately $0.18 per hour of audio.

With support for multiple speakers, long recordings and conversations that switch between languages, Muse Voice Transcribe could give developers a more flexible option for building real-time voice applications.

Its support for Hindi, Tamil, Telugu, Kannada and Malayalam also makes the launch particularly relevant for Indian users and developers building multilingual AI experiences.

40

Leave a Reply

Your email address will not be published. Required fields are marked *

logo

Get the latest news instantly

You can change your preferences anytime.