News in Short
- ElevenLabs has launched its new v4 and v4 Turbo speech models.
- The models support more than 90 languages.
- ElevenLabs says v4 offers greater control over voice expression.
- Users can reportedly clone a voice using just 10 seconds of audio.
- The model can maintain voice identity across longer pieces of text.
ElevenLabs has introduced two new AI speech models as competition in voice technology continues to grow.
The company has launched ElevenLabs v4 and v4 Turbo. The new models focus on expressive speech, voice control and faster responses for AI voice agents.
ElevenLabs says the v4 generation uses a new architecture. This architecture also enables faster voice cloning and greater control over how voices deliver content.
The company launched its previous v3 model last year. Now, the new generation expands its capabilities across languages and conversational applications.
ElevenLabs v4 Adds More Control Over Voice Expression
One of the main changes in v4 is greater control over how an AI voice delivers speech.
The model can consider the context of the text while reading it aloud. Therefore, it can adjust expressions based on what is happening within the content.
ElevenLabs previously introduced inline tags with v3. With v4, users can stack multiple tags and specify a sequence of expressions.
This gives creators more control over the delivery of generated speech. It can also help maintain a voice’s identity across longer sections of text.
The company says v4 can also clone a voice using only 10 seconds of audio.
Support Expands to More Than 90 Languages
ElevenLabs has also expanded the number of languages supported by its speech models.
The previous generation supported 70 languages. With v4, the company says support has grown to more than 90 languages.
According to ElevenLabs, some of the biggest quality improvements appeared in Japanese, Brazilian Portuguese, Mandarin and Cantonese.
This expansion could make the models more useful for creators and businesses producing voice content for different markets.
v4 Targets AI Voice Agents
ElevenLabs is also focusing heavily on conversational AI applications.
The company says the new models offer lower latency, which can help voice agents respond more naturally during conversations.
Furthermore, v4 can begin generating audio as soon as the underlying large language model starts producing an answer. This can reduce the wait before users hear the beginning of a response.
The company also says the model can handle situations such as confrontations, escalations and holds differently. These capabilities could be useful for customer service and enterprise calling applications.
ElevenLabs has expanded its enterprise calling business rapidly over the past year. More than 55 percent of its business now comes from large companies, according to the company.
ElevenLabs Faces Growing Voice AI Competition
The launch comes as competition in AI-generated speech continues to increase.
Startups including Cartesia, Deepgram, Fish Audio, Boson and WellSaid Labs have developed their own expressive speech technologies.
Meanwhile, major technology companies such as Google and OpenAI have also continued to improve their voice models.
As a result, ElevenLabs is pushing v4 toward both creative applications and enterprise voice agents.
ElevenLabs Continues Global Expansion
The new speech models arrive as ElevenLabs expands its business and workforce.
The company raised $500 million in February from Sequoia in a funding round that valued it at $11 billion. There have also been reports of another potential funding round that could value the company at $22 billion.
According to the company, its annualised revenue run rate has grown from around $330 million at the start of 2026 to more than $600 million.
ElevenLabs has also expanded its workforce across markets including India, Europe and Brazil. Its headcount has crossed 800 employees.
Meanwhile, co-founder and CEO Mati Staniszewski recently said the company is targeting an IPO “in the next years”, without giving a specific timeline.