ElevenLabs has launched two speech models, ElevenLabs v4 and v4 Turbo, with more expression controls, lower latency for voice agents and support for 90 languages. The company said users can clone a voice from 10 seconds of audio.
The models use a new architecture that ElevenLabs says enables better control and faster cloning. They can track text context while reading aloud and adjust expression over longer passages. Building on inline expression tags introduced with v3, v4 lets users combine multiple tags and have the model follow them in sequence.
The previous version supported 70 languages. ElevenLabs said it saw the largest quality improvements in Japanese, Brazilian Portuguese, Mandarin and Cantonese.
The company says v4 is suited to voice agents because it can begin generating audio as the language model behind it starts generating a response, supporting more fluid conversations. It can also handle confrontations, escalations and holds differently to help resolve issues.
ElevenLabs said more than 55% of its business comes from large companies. The startup raised $500 million earlier this year in a Sequoia-led round that valued it at $11 billion. Its annualized revenue run rate rose from roughly $330 million at the start of the year to more than $600 million. The company has more than 800 employees. CEO Mati Staniszewski told TechCrunch the company is aiming for an IPO “in the next years,” without giving a timeline.
Comments
0No comments yet. Be the first to comment.