AI

ElevenLabs v4 needs 10 seconds of your voice to make it speak 90 languages

Susan Hill
Add us on Google

ElevenLabs has released Eleven v4, a text-to-speech model that reads written text aloud with a laugh, a sigh or an angry French accent wherever the writer asks for one. Its instant voice cloning needs a recording shorter than a voicemail, and the cloned voice can then speak dozens of languages while keeping its own timbre. Anyone with a free ElevenLabs account can use it.

For people who make things with sound, that removes whole days of work. A podcaster can dub an episode into Japanese in their own voice, an indie game developer can give every minor character a distinct delivery without booking a studio, and an audiobook narrator can write a pause or a chuckle straight into the script. The flip side is just as concrete: a voice note, a clip from a video call or a few seconds of a public post is now enough raw material for a convincing copy of someone.

The main change is control, and the clone itself takes just 10 seconds of audio. Earlier ElevenLabs models read text cleanly but had to guess at the mood. Version 4 takes stage directions written inline in square brackets, such as [laughs] or [said angrily in French accent], and even background cues like [light rain] or [phone buzzing]. The company says the model also reads tone and pacing from the surrounding context, so a line in a tense scene comes out tense without a tag. Scenes with several speakers keep each voice consistent over long passages, a weak point of synthetic narration, where voices tended to drift over the course of a chapter.

Language coverage grows from 70 languages in the previous version to more than 90, and ElevenLabs singled out Japanese, Brazilian Portuguese, Mandarin and Cantonese as the biggest quality gains, according to TechCrunch. A cloned voice keeps its identity when it switches language, so a Spanish speaker’s clone reading German still sounds like that person, with a native German accent. A second model, v4 Turbo, is built for voice assistants and call-center agents. It starts speaking in about 150 milliseconds, roughly the pause between two people taking turns in a conversation, and it can begin talking before the chatbot behind it has finished writing its answer.

ElevenLabs is not alone in this market. Google, OpenAI, Cartesia, Deepgram and Fish Audio all sell synthetic voices to the same developers. Within hours of the launch, the independent benchmarking firm Artificial Analysis placed v4 first on its voice leaderboard, where listeners rate anonymous samples side by side. ElevenLabs also says about three in four listeners preferred v4 in blind comparisons with rivals, a figure that comes from its own testing.

The caveats start with that 10-second figure. ElevenLabs says its clones require the consent of the person whose voice is uploaded, it runs a public classifier that can flag audio made with its tools, and it blocks clones of some political figures outright. Those checks stop casual misuse, but they rely on the uploader telling the truth, and a scammer needs only a short sample of a relative’s voice to stage a fake emergency call. The free tier is also small: about 10 minutes of generated audio a month, according to Unite.AI’s breakdown of the plans, with paid tiers starting at $6 a month.

Eleven v4 and v4 Turbo went live on September 28 in ElevenLabs’ creator app, its agent platform and its developer API. The London-based company raised $500 million from Sequoia in February at an $11 billion valuation, and TechCrunch reports its annualized revenue has passed $600 million. Chief executive Mati Staniszewski told TechCrunch the company is aiming for a stock market listing in the coming years, without committing to a date.

For creators, a voice that laughs on cue in 90 languages is a recording studio in a browser tab. For everyone else, it is one more reason to hang up and call back when a familiar voice on the phone asks for money.

Tags: , , , , ,

Add us on Google

Discussion

There are 0 comments.