ElevenLabs Launches Faster, More Expressive AI Voice Models for 90+ Languages
ElevenLabs has introduced its new v4 and v4 Turbo text-to-speech models, combining faster generation, stronger pronunciation, emotional controls and ultra-short voice cloning to make AI voices more...
ElevenLabs has launched its latest text-to-speech models, v4 and v4 Turbo, aimed at making AI-generated voices faster, more expressive and easier to control.
The company says the models support more than 90 languages, while v4 has achieved an Elo rating of 1,319 in independent benchmarking for quality and pronunciation. The company’s latest technology also focuses heavily on real-time applications, with v4 Turbo delivering response times of around 100 milliseconds.
One of the biggest changes is improved voice cloning. Users can now create a voice clone using as little as 10 seconds of audio, potentially making personalized AI voices more accessible to creators, businesses and developers.
The models also introduce inline controls for emotions and audio effects, allowing developers to influence how a voice sounds and reacts within generated speech. Early demonstrations have highlighted natural accents and more expressive delivery.
For marketers, this could enable more engaging AI voiceovers, multilingual campaigns, advertisements, videos, virtual assistants and conversational experiences.
ElevenLabs is initially offering promotional pricing at $22 per million characters for v4 and $11 for v4 Turbo for two weeks.
The update signals a broader shift toward AI voices that are not only realistic, but also faster, emotionally expressive and controllable.
Akash Takes: AI voice is moving from simply sounding human to understanding how humans communicate. Faster responses and controllable emotion could make synthetic voices increasingly useful in real-time marketing, customer conversations and AI-powered content production.


