ElevenLabs has introduced Eleven v4, a new Text-to-Speech model designed to read scripts like an actor: understanding who is speaking, what happened before a line and how it should land.
Direction inside the script
The model supports instructions such as [laughs], [whispers], [pause] and audio events directly in the text. The company says it supports more than 90 languages, multiple speakers and soundscapes.
Consistency across long projects
Context stitching is designed to keep pacing and delivery stable across long-form audio such as audiobooks. Lines can be regenerated without vocal drift.
For voice agents
Eleven v4 Turbo targets real-time conversations, with median inference latency of about 100 milliseconds and time to first speech of about 150 milliseconds.
The shift is from basic text-to-speech to a controllable voice-direction layer inside the script. Voice cloning still requires responsibility: ElevenLabs says professional cloning requires verified consent.
Source: ElevenLabs.

