ElevenLabs has released Eleven v4, a speech model designed to follow performance directions and preserve a voice across long productions. It responds to tags or plain-language cues for effects such as laughter and whispering, while phonetic spellings give creators more control over names and technical terms.

The model accepts up to 10,000 characters per request, roughly ten minutes of audio. Longer projects are divided into segments, with the system intended to maintain pacing and character delivery across transitions and regenerated lines. Language support has expanded from about 70 to more than 90, and professional voice clones return after being unavailable in v3.

A Turbo version targets real-time voice agents and begins generating speech in about 150 milliseconds, according to ElevenLabs. The company’s latency and quality claims still need testing in varied workflows. For audiobook, localization and conversational-app teams, the concrete improvement is tighter direction control without rebuilding a voice between scenes.