Runway Brings ElevenLabs v4 Voice Into Its Visual Production Workflow

Runway now offers ElevenLabs v4 inside its creative suite, bringing a more expressive text-to-speech model to video generation workflows.

·
·
  • Runway integrated ElevenLabs v4 into its creative suite for narration and character dialogue.
  • v4 uses a new architecture tuned for tone, pacing, emotion, and speaker identity consistency.
  • Voice cloning now needs only 10 seconds of reference audio.
  • Language coverage expanded from 70 to 90, with big gains in Japanese and Mandarin.
  • Single generations support up to 10,000 characters with context stitching for long scripts.
  • Available now in the Runway app and via ElevenAPI.

Runway adds ElevenLabs v4 voice generation

Runway has added ElevenLabs v4 to its creative platform, placing voice generation alongside its image and video tools. Creators can now produce narration and character dialogue within the Runway workflow instead of exporting visuals to a separate voice service.

ElevenLabs released two models alongside the integration. Eleven v4 focuses on expressive, rendered speech for narration and dialogue. Eleven v4 Turbo targets voice agents and other applications that depend on low latency. Runway currently exposes the standard v4 model.

V4 expands performance and script length

ElevenLabs says v4 uses a new architecture that interprets tone, pacing, emotion, character and surrounding context while preserving the selected speaker’s identity. In practical terms, the model can vary delivery across dramatic, conversational, comedic and subdued passages without requiring separate audio processing for each style.

The release includes several changes relevant to video and audio production:

  • Short voice samples: A voice can be cloned from 10 seconds of recorded audio.
  • Longer generations: One request can contain up to 10,000 characters. Context stitching helps maintain pacing and delivery across longer scripts.
  • Broader language support: ElevenLabs expanded support from 70 to 90 languages and reports its largest quality gains in Japanese, Brazilian Portuguese, Mandarin and Cantonese.
  • More expressive controls: Inline tags introduced with v3 can now be stacked, allowing the model to follow a sequence of performance instructions.
  • Context-aware dialogue: Multi-speaker generations use the surrounding conversation to shape timing, delivery and responses.

The multi-speaker upgrade directly affects scenes built around conversation. Context-aware delivery can produce more coherent exchanges because each line reflects the preceding dialogue, reducing the detached cadence common when separately generated clips are assembled in an editor.

A 10-second cloning requirement lowers the technical barrier to reproducing a voice. Teams still need permission from the speaker and appropriate rights for the intended distribution, regardless of how little source audio the model requires.

Voice generation moves into the visual workflow

Producing a short film in Runway previously involved rendering clips, generating speech in ElevenLabs and importing the resulting audio into an editing workflow. The integration removes that export-and-import cycle for voice generation. Runway already combines Gen-4.5 video with ElevenLabs audio, and v4 extends that pipeline with longer scripts and more expressive delivery.

The partnership also runs through ElevenLabs’ products. Its ElevenCreative platform provides access to Runway video models, while Runway now hosts ElevenLabs’ latest voice model. Both companies are assembling broader production environments around their core generation systems.

Choose the model around latency

Runway’s v4 integration suits rendered projects in which vocal performance carries more weight than immediate response time. Common uses include animated shorts, narrated explainers, audiobooks, character performances, dubbing and localized versions of existing clips.

Model Primary use Key detail
Eleven v4 Narration, dubbing and character dialogue Available in Runway and supports generations of up to 10,000 characters
eleven_v4_turbo Voice agents and real-time applications Targets median inference latency of about 100 milliseconds

The Turbo latency figure covers model inference. Network transit, application processing, buffering and playback can increase end-to-end response time, so developers should measure latency inside the complete application stack.

Runway credits cover the integration

The integration is live in the Runway app. Voice generations consume Runway credits, allowing existing customers to use v4 without maintaining a separate ElevenLabs subscription for work performed inside Runway. The available information does not specify the exact credit cost per generation or character, so production teams will need to check the in-app rate before estimating batch workloads.

Developers seeking programmatic access can use ElevenAPI directly. ElevenLabs has also made both models available through ElevenAgents and ElevenCreative, with access included in its free account tier subject to the applicable usage limits.

Benchmark claims need production testing

ElevenLabs says Artificial Analysis ranked v4 first and that roughly 75% of listeners preferred it over competing models in blind head-to-head tests. The usefulness of that result depends on the prompt set, languages, selected competitors, voices and evaluation method. Teams working with specialized accents, long dialogue scenes or strict identity requirements will get a more relevant comparison from their own scripts and speakers.

For existing Runway users, the immediate change is operational: narration and dialogue can remain inside the visual production workflow. For teams moving from ElevenLabs v3, v4 adds longer generations, broader language coverage, sequenced expression controls and more context-aware conversations.

Trending
  • No trending articles

Comments

avatar

Next Reads