Runway Turns Gen-4.5 Into a Multi-Model Creative Router for Teams
Runway's Fall 2026 drop bundles Fish Audio S2.1 Pro, MiniMax H3 Max, Cartesia Sonic 3.6, plus Flux video upscaling and editing into one workspace.
- Runway Fall 2026 collection adds five third-party models to its unified creative workspace, available now at app.runwayml.com.
- Fish Audio S2.1 Pro delivers expressive TTS with inline emotion tags at roughly $15 per million characters.
- MiniMax H3 Max keeps H3 quality but runs faster, strong on Chinese narration and expressive sound tags.
- Cartesia Sonic 3.6 targets real-time agents with sub-100ms first audio and multi-language prosody control.
- Flux Video Upscale regenerates frames to 2K or 4K in Precise or Creative modes, priced $0.07 to $0.10 per mp-second.
- Flux Video Edit joins Aleph 2 for prompt-based object removal, replacement, and full-scene restyling of existing footage.
Runway turns its creative suite into a model router
Runway has folded several third-party audio and video models into the same workspace as its Gen-4.5 and Aleph 2 models. The Fall 2026 collection, which Runway describes as its largest model release, adds speech synthesis from Fish Audio, MiniMax, and Cartesia alongside Black Forest Labs tools for video upscaling and editing.
For developers and production teams, the release reduces the number of vendor integrations, billing systems, and creative interfaces required for a multi-model pipeline. Teams can compare models per shot or voice track, then keep the selected output inside Runway’s app or, where supported, its developer API.
Three voices, distinct tradeoffs
| Model | Strengths | Useful controls | Good fit |
|---|---|---|---|
| Fish Audio S2.1 Pro | Expressive, cost-conscious speech generation with access to Fish Audio’s large community voice catalog | Inline delivery tags such as (whisper) and (excited) |
Long-form narration, voiceovers, and projects generating large volumes of speech |
| MiniMax H3 Max | Fast generation, strong Chinese-language output, and expressive character delivery | Sound tags for details such as laughter and breathing | Mandarin content, character dialogue, and latency-sensitive generation |
| Cartesia Sonic 3.6 | Streaming multilingual speech with low time to first audio | Speed and emotion controls | Voice agents, interactive characters, and conversational applications |
The cited Artificial Analysis snapshot gives Fish Audio S2 Pro an Elo score of 1,128 at $15 per million characters. Elo summarizes relative preference in comparative evaluations, and leaderboard positions change over time. The benchmark covers S2 Pro, so it provides directional evidence for S2.1 Pro rather than a direct measurement of the model added to Runway. Runway’s own usage price may also differ from the benchmarked provider rate.
Runway positions MiniMax H3 Max as matching H3 quality with faster generation. Cartesia’s public documentation reports roughly 90 milliseconds to first audio under ideal conditions, making network latency, region, text length, and buffering important variables in production tests.
Flux rebuilds low-resolution frames
Black Forest Labs’ Flux Video Upscale uses a super-resolution model to synthesize detail at the target resolution. The model targets defects that become conspicuous on larger displays, including soft eyes, smeared faces, and grid artifacts in textures such as water, foliage, and grass.
| Mode | Steps | Listed rate | Tradeoff |
|---|---|---|---|
| Precise | 4 | $0.07 per megapixel-second | Faster processing with stronger identity and reference retention |
| Creative | 8 | $0.10 per megapixel-second | More aggressive repair and detail generation, with greater risk of changing faces or products |
Steps are iterative model passes, with Creative allocating more compute and more freedom to generate texture. Identity-critical footage, branded products, and shots tied to reference images are safer candidates for Precise mode.
A linear scale factor of 1.5x, 2x, or 3x applies to both frame dimensions. For a 1280 × 720 source, those settings produce 1920 × 1080, 2560 × 1440, and 3840 × 2160 output respectively, subject to the service’s resolution limits.
Megapixel-second pricing multiplies the applicable rate by the metered megapixels and clip duration. Teams should confirm whether Runway meters input or output resolution before forecasting costs, since that choice materially changes the total for 4K delivery.
Editing becomes a per-shot model choice
Flux Video Edit transforms existing footage from a text prompt. It can remove, add, or replace objects, change a setting, and restyle an entire clip while preserving the source video’s underlying motion.
The model overlaps with Aleph 2, Runway’s first-party video transformation model. Access to both in one interface allows teams to test the same shot against each model and compare prompt adherence, identity retention, temporal consistency, and unwanted changes before selecting an output.
Runway takes the routing layer
Runway has spent the year building a multi-model workbench around its first-party tools. Its expanding catalog now covers several stages of image, video, and audio production:
- First-party models: Gen-4.5 and Aleph 2.
- Third-party video models: MiniMax Hailuo 3.0, Seedance 2.5, and Wan 3.0.
- Image and audio models: Nano Banana 2, ElevenLabs v3, Fish Audio S2.1 Pro, MiniMax H3 Max, and Cartesia Sonic 3.6.
- Post-production access: an Adobe plugin that connects the model catalog to Premiere Pro and After Effects.
MiniMax Hailuo 3.0 is also available through Runway Dev for text-to-video, image-to-video, and video-to-video generation, with support for keyframes and reference images, video, and audio. The broader catalog gives teams one routing layer for choosing models according to language, latency, visual consistency, and cost.
A three-model workflow
- Generate Mandarin narration with MiniMax H3 Max, or use Cartesia Sonic 3.6 when playback must begin with minimal delay.
- Transform the source clip with Flux Video Edit and Aleph 2, then select the version with stronger temporal and identity consistency.
- Upscale the approved shot with Flux Video Upscale, using Precise mode for faces, products, and reference-dependent footage.
Limits behind the unified UI
- Clip length: Flux Video Upscale accepts up to 20 seconds per call, so longer footage requires chunking and careful joins.
- Identity drift: Creative mode may alter faces, logos, products, or other reference-critical details.
- Model-specific behavior: Language coverage, streaming protocols, output formats, and rate limits can vary behind Runway’s shared interface.
- Voice governance: Production use requires checks on speaker consent, cloning rights, voice storage, retention policies, and commercial licensing.
- API coverage: Runway says most models are available through Runway Dev. Teams should verify each model’s endpoint, schema, regional access, concurrency limits, and billing before automating a pipeline.
Runway says the Fall 2026 collection is live in its app, with most of the included models also accessible through the Runway Dev API.