vLLM-Omni Unifies Text, Speech, and Video Serving With 91% Faster Completion

vLLM-Omni extends the popular inference engine with a stage-based orchestrator, cutting job completion time by up to 91.4% on Qwen3-Omni serving.

·
·
·
vLLM-Omni Unifies Text, Speech, and Video Serving With 91% Faster Completion
  • vLLM team released vLLM-Omni, a unified runtime for text, speech, image, video, and action generation.
  • Stage-graph architecture: orchestrator, specialized AR and diffusion engines, OmniConnector, and session controller.
  • Reports up to 91.4% job completion time reduction and 10x speedup on Qwen3-Omni Thinker.
  • Single OpenAI-compatible endpoint via vllm serve --omni, reusing PagedAttention and continuous batching.
  • 0.30.0 adds full-duplex sessions, world-model serving, and MiniMax H3 realtime inference on Blackwell and Ascend.
  • Open source at github.com/vllm-project/vllm-omni with H100 nightly CI benchmarks.

vLLM-Omni unifies multimodal model serving

The vLLM team has published a technical report for vLLM-Omni, an open-source runtime that coordinates text, speech, image, video, and action generation through one control plane. The project combines autoregressive and diffusion stages behind a single OpenAI-compatible service, reducing the model-specific integration code required to operate multimodal systems.

For developers, the practical changes include one deployment entry point, stage-aware scheduling, cross-stage data transfer, and stateful sessions for workloads such as voice agents, world models, and robot control loops.

One decode loop no longer fits

Most LLM servers optimize autoregressive decoding, where a model generates one token at a time while reusing cached attention data. Newer generative systems mix that pattern with diffusion, which repeatedly refines an image or video representation, and with persistent sessions that retain state across many short inference calls.

Models such as Qwen3-Omni use a Thinker-Talker architecture with connected autoregressive pipelines. Other systems add diffusion transformers, audio encoders, video generators, or control-policy components. Operating these models has typically required separate servers and custom handoff code, leaving each deployment tightly coupled to one architecture.

A stage graph routes the work

vLLM-Omni represents a model as a directed graph whose nodes are independently served stages. Each stage can use an execution engine, batching policy, parallelism strategy, and hardware allocation suited to its workload.

Component Responsibility
Orchestrator Advances requests through the graph and coordinates dependencies between stages.
Specialized engines Run each stage, using vLLM for autoregressive generation and a dedicated engine for diffusion transformers.
OmniConnector Transfers key-value cache data and multimodal tensors between stages.
Session controller Maintains long-lived state for duplex conversations, world-model rollouts, and robot loops.

Autoregressive stages retain vLLM’s existing optimizations, including PagedAttention, continuous batching, CUDA graphs, and its scheduler. During prefill, the initial pass that builds attention state, placeholder tokens for media inputs are replaced with encoder embeddings. The key-value cache then stores attention data for every token regardless of modality, allowing the scheduler to handle text and multimodal sequences through the same mechanism.

The API stays familiar

Running vllm serve with the --omni flag launches the registered stage graph behind one OpenAI-compatible endpoint. Requests use vLLM’s existing multimodal message structure, extended with audio, video, and a modalities field. Callers submit a completion request without invoking individual stages.

code
vllm serve Qwen/Qwen3-Omni --omni

A text-and-audio response can be requested through the chat completions endpoint:

cpp
curl http://localhost:8000/v1/chat/completions \
  -H "Content-Type: application/json" \
  -d '{
    "model": "Qwen/Qwen3-Omni",
    "modalities": ["text", "audio"],
    "messages": [
      {"role": "user", "content": "Say hello."}
    ]
  }'

Disaggregation drives the largest gains

The report’s largest performance claim is a reduction in job completion time of up to 91.4% for evaluated any-to-any multimodal workloads. Job completion time measures the full interval from request arrival to finished output, including every stage in the pipeline. For Qwen3-Omni’s Thinker component, the reasoning and text-generation branch, the team reports a speedup of more than 10 times from additional vLLM-Omni optimizations.

The measurements come from the project’s public multimodal nightly continuous-integration suite on NVIDIA H100 GPUs. The evaluation covers several workload shapes:

  • Qwen3-Omni speech serving
  • Text-to-speech generation
  • Image and video diffusion
  • Cosmos3 and MiniMax video workloads
  • MiniCPM-o duplex health metrics

The reported percentages describe the best measured cases within that suite. Throughput, latency, and memory use will vary with model topology, request mix, batch size, output length, and hardware allocation.

Sessions carry state across turns

Release 0.30.0 introduces engine-owned sessions for unified full-duplex serving, according to the release notes. Full-duplex operation lets the server receive input while output generation continues, supporting voice-agent behavior such as interruption and barge-in. The release also separates duplex orchestration from turn-based execution and adds shared model-plugin interfaces.

MiniCPM-o 4.5 uses the new session framework, while AURA adds multimodal interaction and multi-turn history. The release also includes native transfer of key-value cache data and multimodal payloads, interactive world-model serving, streaming video generation, and real-time MiniMax-H3 four-step inference on NVIDIA Blackwell and Ascend NPU 950 hardware.

World models and robot policies need persistent execution because each inference step depends on an evolving scene or environment state. Engine-owned sessions keep that state attached to the serving runtime across repeated calls, avoiding repeated reconstruction by an external application.

Where Omni fits best

Workload Why the architecture helps
Speech-to-speech assistants Coordinates Thinker-Talker pipelines and preserves conversational state.
Image and video generation Runs language and diffusion stages within one orchestrated graph.
Full-duplex voice agents Supports simultaneous input, output, interruption, and multi-turn sessions.
World models and robot policies Retains state across frequent inference steps and supports specialized stage hardware.

Teams serving plain text can continue using vanilla vLLM and avoid the additional orchestration layer. vLLM-Omni is most useful when a request crosses model components with different execution patterns or requires persistent session state.

The unified flow depends on registered model pipelines. Custom architectures still require pipeline definitions and compatible stage implementations. The API is also evolving: the legacy stage-configuration path has been removed, and --stage-configs-path and the stage_args YAML loader are no longer supported. Current deployments use registered pipelines through vllm serve --omni. Version pinning and upgrade tests can protect production deployments from minor-release changes.

The stage boundary is the key abstraction

Treating each model component as a stage gives the runtime control over batching, parallelism, data routing, and accelerator placement at the point where those requirements diverge. Autoregressive stages can use continuous batching, diffusion stages can follow their iterative schedules, and session-oriented components can retain state without forcing every part of the model into one worker shape.

The source is available in the GitHub repository, with setup and model-support details in the project documentation. Teams already operating vLLM can extend the same serving environment to supported speech, image, video, and action-generation pipelines through the Omni orchestrator and its registered stages.

Trending
  • No trending articles

Comments

avatar

Next Reads