PyTorch and Red Hat's Helion Beats CUTLASS on H100 by 17.8%

PyTorch's Helion kernel DSL now powers a vLLM linear backend that beats CUTLASS and DeepGEMM on Hopper, with 10%+ end-to-end throughput gains.

·
·
·
PyTorch and Red Hat's Helion Beats CUTLASS on H100 by 17.8%
  • PyTorch and Red Hat shipped a Helion-based linear backend for vLLM targeting Hopper GPUs
  • One Helion GEMM kernel covers Standard, Split-K, and Swap-AB variants, chosen per shape by the autotuner
  • Kernel-level geo mean speedups of 1.11x to 1.178x over CUTLASS, DeepGEMM, and FlashInfer
  • End-to-end vLLM serving sees consistent gains, exceeding 10% throughput on some workloads
  • Hybrid dispatch sends small shapes (up to 32 tokens) to Helion under CUDA Graphs, falls back otherwise
  • Available now in the redhat-et/vllm-helion fork via --linear-backend helion

Helion’s autotuned kernels speed up vLLM on Hopper

PyTorch and Red Hat have released a Helion linear backend for vLLM. On an NVIDIA H100, its high-level quantized matrix-multiplication kernels delivered geometric mean speedups of 11.0% to 17.8% over CUTLASS, DeepGEMM, and FlashInfer. The integration moves shape-specific kernel selection into an autotuner, reducing the need for separate implementations and hand-written dispatch rules.

The published implementation targets small-token inference on NVIDIA Hopper GPUs. It supports three 8-bit quantization formats and uses existing vLLM kernels as fallbacks for larger workloads. Code, tuning scripts, and pre-tuned configurations are available in a Red Hat-maintained vLLM fork.

Autotuning replaces manual dispatch

vLLM routes transformer linear layers through a backend that performs general matrix multiplication, or GEMM. Quantized GEMMs multiply low-precision activations and weights, then apply scaling factors to recover the intended numerical range. vLLM already dispatches these operations to specialized libraries such as CUTLASS, DeepGEMM, and FlashInfer.

Helion is a Python-embedded kernel DSL built around tiled GPU programs. Its ahead-of-time autotuner explores memory layouts, execution schedules, tile sizes, and algorithm choices for each target shape. The selected configurations can then be stored and reused during inference.

The initial vLLM integration covers the following formats:

Format Activation scaling Weight scaling
FP8_Dynamic Per token Per channel
W8A8_INT8 Per token Per channel
Block_FP8 1 × 128 blocks 128 × 128 blocks

vLLM’s broader linear interface also handles formats such as INT4 and NVFP4. Those formats are outside this Helion release.

One kernel, three execution plans

GPU GEMM performance depends heavily on the matrix dimensions. For an operation written as C = A @ B, M commonly tracks the number of tokens processed together, N represents the output width, and K is the reduction dimension. Decode workloads often have a small M, leaving too little parallel work for a conventional GEMM schedule.

The Helion implementation keeps three execution plans in one kernel definition:

Plan How it works Useful conditions
Standard Tiles the original matrix multiplication directly. Shapes with enough work across M and N.
Split-K Partitions K across thread blocks and combines their partial results. Shapes with limited parallelism across the output dimensions.
Swap-AB Computes (B.T @ A.T).T to expose a more favorable tile layout. Shapes with a small M, including decode-heavy workloads.

Existing backends commonly implement these plans separately and choose among them with fixed heuristics. vLLM’s current Block_FP8 path, for example, selects Swap-AB on Hopper when M is below 32.

Helion exposes the plan through hl.register_tunable. In the reference kernel, split_k can take power-of-two values up to 256, while swap_ab is a Boolean option. Offline tuning benchmarks the candidates and records the winner for each target shape.

Small shapes take the Helion path

Helion dispatch and kernel launch can add enough CPU-side overhead to offset its GPU gains when execution occurs outside CUDA Graph replay. CUDA Graphs record a sequence of GPU operations and replay it with less recurring CPU work, making them well suited to stable inference shapes.

The backend therefore uses hybrid dispatch:

  1. Token counts up to max_helion_size use Helion during CUDA Graph replay. The published configuration sets the threshold to 32.
  2. Larger shapes continue through the default CUTLASS or DeepGEMM path.

This design concentrates tuning on the small-token decode regime and reduces the number of configurations that must be generated, validated, and distributed. CUTLASS and DeepGEMM remain part of the runtime path for workloads beyond the threshold.

H100 gains, from kernel to server

The kernel benchmarks ran on an NVIDIA H100 with 80 GB of HBM3 memory. The evaluated workloads included Qwen3 models from 1.7B to 32B parameters and Qwen3.8-27B. Helion produced the following geometric mean speedups:

Format Baseline Speedup Relative gain
FP8_Dynamic CUTLASS 1.110× 11.0%
W8A8_INT8 CUTLASS 1.178× 17.8%
Block_FP8 FlashInfer 1.149× 14.9%
Block_FP8 DeepGEMM 1.177× 17.7%
Box plots comparing Helion GEMM speedups with CUTLASS, FlashInfer, and DeepGEMM
Kernel speedups vary by matrix shape, supporting the use of shape-specific tuning rather than one global configuration.

End-to-end vLLM serving tests used the ShareGPT dataset. The Helion backend improved throughput across the evaluated model and quantization combinations, with gains above 10% for some workloads.

Heatmap comparing vLLM throughput across models, quantization formats, and batch sizes
Reported end-to-end throughput gains across the tested model sizes and serving configurations.

The published evidence covers one GPU model, selected Qwen workloads, and three quantization formats. Results for other Hopper GPUs, model architectures, tensor-parallel configurations, and request distributions require separate measurement.

Speed shifts work offline

Fine-grained autotuning reduces manual kernel work but introduces operational costs:

Cost Effect Mitigation
Tuning time Searching detailed shape-specific configurations can take hours. Run tuning offline and reuse the resulting files.
Cold-start latency CUDA Graph capture can trigger Helion JIT compilation during vLLM startup. Cache compiled artifacts for warm starts.
Configuration upkeep Pre-tuned model files are large and difficult to validate exhaustively. Limit bundled configurations to common models and retune target deployments.

The authors reduce search time with LLM-guided autotuning. Their LLMSeededLFBOTreeSearch setup uses Claude Opus 4.8 to propose initial candidates, followed by a numerical tree search that refines those configurations.

Run it from Red Hat’s fork

The implementation, pre-tuned configurations, and autotuning utilities are available in the Red Hat fork. Within that repository, --linear-backend helion selects the new backend:

haskell
vllm serve \
    --model "$MODEL" \
    --max-num-seqs 32 \
    --tensor-parallel-size 1 \
    --no-enable-prefix-caching \
    --linear-backend helion

The other options in this example set the maximum sequence count, use one tensor-parallel worker, and disable prefix caching. They describe the published serving setup; the backend selector is the final flag.

Models without a matching bundled configuration require an offline tuning run against their expected matrix shapes. Production evaluation also needs to account for first-start compilation, cache persistence, actual request batching, and the possibility that application-level bottlenecks will reduce the kernel-level gain.

The portability bet

Helion moves optimization policy from a collection of specialized kernels and dispatch heuristics into a PyTorch-native kernel definition plus an automated search process. The current vLLM integration remains deliberately narrow, covering Hopper, three quantization formats, and small-token CUDA Graph workloads.

Initial work with Helion’s CuteDSL backend has shown competitive GEMM performance on NVIDIA Blackwell GPUs. Development is also underway for AMD GPUs, TPUs, and a Helion backend for mixture-of-experts workloads.

Hopper deployments using the supported formats can exchange offline tuning time and configuration maintenance for measured decode-throughput gains. The hybrid dispatcher confines Helion to the small-token regime where detailed tuning is designed to pay off. Established libraries continue to handle larger inputs, providing a controlled path for testing autotuned kernel authoring in an existing inference stack.

Trending
  • No trending articles

Comments

avatar

Next Reads