Together AI's ParallelKernelBench Reveals Top Models Fail 70% of Multi-GPU Tasks
Together AI's ParallelKernelBench reveals frontier models solve under a third of 87 real multi-GPU kernel problems, exposing a critical blind spot in AI coding benchmarks.

- New benchmark: Together AI releases ParallelKernelBench, 87 multi-GPU CUDA kernel problems drawn from real production codebases.
- Frontier models struggle: GPT-5.5, the top performer, solves only 28/87 problems correctly in one shot, and only 22 beat the naive PyTorch + NCCL baseline.
- Root cause: Models fail not on syntax but on reasoning about rank coordination, data partitioning, and GPU-to-GPU transfer mechanism selection.
- Agentic loops help modestly: Gemini 3 Pro with iterative feedback improved from 24 to 35 correct solutions, but plateaued after ~20 steps.
- Surprise wins: A few generated kernels beat all publicly available implementations, including one for NVIDIA NeMo-RL's GRPO training loop with no prior optimized reference.
- Open source: Code on GitHub, problems on HuggingFace, free to use and contribute to.
Every major coding benchmark for LLMs tests single-GPU CUDA kernels. But production AI infrastructure doesn't run on one GPU , it runs on clusters, where the real bottleneck is how fast data moves between GPUs, not how fast a single chip computes. ParallelKernelBench (PKB) is a new open-source benchmark from Together AI's Frontier Performance team that tests exactly this, and the results are a wake-up call.
The gap nobody was measuring
In production, communication overhead can account for over 20% of inference latency, and that gap keeps widening as compute scales faster than interconnect bandwidth. Yet LLMs have made progress on GPU kernel generation, but that progress has mostly been measured on a single GPU. PKB is the first benchmark to directly address this mismatch.
ParallelKernelBench offers a benchmark and evaluation framework for multi-GPU kernel generation, including 87 problems from real codebases where the task is replacing PyTorch + NCCL with a CUDA kernel that moves data directly over NVLink. NCCL (NVIDIA Collective Communications Library) is the standard library for GPU-to-GPU communication , PKB asks models to bypass it entirely and write lower-level kernels that talk directly over NVLink, the high-bandwidth interconnect between GPUs in a server node.
What makes multi-GPU so much harder
It's not just about knowing CUDA syntax. The challenge is fundamentally different from single-GPU work in three ways:
- The design space explodes. Practitioners compose tensor, expert, data, context, and sequence parallelism to fit the hardware, and each combination creates a different communication pattern.
- The performance model changes. On a single GPU, you optimize for compute throughput and memory bandwidth. On multiple GPUs, the bottleneck is the interconnect , and the roofline model (a framework for estimating peak achievable performance) looks completely different.
- New low-level choices appear. How do you move data between GPUs? Through the copy engine, TMA (Tensor Memory Accelerator), SM load/store, or NVLS (NVLink SHARP, a hardware-accelerated collective operation)? Each choice has different performance characteristics, and picking the wrong one leaves performance on the table.
87 problems, pulled from real systems
To make sure the 87 problems cover the real space of production parallelism types, the team built them from a taxonomy of distributed workloads , identifying the major ways models get sharded (tensor, context, data, expert, sequence, and FSDP/ZeRO) and choosing problems from the codebases of systems like Megatron-LM, DeepSpeed, DeepEP, TensorRT-LLM, and NeMo-RL, as well as non-LLM workloads like GNN routing, distributed FFTs, and Gaussian splatting.
Each problem provides a task, hardware topology, and PyTorch + NCCL reference; the model generates a custom CUDA kernel that is evaluated for correctness, wall-clock speedup, and communication roofline. Because PKB references are written in standard PyTorch + NCCL, the benchmark is not tied to any single hardware generation.
The numbers are brutal
Fewer than 31% of the 87 problems were solved correctly , and only a subset of those offered performance improvements over baseline implementations. Here's how the top models stack up at pass@1 (single attempt) vs. pass@3 (best of three):
| Model | Correct (pass@1 → @3) | Faster than baseline (fast1@1 → @3) |
|---|---|---|
| GPT-5.5 | 28 → 36 | 22 → 27 |
| Gemini 3 Pro | 24 → 30 | 12 → 19 |
| Claude Opus 4.7 | 20 → 31 | 12 → 20 |
| GLM-5.1 | 6 → 7 | 2 → 3 |
| DeepSeek V4 Pro | 2 → 3 | 0 → 1 |
Even the best result , GPT-5.5 with three attempts , tops out at 31% on the fast1@3 metric, which counts solutions that are both correct and faster than the naive PyTorch + NCCL baseline.
Where models succeed, and where they fall apart
Successes are concentrated in familiar patterns: collective primitives, tensor-parallel GEMMs, and Ulysses-style context parallelism. These are the parts of the multi-GPU stack most visible in open-source code, so models likely have stronger priors for them.
The failure patterns are revealing. The failures suggest a deeper issue than CUDA syntax. Weaker models often fail to compile, but stronger reasoning models frequently produce kernels that compile and return incorrect results. The hard part is reasoning about rank coordination, data partitioning, and collective ordering.
Generated kernels also use a very narrow set of communication mechanisms. Most rely on copy engines or SM load/store instructions, while more specialized mechanisms such as TMA and NVLS are almost absent. In many cases, models do not choose the mechanism needed for peak performance. This is a training data problem as much as a reasoning problem , these newer hardware primitives are simply underrepresented in the open-source code that models train on.
Agentic loops help, but not enough
The team also tested an agentic setup: wrapping Gemini 3 Pro in a loop with access to a terminal, compiler output, correctness tests, and speed measurements , letting it iterate on its own kernels. Gemini 3 Pro improved from 24 correct solutions to 35 out of 87, with 26 beating the baseline. But after roughly 20 refinement steps, performance plateaued. Feedback helps models fix syntax errors and shape mismatches, but the deeper reasoning failures around rank coordination and communication ordering remain unsolved.
The surprise: a few generated kernels beat everything public
The most striking finding isn't the failures , it's the occasional win. A few cases saw models produce kernels faster than anything publicly available, including one for NVIDIA NeMo-RL's GRPO training loop, which has no prior optimized public reference.
Three standout examples were verified for correctness over 4 H100 GPUs and 100 randomized runs:
- NeMo vocab-parallel log-prob with top-k/top-p filtering (Gemini 3 Pro): The generated kernel skips full-vocabulary NCCL gathers entirely, using symmetric memory to permute shards inline while fusing log-softmax, token extraction, and target gather into a single warp-shuffle reduction.
- Hyena context parallelism forward pass (GPT-5.5): For the Hyena operator (a long-convolution alternative to attention), the kernel packs inputs into one symmetric allocation and streams remote slices over NVLink, computing gating and reindexing in a unified pass.
- SAM 3 all-gathered mask IoU suppression (GPT-5.5): Cross-GPU duplicate suppression for video segmentation, collapsing variable-length all-gather collectives plus a dense matmul into a pipeline of symmetric-memory kernels that bitpack masks and compute pairwise overlap with hardware popcount.
Such wins highlight the potential of AI-driven kernel optimization, especially in niche areas where no optimized public references exist. However, these successes remain exceptions rather than the norm.
What this means for the field
PKB exposes a fundamental assumption that needs updating: that progress on single-GPU kernel generation benchmarks like KernelBench translates to multi-GPU capability. It doesn't. The reasoning required for distributed kernels , coordinating ranks, ordering collectives, choosing the right NVLink transfer mechanism , is qualitatively different, and current models lack the training data and reasoning patterns to handle it reliably.
The benchmark is free to use. The code is on GitHub, the problems are on HuggingFace, and the paper is available on alphaXiv. PKB currently covers intra-node NVLink; the team plans to extend to inter-node fabrics (RoCE, InfiniBand) and other accelerators. It also already accepts Triton solutions alongside raw CUDA, leaving the door open for higher-level kernel languages to close the gap.