Together AI's ThunderAgent Doubles GPU Throughput Running 192 Agents at Once
ThunderAgent fixes KV cache thrashing in agentic inference, delivering 2.5x higher throughput and ~10x lower latency -- accepted as an ICML 2026 Spotlight paper.

- ThunderAgent is a program-aware scheduler for agentic LLM inference, accepted as an ICML 2026 Spotlight paper (top 2.2%).
- It solves KV cache thrashing -- the cascade of cache evictions that cripples engines like vLLM and SGLang when hundreds of agents run concurrently.
- On a single 8xH100 node, it delivers 2x throughput (803 vs 390 tok/s) and 6x lower latency (10.6s vs 65s) over SGLang at batch 192.
- Multi-node scaling is near-linear: 671 to 2,248 steps/min from 16 to 64 GPUs, with a lead that widens from 1.79x to 2.39x over SGLang Gateway.
- Integration requires only adding a
program_idfield to existing OpenAI-compatible API calls -- works with vLLM, SGLang, offloading, and speculative decoding. - Already adopted by NVIDIA Dynamo and SkyRL; open-source on GitHub under MIT license.
Running hundreds of AI agents concurrently is one of the hardest infrastructure problems in ML right now. Every agent alternates between heavy GPU work — generating tokens — and idle waiting while a tool runs. That rhythm is brutal for inference engines, and Together AI's ThunderAgent is the first system to fix it at the scheduler level. It was accepted to ICML 2026 as a Spotlight paper, placing it in the top 2.2% of submissions.
The GPU memory problem at high concurrency
When a model generates tokens, it stores intermediate attention computations — keys and values — in GPU memory so it doesn't have to recompute them for every new token. This KV cache is precious and finite.
In a single-agent setup, that's manageable. Run 192 agents at once and you have a serious problem. An agentic workflow alternates between two phases: the GPU-heavy reasoning phase where the model generates tokens, and the GPU-idle acting phase where the agent waits for a tool like a compiler to return. While Agent A waits on a tool call, its KV cache sits in memory doing nothing. The engine, seeing memory pressure, evicts it using a least-recently-used (LRU) policy. When Agent A's tool returns and it needs to resume, the engine recomputes the entire conversation history from scratch, which in turn evicts Agent C's cache. At high concurrency, this cascade of evictions and recomputations collapses both throughput and latency.
Why the obvious fixes fall short
- More GPU nodes: Existing multi-node routers such as SGLang Gateway pin each agent to a fixed node to preserve cache locality. Because agentic context lengths grow unpredictably, some nodes get agents with long contexts that exhaust memory while others sit idle. The overloaded nodes still thrash.
- KV cache offloading to CPU or disk: Strategies like LMCache and HiCache expand total cache capacity but only delay thrashing. When the working set of concurrent agents exceeds all available storage tiers, evictions resume and the same cycle repeats.
The root cause is structural: request-level engines never see that a series of LLM calls belongs to one longer workflow. Every call looks like an independent request to engines like vLLM or SGLang. There is no concept of a workflow, just a queue of isolated calls.
Treating workflows as schedulable programs
ThunderAgent adds that missing view. It is a lightweight scheduling layer that sits between agentic clients and inference backends, abstracting each agentic workflow as a schedulable program and tracking its execution phase, KV cache footprint, and node placement. Think of it as an OS process scheduler, but for agent workflows instead of CPU threads.

With that abstraction in place, ThunderAgent can apply program-level admission control — something no request-level engine can do:
- Monitor memory pressure on each node continuously.
- Pause low-priority workflows under pressure, reducing the number of programs competing for cache at once.
- Boost cache hit rates for the remaining active workflows — fewer programs means less eviction contention.
- Route resumed workflows through a global waiting queue to whichever node has the most available capacity, rather than pinning to a fixed node.
ThunderAgent also treats KV cache capacity across GPU HBM, CPU RAM, and disk as a unified pool, making it compatible with existing offloading strategies like HiCache. It layers on top of your existing stack rather than replacing it.
Benchmark results
At batch size 192 on a single 8xH100 node with HiCache offloading, SGLang delivers 390 tokens/s with a mean latency of 65s. ThunderAgent on the same hardware delivers 803 tokens/s with a mean latency of 10.6s — roughly 2x the throughput and 6x lower latency.
The multi-node results are more striking. Throughput scales near-linearly from 671 to 2,248 steps/min as the cluster grows from 16 to 64 GPUs. The speedup over SGLang Gateway also widens with cluster size, from 1.79x at 2 nodes to 2.39x at 8 nodes. Most distributed systems see diminishing returns as nodes are added; ThunderAgent's advantage grows.

Across diverse agentic workloads including SWE-Agent, OpenHands, and ToolOrchestra, ThunderAgent improves vLLM throughput by 1.5–3.6x. The gains generalize across the major open-source agent frameworks, not just Together AI's internal pipeline.
Who is already using it
ThunderAgent is integrated into NVIDIA Dynamo 2.0, where the program abstraction operates as a first-class scheduling unit. It is also integrated into SkyRL for agentic RL training, with a published training recipe that accelerates SWE Agent rollout by 3x on 40 H100 GPUs.
The primary use case is large-scale synthetic data generation — the kind of workload needed to train the next generation of tool-using models. Together AI uses it to generate agentic datasets like CoderForge at high concurrency. The system generalizes to any workload running many parallel agent workflows: RL rollouts, automated coding pipelines, multi-agent research infrastructure.
Adding it to your stack
The only client-side change is adding a program_id field to your existing OpenAI-compatible API calls:
# Before: standard OpenAI call
client.chat.completions.create(
model="Qwen/Qwen3-32B",
messages=messages,
)
# After: ThunderAgent-aware call
extra_body = {"program_id": "unique_workflow_id"}
client.chat.completions.create(
model="Qwen/Qwen3-32B",
messages=messages,
extra_body=extra_body,
)ThunderAgent supports vLLM and SGLang backends and is compatible with existing optimizations like quantization and speculative decoding. Install it, point it at your existing backend, and route traffic through port 9000 instead of 8000. The backend itself does not change.
One caveat: the gains are most dramatic at high concurrency. Below roughly 64 parallel workflows, the baseline engines are likely adequate. ThunderAgent is purpose-built for the scale where KV cache thrashing becomes the dominant bottleneck — RL training rollouts and large synthetic data pipelines, not single-user chatbots.
The code is open-source under MIT and the paper is on arXiv. The core insight — treating agent workflows as schedulable programs rather than bags of independent requests — is simple enough that it's surprising no one shipped it sooner, and consequential enough that it unlocks a whole class of workloads that were previously too expensive to run at scale.