Prime Intellect Launches Prime Inference, Serving 600B Tokens Daily

Prime Intellect launches its inference platform serving 600B tokens daily, with NVFP4 KV compression, prefill/decode disaggregation, and sub-second TTFT on GB200 NVL72.

·
·
·
  • Prime Intellect launches Prime Inference, processing ~600B tokens/day across multiple datacenters
  • GLM-5.3 on GB200 NVL72 serves 66 concurrent sessions at 101 tokens/sec/user
  • NVFP4 KV compression shrinks cache rows from 576 to 352 bytes, 50% more capacity
  • Prefill/decode disaggregation via Dynamo cut p90 inter-token latency by nearly 40%
  • BLHNC block-major KV layout reduced NVLink transfer descriptors ~10x, cut transfer time 47%
  • OpenAI-compatible endpoint at api.pinference.ai/api/v1, serverless and reserved capacity available

Prime Intellect has launched Prime Inference, the production serving stack behind its reinforcement learning workloads and dedicated customer deployments. The platform processes nearly a trillion tokens every day on its own compute. It is now offering the same infrastructure through serverless endpoints for variable traffic and reserved capacity for sustained workloads on frontier open-weight models.

Prime built the stack to support reinforcement learning rollouts, synthetic data generation, evaluations, and long-running coding agents. Its first public deployment, GLM-5.3, is available through OpenRouter and Prime’s API. Prime reports that the endpoint ranks among the fastest OpenRouter deployments for the model and has a near-zero tool-call error rate.

A familiar API, two capacity models

The service implements an OpenAI-compatible API, allowing existing clients to connect by changing their base URL to https://api.pinference.ai/api/v1. Prime also provides a command-line interface for direct testing:

code
prime inference chat 'z-ai/glm-5.3' "Write a haiku about KV caches."
  • Public model: Low-latency GLM-5.3 serving backed by production service-level agreements
  • Hardware: NVIDIA Blackwell accelerators, with Vera Rubin infrastructure planned
  • Resilience: Automatic failover across data centers when a deployment becomes unhealthy
  • Capacity: Serverless endpoints for variable demand and reserved deployments for sustained workloads
  • Software: NVIDIA Dynamo, vLLM, Mooncake, and FlashInfer, with selected improvements contributed upstream

Prime plans to add batch and asynchronous inference for lower-cost offline jobs. It also intends to support one-click deployment of fine-tuned models from Prime training runs onto reserved capacity.

Why agents strain inference servers

A typical turn in Prime’s long-context agent workload adds about 6,000 tokens to a 140,000-token prompt while reusing most of the conversation history. The server stores intermediate attention data from that history in a key-value cache, or KV cache, so it does not need to process the entire conversation again.

Under concurrent load, returning sessions with reusable caches share GPUs with new requests carrying long, uncached prompts. Processing a cold prompt can delay token generation for active sessions, while moving large KV caches between workers can consume memory bandwidth and stall requests.

Split the prompt from generation

Prime separates prompt processing, called prefill, from token generation, called decode. Dedicated GPU groups handle each phase. NVIDIA Dynamo routes requests, vLLM runs the model, and NVIDIA’s NIXL transfer library moves the computed KV cache to a decoder after prefill completes.

Prime’s tests found that this design reduced 90th-percentile inter-token latency by nearly 40%. On GB200 NVL72 systems running GLM-5.3, one prefill group for every four decode groups supported 66 concurrent sessions per prefill group at 101 generated tokens per second per active user.

Performance frontier for GLM-5.3 on GB200 NVL72 across prefill-to-decode ratios
GLM-5.3 throughput and latency across prefill-to-decode configurations on GB200 NVL72.

Four-bit caches make room

Tensor parallelism across four GPUs, known as TP4, produced the lowest inter-token latency during decoding. Each GPU rank still held a full copy of every request’s latent KV cache, limiting the number of tokens that could remain cached in memory.

Prime compressed GLM-5.3’s multi-head latent attention cache with NVFP4, a four-bit format that stores one FP8 scale for each group of 16 values. Each cache row shrank from 576 to 352 bytes. At the same memory budget, decoder capacity increased by about 50%, from 1.09 million to 1.63 million cached tokens.

A staged implementation first unpacked NVFP4 values into FP8 buffers and then ran attention, adding an extra GPU memory round trip for every layer. Prime wrote a native sparse-attention kernel that loads only the rows selected by the sparse indexer and unpacks them on-chip. NVFP4 remains the storage format, while the attention calculation uses FP16 operands with FP32 accumulation. Three implementation choices drove the improvement:

  1. A three-slot pipeline that lets one warp group unpack rows while other groups score and accumulate earlier stages
  2. Dynamic sizing of streaming multiprocessor clusters according to the token batch size
  3. Zero-filled empty entries in the 2,048-position list, cutting a 35-token attention launch from about 41 microseconds to 20.6 microseconds
Attention-kernel latency for 15 query tokens per launch on GB200
Attention path Latency
Native NVFP4 About 12 microseconds
Staged NVFP4 17.7 microseconds
FP8 13.7 microseconds

Prime is contributing the native NVFP4 kernel to FlashInfer as an experimental operation.

Latency measurements across successive versions of Prime's NVFP4 decode kernel
Successive kernel changes reduced NVFP4 attention latency on the decode path.

Prime’s early tests found that transferring KV data across multi-node NVLink added about 292 milliseconds to time to first token compared with InfiniBand. A single request containing 200,000 tokens could trigger 32,000 small device-to-device copies on one TP4 rank, flooding the interconnect with transfer descriptors.

The vLLM community addressed the fragmentation with BLHNC, a block-major KV layout that stores data from multiple layers contiguously for each block. Although the layout uses blocks one-sixteenth the previous size, it reduced the descriptor count from 19,559 to about 1,940. Mean transfer time fell from 146 to 78 milliseconds, a reduction of roughly 47%.

More cache, shorter queues

On prefill GPUs, Prime selected DEP8, which combines eight data-parallel attention ranks with expert parallelism. The alternative TEP8 layout replicated each request’s KV data across all eight ranks under Prime’s MLA cache configuration. DEP8 allowed different ranks to cache different requests, providing about five times more usable prefix-cache capacity on the same hardware.

Prime’s scheduler often had cached KV data ready before the corresponding request could enter a running batch. Reducing the prefill budget from 8,000 to 4,000 tokens per GPU shortened that delay. Median queue time fell from 550 to 110 milliseconds, reducing median time to first token by about 20%.

Grammar guards for tool calls

Dynamo could silently discard calls to undeclared tools and return a normal stop condition, leaving an agent with no action to execute and no error to handle. Such failures are especially damaging in coding and research agents, where each step may depend on a valid structured tool call.

Prime contributed a structural-tag builder to Dynamo so vLLM’s xgrammar engine can enforce the tool-call schema during generation. The grammar masks tokens that would create invalid structure before the model emits them. The team also fixed parser bugs that corrupted literal < characters in generated code and file contents.

The open-model feedback loop

Prime’s broader goal is to connect production serving with training, allowing traces from deployed agents to become data for later reinforcement learning runs. That loop requires an inference layer capable of serving long contexts reliably while preserving the structured outputs that agents depend on.

The engineering results show how memory layout, cache compression, scheduling, transfer granularity, and constrained decoding shape agent performance alongside GPU compute. Open-weight teams can now access those optimizations through a public GLM-5.3 endpoint or reserved deployments. The published measurements remain vendor-run and model-specific, so performance will vary with context length, concurrency, prompt reuse, and model architecture.

Trending
  • No trending articles

Comments

avatar

Next Reads