Artificial Analysis' AA-AgentPerf-Local Shows RTX 5090 Beating 128GB Systems by 5x
Artificial Analysis open-sources AA-AgentPerf-Local, replaying real agent trajectories on laptops and workstations to benchmark local inference performance.
- Artificial Analysis open-sourced AA-AgentPerf-Local, an Apache-2.0 tool for benchmarking local agent inference.
- Replays 8 recorded agent tasks with 168 turns, context grows to ~56K tokens, exact token counts per turn.
- Initial leaderboard covers DGX Spark, Ryzen AI Halo, MacBook Pro M5 Pro, and RTX 5090.
- RTX 5090 finishes Qwen3.8-27B in 4.9 min versus 24.2 min for DGX Spark, driven by 1,792 GB/s bandwidth.
- DGX Spark beats Ryzen AI Halo by 1.4-1.7x despite similar specs, thanks to CUDA maturity and prefill compute.
- Works with llama.cpp, vLLM, SGLang, LM Studio; optional Docker mode runs real tool calls.
AA-AgentPerf-Local benchmarks recorded agent conversations on local hardware
Artificial Analysis has released AA-AgentPerf-Local, an open-source benchmark for measuring how quickly local hardware serves AI agents. Its initial leaderboard compares four systems across long, multi-turn workloads that capture growing context, cache reuse, prefill delays, and token generation.
The GitHub repository is licensed under Apache 2.0. The benchmark sends recorded conversations to an OpenAI-compatible model server, then reports throughput and latency. Answer quality falls outside its scope.
Replaying the agent loop
AA-AgentPerf-Local sends the full conversation history with every model request, matching the structure of common agent APIs. The context reaches roughly 56,000 tokens near the end of a session, giving inference servers repeated prefixes to cache while forcing them to process each turn’s new input.
| Setting | Default behavior |
|---|---|
| Workload | Eight recorded agent tasks containing 168 model turns |
| Output length | Each turn generates its recorded number of tokens, keeping the work identical across systems |
| Context window | 65,536 tokens at batch size 1 for a full run and every catalog profile |
| Tool execution | Skipped by default to isolate inference; --tool-mode live runs recorded shell commands inside Docker containers |
| Benchmark client | Python by default, with an experimental Rust client for high-concurrency tests |
An attached run uses run --base-url to test an existing OpenAI-compatible server. A managed run uses managed-run to download a pinned model, launch the configured server, and execute the benchmark automatically.
Ollama cannot honor ignore_eos, which the exact output policy requires. The harness detects Ollama before execution and switches to a recorded policy that reports normalized end-to-end latency. Those estimates are not directly comparable with exact-policy results.
The RTX 5090 leads inside 32 GB
The launch leaderboard covers the NVIDIA DGX Spark with 128 GB of unified memory, AMD Ryzen AI Halo with 128 GB, MacBook Pro M5 Pro with 64 GB, and NVIDIA GeForce RTX 5090 with 32 GB of VRAM. Together, the configurations span CUDA, ROCm, Vulkan, and Metal.
Artificial Analysis tested Qwen3.5-9B, Qwen3.8-27B, Qwen3.6-35B-A3B, and Ling 3.0 Flash, which has 124 billion total parameters and 5.1 billion active parameters. Every model used 4-bit quantization.
For Qwen3.8-27B, the benchmark results show the RTX 5090 completing the workload several times faster than the unified-memory systems.
| System | Memory | Completion time |
|---|---|---|
| NVIDIA GeForce RTX 5090 | 32 GB | 4.9 minutes |
| NVIDIA DGX Spark | 128 GB | 24.2 minutes |
| AMD Ryzen AI Halo | 128 GB | 34.5 minutes |
| MacBook Pro M5 Pro | 64 GB | 37.6 minutes |
The RTX 5090 has 1,792 GB/s of memory bandwidth, while the unified-memory systems provide 256 to 307 GB/s. That advantage drives its lead when the model, runtime state, and key-value cache fit within 32 GB.
The DGX Spark and Ryzen AI Halo launched at the same $4,000 price and pair 128 GB of unified memory with similar bandwidth. The Spark ran 1.4 to 1.7 times faster on three of the four models, despite a bandwidth advantage of only 7 percent. Artificial Analysis attributes the gap to stronger low-precision compute during prefill and a more mature CUDA stack for mixture-of-experts routing and speculative decoding.
Long inputs expose the prefill cost
Each agent request begins with prefill, the stage in which the model reads prompt tokens and builds the key-value cache used during generation. Prefix caching can reuse earlier conversation tokens, while new user messages and tool output still require processing before the model emits its next token.
Across the recorded workload, the servers retrieved 73 to 93 percent of prompt tokens from the cache. Processing the remaining new input still consumed 22 to 41 percent of total completion time for Qwen3.8-27B. After a tool returned a large file, prefill delayed generation by as much as 24 seconds on the Ryzen AI Halo and 39 seconds on the MacBook Pro, compared with roughly 2 seconds on the RTX 5090.
Every leading configuration used speculative decoding through MTP, DFlash, or DSpark. These methods predict multiple tokens and validate them together, reducing the number of full model passes. They increased effective Qwen3.8-27B decode speed by 30 to 120 percent beyond the single-token rate implied by raw memory bandwidth. The repository publishes all 14 serving configurations.
Capacity, bandwidth, and active weights shape the choice
- Capacity: The model weights, 65,536-token context, and key-value cache must fit in available memory. Systems with 128 GB can host models that exceed the RTX 5090’s 32 GB ceiling.
- Bandwidth: Models that fit on the RTX 5090 benefit from its substantially higher memory throughput.
- Prefill compute: Large tool responses and growing conversations reward hardware with strong low-precision processing.
- Software stack: Kernel quality, mixture-of-experts routing, caching, and speculative decoding can create performance gaps between systems with similar memory specifications.
- Benchmark mode: Default runs isolate model inference, while live tool mode captures more of the agent’s end-to-end execution path.
Mixture-of-experts models also changed the speed calculation. Qwen3.6-35B-A3B activates only a subset of its total parameters for each token and ran 2.5 to 3.3 times faster than the dense 27B model on every tested system. The result shows why active parameter count belongs alongside total model size when estimating local inference performance.
Running the benchmark
Installation uses a single command on macOS, Linux, or Windows:
uv tool install agentperf-localDevelopers can connect the client to an existing server or let managed mode provision a pinned model and server configuration. Comparable leaderboard submissions require the same model, quantization, context length, output policy, and tool mode.
The launch dataset covers four machines, four 4-bit models, and batch size 1. It does not measure answer quality, energy use, thermal throttling, acoustic output, or sustained multi-user serving, so deployment decisions may require additional tests for those constraints.