Sapiens AI's Agnes 2.5 Pro Beta Triples Agentic Performance at Bargain Prices
Singapore lab Sapiens AI pushes its Agnes 2.5 Pro model to 49 on Artificial Analysis Intelligence Index through agentic gains, but at roughly double the token cost.

- Agnes 2.5 Pro Beta scores 49 on Artificial Analysis Intelligence Index, up 9 points from Alpha.
- Agentic Index jumps from 25 to 44, with τ³-Banking nearly tripling from 12% to 36%.
- GDPval-AA v2 Elo rises to 1456, ahead of MiniMax-M3 and near GPT-5.5.
- AA-Omniscience gains come from abstention: attempts drop from 94% to 45%, accuracy halves to 17%.
- Uses 50k output tokens per task, roughly 2x the Alpha version's consumption.
- Pricing: $0.10/$0.30/$0.01 per 1M input/output/cache tokens, 1M context, via Agnes AI API.
A Singapore-based lab just moved from mid-pack to knocking on the frontier's door. Agnes 2.5 Pro Beta, built by Sapiens AI, jumped nine points on the Artificial Analysis Intelligence Index, landing at 49 versus 40 for the previous Alpha release. That puts it just under Gemini 3.5 Flash (high) and GPT-5.6 Luna (max), both at 52, and above MiniMax-M3 at 45.
The interesting part is where the gains came from and what they cost. Almost all the improvement is concentrated in agentic tasks, and the model roughly doubled its token consumption to get there.
Agentic work is doing the heavy lifting
The Artificial Analysis Agentic Index, which tracks how well models handle tool use and multi-step workflows, climbed from 25 to 44. Agnes now sits just behind Gemini 3.7 Flash (high) at 45 and comfortably ahead of both Gemini 3.5 Flash (high) at 40 and MiniMax-M3 at 36.
Two specific agentic benchmarks tell the story:
- τ³-Banking, which tests tool-calling in a simulated banking environment, nearly tripled from 12% to 36%.
- GDPval-AA v2, which scores models on real-world knowledge work using an Elo system with a human baseline of 1000, jumped from 1171 to 1456. That beats MiniMax-M3 at 1384 and trails GPT-5.5 (xhigh) at 1489 and Gemini 3.7 Flash (high) at 1527.
Frontier reasoning evaluations moved more modestly. Humanity's Last Exam went from 34% to 38%, GPQA Diamond from 88% to 91%, and CritPt (a physics reasoning benchmark) from 11% to 16%. Solid, but not the headline story.
The Omniscience number is doing something sneaky
On paper, the AA-Omniscience score improved from -25 to -11. That benchmark rewards correct answers, penalizes hallucinations, ignores refusals, and produces scores between -100 and 100.
Look closer and the improvement is misleading. Agnes 2.5 Pro Beta attempted only 45% of questions versus 94% for the Alpha version. Its hallucination rate did drop meaningfully from 88% to 33%, but its actual accuracy on AA-Omniscience halved from 33% to 17%. In practical terms, the model learned to shut up more often rather than learning more facts. That behavior helps in production settings where wrong answers cost more than missing answers, though it reflects better calibration rather than new knowledge.
Intelligence at a token premium
Reasoning models pay for capability in output tokens, and Agnes 2.5 Pro Beta pays a lot. It burns 50k output tokens per Intelligence Index task, more than double the 24k the Alpha version used. That exceeds Qwen3.8 27B (xhigh) at 47k and GLM-5.3 (max) at 41k.
Pricing softens the blow. Sapiens AI charges $0.10 per million input tokens, $0.30 per million output tokens, and $0.01 per million cache-hit tokens, well below the median of $0.25 input and $0.90 output for comparable reasoning models. On the artificialanalysis.ai listing, output throughput clocks in around 160 tokens per second, above the 102 t/s median for peers in its price band.
Specs and availability
| Attribute | Value |
|---|---|
| Context window | 1M tokens |
| Max output | 65k tokens |
| Input modalities | Text, image |
| Output modality | Text |
| Cache discount | 90% |
| Access | Agnes AI first-party API only |
The API is OpenAI-compatible, so migration usually means swapping the base URL and key rather than rewriting client code.
Why Singapore matters here
Sapiens AI runs the Agnes AI platform, an omni-modal gateway that previously drew attention for being the first Singapore-born model on the leaderboard, and for offering free access across text, image, and video models to more than three million users. The lab trains its full-modality foundation models in house rather than fine-tuning open weights from Meta or Alibaba.
The takeaway for anyone shopping models: if you need cheap agentic capability with real tool-use chops and can tolerate verbose reasoning traces, this is a serious option at a price point where most competitors sit two tiers lower on capability. If you need reliable factual recall or minimal token spend, the tradeoffs are real, and the Omniscience abstention behavior is worth testing on your own knowledge-heavy prompts before you commit.