Singapore's Agnes AI Builds a Reasoning Model That Uses 72% Fewer Tokens
Agnes AI's new reasoning model scores 39 on the intelligence index with top-tier coding performance, at just $0.45/$0.90 per 1M tokens

- Agnes 2.5 Pro Alpha launched from Singapore-based Agnes AI (Sapiens AI), a proprietary reasoning model built from scratch -- not a fine-tune of an open-source base.
- Scores 39 on the Artificial Analysis Intelligence Index, mid-pack among reasoning models, with a Coding Index of 58.8 that beats DeepSeek V4 Flash, GPT-5.4 mini, and Qwen3.7 Plus.
- Priced at $0.45/$0.90 per 1M input/output tokens, blending to $0.18/1M with caching -- second cheapest proprietary model at its intelligence tier.
- Uses 72% fewer reasoning tokens than GPT-5.4 mini (22k vs 78k per task), making per-task costs lower than raw pricing suggests.
- Weak on open-ended factual recall: AA-Omniscience score of -26.3 and 13.3% non-hallucination rate -- avoid unaugmented knowledge Q&A tasks.
- Available now via OpenAI-compatible API at platform.agnes-ai.com using model name
agnes-2.5-pro-alpha.
Singapore's Agnes 2.5 Pro Alpha is the latest entry from Agnes AI (operating under Sapiens AI), a homegrown Singapore lab that has been quietly building full-stack foundation models from scratch since mid-2025. This is not a fine-tune of an open-source base. Unlike many AI platforms relying heavily on overseas open-source models, Agnes built its own model architecture from the ground up, and 2.5 Pro Alpha is the most capable model that strategy has produced yet.
A new face on the reasoning leaderboard
Agnes 2.5 Pro Alpha scores 39 on the Artificial Analysis Intelligence Index, placing it well above average among other reasoning models in a similar price tier (median: 16). That index is a composite of nine evaluations spanning real-world agentic work, coding, scientific reasoning, and long-context tasks. A score of 39 puts the model mid-pack among all reasoning models tested, below the current proprietary frontier but solidly competitive for its price band.
Agnes 2.5 Pro Alpha is a proprietary model and Sapiens AI has not disclosed the model size or parameter count. What they have disclosed is a clean benchmark profile with a few clear strengths and one notable weakness.
Where it shines -- and where it doesn't
Coding is the headline strength. Agnes 2.5 Pro Alpha's Coding Index of 58.8 sits above similarly priced models including DeepSeek V4 Flash (56.2), GPT-5.4 mini (56.1), and Qwen3.7 Plus (55.9), landing just behind Nex-N2-Pro (59.1). On real-world work tasks, it scores 1,168 on GDPval-AA v2 (a benchmark that rates model performance against a human baseline of 1,000), putting it level with GPT-5.4 mini and ahead of Qwen3.7 Plus.
Academic reasoning is competitive too. Agnes 2.5 Pro Alpha scores 32% on Humanity's Last Exam (HLE) -- a notoriously hard benchmark of expert-level questions -- and 88% on GPQA Diamond (graduate-level scientific reasoning). Its HLE result leads similarly priced proprietary models including GPT-5.4 mini (27%) and GPT-5.4 nano (26%).