Alibaba's CommerceAgentBench Reveals Top AI Agents Fail 38% of Real Business Tasks

Alibaba's Accio team open sourced a 107-task benchmark that forces agents to complete real e-commerce workflows inside stateful replicas of Shopify, Gmail, Stripe and more.

·
·
Alibaba's CommerceAgentBench Reveals Top AI Agents Fail 38% of Real Business TasksPRO
  • Accio at Alibaba International open sourced CommerceAgentBench, a 107-task stateful agent benchmark.
  • Tasks span 53 CLI, 28 browser, 16 file, and 10 API/MCP workflows in Dockerized real-app replicas.
  • Best score: Claude Opus 5 at 61.7% pass rate; Gemini 3 Flash trails at 29.0%.
  • A task passes only when every required verifier check passes, so partial work scores zero.
  • Replicas include Shopify, Alibaba.com, Freightos, Gmail, Stripe, Jira, Notion, Amazon SP-API.
  • Apache 2.0 + CC-BY-4.0 licensed, with a live leaderboard and community mock contributions welcome.

A new open source benchmark from Alibaba International is trying to answer a question most evaluations sidestep: can an AI agent actually finish a job, not just describe how to do one? CommerceAgentBench, released by the Accio team, drops agents into containerized replicas of real business software and grades them on whether the final state of the system matches what the task required.

The suite is Apache-licensed, ships a Docker runtime, and comes with reference results across 13 frontier models. Even the strongest model on the leaderboard finishes fewer than two thirds of the tasks.

What ships in the repo

CommerceAgentBench evaluates whether an agent can complete long-horizon business workflows rather than answer questions about them. Tasks cover browser operations, native-style CLI tools, API/MCP workflows, document and spreadsheet production, public-web research, supplier analysis, product publishing, logistics, and commerce operations. Every task runs in a fresh container and is graded by its own deterministic or LLM-assisted verifier.

The composition breaks down as follows:

  • 107 tasks total: 53 CLI, 28 browser, 16 file, and 10 API/MCP
  • Three capability slices: 65 text-only, 20 browser-text-capable, and 22 vision-required
  • Local mock services that model SaaS, commerce, messaging, document, and operational systems without requiring production accounts
  • Every run preserves the resolved configuration, trajectory, verifier result, artifacts, logs, and container metadata

Pro article

This story is for Pro members

You've reached the end of the free preview. Upgrade to AlphaSignal Pro to read the full article - and everything else behind the paywall.

Trending
  • No trending articles

Comments

avatar

Next Reads