BitterSecurity's Decepticon Hits 98% on Web Hacking Benchmark With 16 AI Agents

An open-source LangGraph agent swarm runs full red-team kill chains inside a Kali sandbox, scoring 98% on XBOW's validation benchmarks.

·
·
·
BitterSecurity's Decepticon Hits 98% on Web Hacking Benchmark With 16 AI AgentsPRO
  • Decepticon is an open-source autonomous red-team agent built on LangGraph, now at 5.7k stars on GitHub.
  • Reports 98.08% pass rate (102/104) on XBOW validation benchmarks across easy, medium, and hard tiers.
  • 16 specialist agents cover recon, exploitation, post-exploitation, Active Directory, cloud, reversing, and smart contracts.
  • Runs commands in persistent tmux sessions inside an isolated Kali sandbox, enabling interactive tools like msfconsole and sliver.
  • Tier-based model routing supports Anthropic, OpenAI, Gemini, DeepSeek, Ollama, plus OAuth for Claude Max and ChatGPT Pro.
  • Installs via one-line script or pip install decepticon, with MCP integration for Claude Code and Codex.

Decepticon runs autonomous red-team workflows in Kali sandboxes

Decepticon is an Apache-2.0 project that uses LangGraph, a framework for stateful LLM workflows, to plan and execute authorized security tests. It runs offensive tools inside containerized Kali Linux environments and coordinates reconnaissance, exploitation, privilege escalation, lateral movement, and command-and-control operations.

The repository has attracted more than 5,700 GitHub stars, while the maintainers report a 98.08% pass rate on a public web-exploitation benchmark. The combination matters for developers because Decepticon addresses practical problems that simpler agents often leave unresolved: interactive shell state, tool orchestration, model routing, network separation, context management, and engagement planning.

102 solves, with a narrow scope

Decepticon’s published benchmark results show 102 successful challenges out of 104 in the XBOW validation suite. The public suite tests AI pentesting agents against web-exploitation tasks at three difficulty levels.

Decepticon’s reported XBOW results
Difficulty Solved Pass rate
Easy 45 of 45 100%
Medium 50 of 51 98.04%
Hard 7 of 8 87.5%
Total 102 of 104 98.08%
Chart showing Decepticon solving 102 of 104 XBOW challenges for a 98.08% pass rate
The maintainers’ summary of Decepticon’s XBOW benchmark run.

The maintainers also publish a benchmark comparison covering Strix, PentestGPT, MAPTA, Cyber-AutoAgent, and the commercial XBOW agent. Cross-project comparisons depend on the model, token budget, retry policy, tool access, benchmark revision, and scoring method, so independent reproduction requires those variables to be pinned.

The result measures performance on a specific collection of web vulnerabilities. Its conclusions remain limited to that domain and cannot predict performance in an Active Directory forest, cloud account, smart contract, internal network, or custom application.

Planning, shells, and two networks

Decepticon’s architecture combines three design choices that support longer engagements:

Pro article

This story is for Pro members

You've reached the end of the free preview. Upgrade to AlphaSignal Pro to read the full article - and everything else behind the paywall.

Trending
  • No trending articles

Comments

avatar

Next Reads