BitterSecurity's Decepticon Hits 98% on Web Hacking Benchmark With 16 AI Agents
An open-source LangGraph agent swarm runs full red-team kill chains inside a Kali sandbox, scoring 98% on XBOW's validation benchmarks.
- Decepticon is an open-source autonomous red-team agent built on LangGraph, now at 5.7k stars on GitHub.
- Reports 98.08% pass rate (102/104) on XBOW validation benchmarks across easy, medium, and hard tiers.
- 16 specialist agents cover recon, exploitation, post-exploitation, Active Directory, cloud, reversing, and smart contracts.
- Runs commands in persistent tmux sessions inside an isolated Kali sandbox, enabling interactive tools like msfconsole and sliver.
- Tier-based model routing supports Anthropic, OpenAI, Gemini, DeepSeek, Ollama, plus OAuth for Claude Max and ChatGPT Pro.
- Installs via one-line script or
pip install decepticon, with MCP integration for Claude Code and Codex.
Decepticon runs autonomous red-team workflows in Kali sandboxes
Decepticon is an Apache-2.0 project that uses LangGraph, a framework for stateful LLM workflows, to plan and execute authorized security tests. It runs offensive tools inside containerized Kali Linux environments and coordinates reconnaissance, exploitation, privilege escalation, lateral movement, and command-and-control operations.
The repository has attracted more than 5,700 GitHub stars, while the maintainers report a 98.08% pass rate on a public web-exploitation benchmark. The combination matters for developers because Decepticon addresses practical problems that simpler agents often leave unresolved: interactive shell state, tool orchestration, model routing, network separation, context management, and engagement planning.
102 solves, with a narrow scope
Decepticon’s published benchmark results show 102 successful challenges out of 104 in the XBOW validation suite. The public suite tests AI pentesting agents against web-exploitation tasks at three difficulty levels.
| Difficulty | Solved | Pass rate |
|---|---|---|
| Easy | 45 of 45 | 100% |
| Medium | 50 of 51 | 98.04% |
| Hard | 7 of 8 | 87.5% |
| Total | 102 of 104 | 98.08% |
The maintainers also publish a benchmark comparison covering Strix, PentestGPT, MAPTA, Cyber-AutoAgent, and the commercial XBOW agent. Cross-project comparisons depend on the model, token budget, retry policy, tool access, benchmark revision, and scoring method, so independent reproduction requires those variables to be pinned.
The result measures performance on a specific collection of web vulnerabilities. Its conclusions remain limited to that domain and cannot predict performance in an Active Directory forest, cloud account, smart contract, internal network, or custom application.
Planning, shells, and two networks
Decepticon’s architecture combines three design choices that support longer engagements:
This story is for Pro members
You've reached the end of the free preview. Upgrade to AlphaSignal Pro to read the full article - and everything else behind the paywall.