3,622Anthropic's Claude Code Auto Mode Catches Dangerous Commands 89% of the TimeClaudeDevs·Development·1 hr ago·
3,151OpenAI's Astra Becomes First AI Flagged as Critically Dangerous Before ReleaseOpenAI·Security·1 hr ago·
1,353OpenAI's Codex Security Review Catches 74% of Real Bugs Semgrep MissesOpenAI Developers·Development·22 hrs ago·
1,020MiniMax Rebuilds Code 2.0 on Pi Agent, Slashing Latency by 90%MiniMax_Agent·Development·9 hrs ago·
769Artificial Analysis Patches Intelligence Index to Crown Claude Opus 5Artificial Analysis·Benchmarks·23 hrs ago·
388Epoch Hides a Game's Identity to Stop AI Labs From Cheating BenchmarksEpoch AI·Benchmarks·23 hrs ago·
374Prime Intellect Ships Multi-Agent RL Training Into Its Open-Source StackPrime Intellect·Agents·2 hrs ago·
292Google's Gemini 3.6 Flash Hits 91.2% on the World's Hardest Reasoning TestARC Prize·Benchmarks·23 hrs ago·
125Artificial Analysis Rebuilds Its Image Arena to Rank Models by Real WorkArtificial Analysis·Benchmarks·2 hrs ago·