4,028Anthropic Opens Claude Mythos 5 to Enterprise Teams for Hunting Code VulnerabilitiesClaude·Security·4 hrs ago·
128Artificial Analysis' MLCR-AA Shows Most AI Models Fail Medical ReasoningArtificial Analysis·Benchmarks·5 hrs ago·
113NovaSky's IsoExec Fixes the Hidden Math Bug Corrupting AI Training RunsvLLM·Post Training·6 hrs ago·
158Artificial Analysis' Speech Arena Reveals Voice AI's Uncomfortable Split Brain ProblemArtificial Analysis·Audio·7 hrs ago·
10,865DeepSeek's V4-Flash-Vision-Exp Quietly Challenges Anthropic's Opus on Multimodal Agent TasksDeepSeek·Api·12 hrs ago·
264Sakana AI's Namazu Beats Google Translate and DeepL on Japanese Cultural NuanceSakana AI·Llms·21 hrs ago·
206Alibaba's Qwen-Image-3.0-Pro Jumps 83 Elo Points but Drops Open WeightsArtificial Analysis·Image·1 day ago·
4,974OpenAI's GPT-Image-2 Now Generates Transparent Backgrounds in a Single API CallOpenAI Developers·Image·1 day ago·
115Kaggle's Adversarial Customer Service Benchmark Pits AI Against Identity ThievesKaggle·Benchmarks·1 day ago·