Poolside's Laguna S 2.1 Beats 1.6T Models at Coding on a Single Desktop

Poolside's 118B open-weight coding model beats models 5-25x its size on long-horizon agentic tasks, runs on a single DGX Spark, and ships free under OpenMDW-1.1

·
·
AuthorPoolside
Read2 min
  • Poolside releases Laguna S 2.1, a 118B MoE model (8B active params/token) with a 1M token context window, fully open under OpenMDW-1.1.
  • Scores 70.2% on Terminal-Bench 2.1 and 78.5% on SWE-Bench Multilingual, beating models up to 25x its size including DeepSeek-V4-Pro Max and NVIDIA Nemotron 3 Ultra.
  • On DeepSWE (the hardest long-horizon benchmark), scores 40.4% -- compared to DeepSeek-V4-Pro Max's 9.0% despite having 13x fewer parameters.
  • Trained in under 9 weeks on 4,096 H200 GPUs; key gains came from RL with longer rollout budgets, multi-harness rollouts, and FP8 RL training.
  • Ships in BF16, FP8, INT4, NVFP4, GGUF, and MLX; runs on a single NVIDIA DGX Spark; available via OpenRouter at $0.10/$0.20 per 1M tokens or free at 256K context.
  • Known limitations: overthinking on hard math, occasional tool schema drift in third-party harnesses, and no intermediate effort control yet.

Laguna S 2.1 is Poolside's biggest open-weight model release to date: a 118 billion parameter Mixture-of-Experts (MoE) model that activates only 8 billion parameters per token, supports a 1 million token context window, and is purpose-built for long-horizon agentic coding. The weights are live on Hugging Face under the permissive OpenMDW-1.1 license, free to download and use commercially.

Laguna S 2.1 went from the start of training to launch in under nine weeks, and on long-horizon coding benchmarks it holds its own against models many times its size. The headline claim is simple: the 118-billion-parameter model matches models several times its size on agentic coding and is small enough to run on a single desktop.

Numbers that matter

Poolside ran Laguna S 2.1 across six agentic coding benchmarks. The results are striking for a model this size:

  • Terminal-Bench 2.1: 70.2% -- ahead of DeepSeek-V4-Pro Max (1.6T params, 64.0%), NVIDIA Nemotron 3 Ultra (550B, 56.4%), and Thinking Machines' Inkling (975B, 63.8%)
  • SWE-Bench Multilingual: 78.5% -- beating Tencent Hy3 (295B, 75.8%) and Nemotron 3 Ultra (67.7%)
  • SWE-Bench Pro: 59.4% -- ahead of DeepSeek-V4-Pro Max (55.4%) and Inkling (54.3%)
  • DeepSWE v1.1: 40.4% -- outperforming DeepSeek-V4-Pro Max which scored just 9.0% on this benchmark

DeepSWE still has significant headroom. Its tasks are longer-horizon and hard to partially solve, and the scores actually spread: frontier models range from 54% to 73% on the v1.1 variant, with some 1T+ parameter open models scoring below 10%. That spread makes it a more meaningful signal than saturated benchmarks where top models cluster within a few points of each other.

Thinking mode matters a lot here. Max thinking lifts S 2.1's score on Terminal-Bench 2.1 from 60.4% to 70.2% and on DeepSWE from 16.5% to 40.4%. The tradeoff is token cost: thinking mode uses roughly 129K tokens per trajectory on Terminal-Bench versus 80K without it.

What "agentic" actually means here

Poolside published three unedited case studies showing the model working autonomously. They're worth reading to understand what the benchmarks are actually measuring.

Keep reading

Don't miss what's next in AI

Join 300,000+ engineers and researchers who get the signal, not the noise. Create a free account to read the rest of this story.

  • Full access to in-depth AI research breakdowns
  • Be the first to know what's trending before it hits mainstream
  • Daily curated papers, repos, and industry moves