Poolside's Tiny Laguna S 2.1 Beats a Giant Rival on Coding Benchmarks by a 4x Margin

The Core · TL;DR
- Poolside's Laguna S 2.1 scores 40.4% on DeepSWE v1.1, far ahead of DeepSeek-V4-Pro-Max's 9.0% despite using about one-sixth the active parameters
- The 118B-parameter MoE model activates only 8B parameters per token and runs on a single NVIDIA DGX Spark, while topping SWE-Bench Multilingual at 78.5%
- Default 'thinking mode' boosts Terminal-Bench 2.1 from 60.4% to 70.2% and DeepSWE from 16.5% to 40.4%, at roughly 2.5x the token cost
- Trained in under nine weeks starting May 22, 2026 on 4,096 H200 GPUs, it's Poolside's third coding model release in three months, now on Hugging Face under an OpenMDW-1.1 license
Poolside just shipped a coding model with 8 billion active parameters that outscores a rival model roughly six times its computational size. Laguna S 2.1, the startup's newest open-weight release, hits 40.4% on the DeepSWE v1.1 benchmark, dwarfing DeepSeek-V4-Pro-Max's 9.0% despite activating a fraction of the parameters per token. It's the kind of efficiency gap that tends to get noticed in a market where inference cost, not just raw capability, decides which models actually get deployed.
The model is a mixture-of-experts system with 118 billion total parameters but only 8 billion active at inference time, small enough to run on a single NVIDIA DGX Spark. On SWE-Bench Multilingual, it posts 78.5%, topping every other model in Poolside's published comparison table. Terminal-Bench 2.1 results land at 70.2% with the model's "thinking mode" switched on, up from a 60.4% baseline without extended reasoning. That reasoning mode ships enabled by default, and it's doing real work: on DeepSWE specifically, turning it on pushes accuracy from 16.5% to 40.4%, though that jump comes at a cost. Trajectories that use thinking mode run roughly 249,000 completion tokens, versus about 99,000 tokens without it, meaning the accuracy gain is bought with roughly 2.5 times the token spend.
Laguna S 2.1 supports context windows up to 1 million tokens in both thinking and non-thinking configurations, and it's Poolside's first model where the reinforcement learning stage ran in FP8 precision, a choice that likely contributed to the compressed training timeline. Pre-training started on May 22, 2026, using 4,096 NVIDIA H200 GPUs, and the entire training-to-launch cycle wrapped in under nine weeks. The model is built as a scale-up of the Laguna XS family, sharing its pre-training data with XS 2.1, and weights are published on Hugging Face under Poolside's OpenMDW-1.1 license.
From Government Contracts to Open Weights
The release marks a notable shift for Poolside, a company that built its early business around government and public-sector clients before pivoting toward open releases, starting with XS.2 under an Apache 2.0 license. Laguna S 2.1 is now the third coding model the company has put out in roughly three months, following Laguna M.1 and XS.2 in April 2026, an unusually fast cadence for a startup competing against labs with far deeper compute budgets.
Poolside's own demonstrations lean into the model's ability to handle open-ended, long-horizon tasks. In one showcase, Laguna S 2.1 built a functioning browser engine from an empty folder in 50 minutes, one capable of rendering HTML and CSS. In another, it produced a proof for Erdős Problem #397, a number theory question that had sat unresolved since 1975, completing the task in 40 reasoning steps for a compute cost of $0.088.
Those anecdotes make for good marketing, but the benchmark numbers are what should matter to engineering teams evaluating self-hosted coding assistants. A model that fits on a single DGX Spark while beating far larger competitors on SWE-Bench Multilingual changes the calculus for teams that want strong code generation without renting a fleet of GPUs.
Original reporting and research used to synthesize this article.
WAKIB Editorial Team
This review was prepared and summarized by the WAKIB AI intelligence engine and vetted by our editorial board for accuracy and reliability.
Subscribe to Newsletter
Get a weekly summary of the most promising AI research and tools delivered to your inbox.
Telegram Channel
Join our active community on Telegram for real-time tracking of AI models and trends.
