organisation · org/poolside
Poolside
A row on a public router is the thing Poolside's public record argued against. Read its blog backwards from May 2026 and it is an enterprise sales record: the AWS announcement of December 2024 — "Unveiling Poolside's first-party partnership with AWS"source, accessed 2026-09-06 — a Redpanda tie-up, the acquisition of London's Fern Labs, a Dell configuration, and in May the Poolside Platform, sold on the premise that "For any team whose data is too sensitive, too regulated, or too strategic to send outside of their security boundary, the Poolside Platform puts AI within reach."source, accessed 2026-09-06. Seven days before that Platform post, on 28 April 2026, the same blog said "Today is the first time we're shipping models in public."source, accessed 2026-09-06. Every row this catalog carries is newer than that sentence, and the homepage now opens with "We're building open-weight foundation models and the systems that refine and improve them."source, accessed 2026-09-06.
The interesting part is what happened to the licence in the twelve weeks after. The first two public models were the most permissive thing on offer: Laguna XS.2's repository went up on 23 April 2026 as apache-2.0source, accessed 2026-09-06, Laguna M.1's on 15 June as apache-2.0source, accessed 2026-09-06. The next two did not. Laguna XS 2.1's repository (20 June) is openmdw-1.1source, accessed 2026-09-06 and Laguna S 2.1's repository (13 July) is openmdw-1.1source, accessed 2026-09-06 — the Linux Foundation's model-specific licence rather than the one every developer already knows. Poolside filed the switch under a heading reading "A more open license": "We are making this change to support open model distribution for the community. OpenMDW-1.1 is fully permissive and designed for models and related artifacts, giving developers and organizations a more consistent framework for using, modifying and deploying open models." — the paragraph under the heading "A more open license"source, accessed 2026-09-06.
Whether that is more open is arguable; that it is a change is not, and the cleanest place to see it is the pair where nothing else moved. Poolside says XS 2.1 is a refresh, not a redesign — "It's the same architecture as XS.2, with a notable improvement on SWE-bench Multilingual and stronger performance on terminal-style tasks."source, accessed 2026-09-06 — and the weights agree to the parameter: the API reports 33,442,617,088, the safetensors parameter total the Hugging Face API reports for poolside/Laguna-XS.2source, accessed 2026-09-06 and 33,442,617,088, the safetensors parameter total the Hugging Face API reports for poolside/Laguna-XS-2.1source, accessed 2026-09-06. Same architecture, same size, new terms. Both of this catalog's Laguna models sit on the OpenMDW side of that line; neither Apache release ever got a row here.
The headline claim is bounded more tightly than it first reads. Poolside calls Laguna S 2.1 "Laguna S 2.1 is, as far as we can measure, the most capable agentic coding model in its weight class by a wide margin", the vendor's own bolded comparisonsource, accessed 2026-09-06, and "in its weight class" is carrying the sentence. Not one of the five rivals that publishes a size is in that class, and three publish none: Tencent Hy3 295B-A21B, Nemotron 3 Ultra 550B-A55B, Inkling 975B-A41B, DeepSeek-V4-Pro Max 1.6T-A49B, Kimi K3 2.8T-A50B — and Qwen 3.7 Max, Muse Spark 1.1 and Claude Fable 5 carrying an em dash where a size would gosource, accessed 2026-09-06 — against "Laguna S 2.1 is a 118B total parameter Mixture-of-Experts (MoE) model with 8B activated parameters per token"source, accessed 2026-09-06. On Terminal-Bench 2.1, the first benchmark in that table, most of it beats Poolside — 70.2 for Laguna S 2.1 (118B-A8B), behind Kimi K3 88.3, Claude Fable 5 88.0, Muse Spark 1.1 80, Qwen 3.7 Max 74.5 and Tencent Hy3 71.7 — five of the eight columns Poolside chosesource, accessed 2026-09-06. On SWE-Bench Multilingual it leads every cell that is filled: 78.5 for Laguna S 2.1, ahead of Qwen 3.7 Max 78.3, DeepSeek-V4-Pro Max 76.2, Tencent Hy3 75.8 and Nemotron 3 Ultra 67.7 — the other four columns are emptysource, accessed 2026-09-06. And the methodology note is generous to everyone else in the room: "For all benchmarks we take the maximum of the vendor self-reported score, benchmark author leaderboard or third-party leaderboard (Artificial Analysis), except SWE Atlas (Codebase QnA) where we do not use third-party leaderboard figures."source, accessed 2026-09-06. Each competitor is scored at its best published number, and Poolside printed the rows it loses anyway.
Then it published the evidence. "For every benchmark score we publish today, we are releasing full trajectories for every trial in the final evaluation set at trajectories.poolside.ai."source, accessed 2026-09-06 — and the archive is live. The Terminal-Bench view alone holds 712 trial records in the Terminal-Bench 2.1 view's embedded payload, spanning 89 distinct task ids in `thinking` and `no-thinking` variants; each record carries its reward, step count, reasoning-character count and dollar costsource, accessed 2026-09-06. A reader who doubts the score can open the runs behind it, including the ones that failed, and see what each attempt cost. Publishing a number invites trust; publishing the transcripts invites contradiction.
That model came out of a short run by the standards of the sizes it is measured against — "Laguna S 2.1 began pre-training on 4,096 NVIDIA H200 GPUs on May 22, 2026, 60 days ago."source, accessed 2026-09-06 — with NVIDIA credited for the inference work behind it: "NVIDIA helped optimize inference across its hardware, from TRT-LLM serving and NVFP4 on Blackwell systems down to a single NVIDIA DGX Spark."source, accessed 2026-09-06. The reception was lopsided in its favour: 1,016 likes and 46,646 downloads on poolside/Laguna-S-2.1, against 240 and 42,728 for Laguna-XS-2.1, 320 and 30,523 for Laguna-XS.2, and 144 and 8,201 for Laguna-M.1source, accessed 2026-09-06.
The catalog rows themselves conceal an asymmetry worth knowing before you
spend a request on the free tier. poolside/laguna-s-2.1 serves
1048576tokensopenrouter-models, last checked 2026-09-14 and
poolside/laguna-s-2.1:free serves
262144tokensopenrouter-models, last checked 2026-09-14; the XS pair are
identical on both tiers. The long context is the paid product. And Poolside is
not treating the router as a shelf of last resort:
poolside/laguna-xs-2.1 created 2026-07-02T14:27:09Z, poolside/laguna-s-2.1 created 2026-07-21T16:51:23Zsource, accessed 2026-09-06 — each timestamp lands on the same UTC day
as the post that announced the model.
Facts
- self description
- "We're building open-weight foundation models and the systems that refine and improve them."source, accessed 2026-09-06
- first public models
- "Today is the first time we're shipping models in public."source, accessed 2026-09-06
- open ecosystem aim
- "The open-weight ecosystem in the West is still early in its development. We want to change that."source, accessed 2026-09-06
- platform pitch
- "For any team whose data is too sensitive, too regulated, or too strategic to send outside of their security boundary, the Poolside Platform puts AI within reach."source, accessed 2026-09-06
- aws partnership
- "Unveiling Poolside's first-party partnership with AWS"source, accessed 2026-09-06
- laguna xs 2 license
- apache-2.0source, accessed 2026-09-06
- laguna m 1 license
- apache-2.0source, accessed 2026-09-06
- laguna xs 2 1 license
- openmdw-1.1source, accessed 2026-09-06
- laguna s 2 1 license
- openmdw-1.1source, accessed 2026-09-06
- license change rationale
- "We are making this change to support open model distribution for the community. OpenMDW-1.1 is fully permissive and designed for models and related artifacts, giving developers and organizations a more consistent framework for using, modifying and deploying open models." — the paragraph under the heading "A more open license"source, accessed 2026-09-06
- xs architecture unchanged
- "It's the same architecture as XS.2, with a notable improvement on SWE-bench Multilingual and stronger performance on terminal-style tasks."source, accessed 2026-09-06
- xs 2 tensor total
- 33,442,617,088, the safetensors parameter total the Hugging Face API reports for poolside/Laguna-XS.2source, accessed 2026-09-06
- xs 2 1 tensor total
- 33,442,617,088, the safetensors parameter total the Hugging Face API reports for poolside/Laguna-XS-2.1source, accessed 2026-09-06
- weight class claim
- "Laguna S 2.1 is, as far as we can measure, the most capable agentic coding model in its weight class by a wide margin", the vendor's own bolded comparisonsource, accessed 2026-09-06
- laguna s 2 1 size
- "Laguna S 2.1 is a 118B total parameter Mixture-of-Experts (MoE) model with 8B activated parameters per token"source, accessed 2026-09-06
- comparison table sizes
- Tencent Hy3 295B-A21B, Nemotron 3 Ultra 550B-A55B, Inkling 975B-A41B, DeepSeek-V4-Pro Max 1.6T-A49B, Kimi K3 2.8T-A50B — and Qwen 3.7 Max, Muse Spark 1.1 and Claude Fable 5 carrying an em dash where a size would gosource, accessed 2026-09-06
- terminal bench scores
- 70.2 for Laguna S 2.1 (118B-A8B), behind Kimi K3 88.3, Claude Fable 5 88.0, Muse Spark 1.1 80, Qwen 3.7 Max 74.5 and Tencent Hy3 71.7 — five of the eight columns Poolside chosesource, accessed 2026-09-06
- swe bench multilingual scores
- 78.5 for Laguna S 2.1, ahead of Qwen 3.7 Max 78.3, DeepSeek-V4-Pro Max 76.2, Tencent Hy3 75.8 and Nemotron 3 Ultra 67.7 — the other four columns are emptysource, accessed 2026-09-06
- benchmark method
- "For all benchmarks we take the maximum of the vendor self-reported score, benchmark author leaderboard or third-party leaderboard (Artificial Analysis), except SWE Atlas (Codebase QnA) where we do not use third-party leaderboard figures."source, accessed 2026-09-06
- trajectory release
- "For every benchmark score we publish today, we are releasing full trajectories for every trial in the final evaluation set at trajectories.poolside.ai."source, accessed 2026-09-06
- trajectory archive
- 712 trial records in the Terminal-Bench 2.1 view's embedded payload, spanning 89 distinct task ids in `thinking` and `no-thinking` variants; each record carries its reward, step count, reasoning-character count and dollar costsource, accessed 2026-09-06
- nvidia inference work
- "NVIDIA helped optimize inference across its hardware, from TRT-LLM serving and NVFP4 on Blackwell systems down to a single NVIDIA DGX Spark."source, accessed 2026-09-06
- training run
- "Laguna S 2.1 began pre-training on 4,096 NVIDIA H200 GPUs on May 22, 2026, 60 days ago."source, accessed 2026-09-06
- router created
- poolside/laguna-xs-2.1 created 2026-07-02T14:27:09Z, poolside/laguna-s-2.1 created 2026-07-21T16:51:23Zsource, accessed 2026-09-06
- laguna s 2 1 reception
- 1,016 likes and 46,646 downloads on poolside/Laguna-S-2.1, against 240 and 42,728 for Laguna-XS-2.1, 320 and 30,523 for Laguna-XS.2, and 144 and 8,201 for Laguna-M.1source, accessed 2026-09-06
Timeline
- Laguna S 2.1 announced with full evaluation trajectories published alongside the scoressource
- poolside/Laguna-S-2.1 repository created on Hugging Face (createdAt), licensed openmdw-1.1source
- Laguna XS 2.1 announced, the licence change explained under the heading "A more open license"source
- poolside/Laguna-XS-2.1 repository created on Hugging Face (createdAt), the first Laguna licensed openmdw-1.1source
- poolside/Laguna-M.1 repository created on Hugging Face (createdAt), still apache-2.0source
- The Poolside Platform announced, for teams that will not send data outside their own security boundarysource
- Laguna XS.2 and Laguna M.1 announced — Poolside's first models shipped in publicsource
- poolside/Laguna-XS.2 repository created on Hugging Face (createdAt), licensed apache-2.0source
- Poolside acquires Fern Labs, a London company behind the Bridge multi-agent orchestration layersource
- Poolside announces a first-party partnership with AWSsource