Laguna S 2.1 Free
Laguna S 2.1 Free is the free version of Laguna S 2.1 on AI Gateway, with a context window of 256K tokens. It serves the same open-weight agentic coding model as the paid version, at a smaller window. Call it with poolside/laguna-s-2.1-free. Your use is subject to Poolside's Terms & Privacy Policies.
import { streamText } from 'ai'
const result = streamText({ model: 'poolside/laguna-s-2.1-free', prompt: 'Why is the sky blue?'})Playground
Try out Laguna S 2.1 Free by Poolside. Usage is billed to your team at API rates. Free users (those who haven't made a payment) get $5 of credits every 30 days.
Laguna S 2.1 Free
Providers
Route requests across multiple providers. Copy a provider slug to set your preference. Visit the docs for more info. Using a provider means you agree to their terms, listed under Legal.
| Provider |
|---|
P50 throughput on live AI Gateway traffic, in tokens per second (TPS). Visit the docs for more info.
P50 time to first token (TTFT) on live AI Gateway traffic, in milliseconds. View the docs for more info.
Direct request success rate on AI Gateway and per-provider. Visit the docs for more info.
More models by Poolside
| Model |
|---|
About Laguna S 2.1 Free
Laguna S 2.1 Free became available on AI Gateway on July 21, 2026, alongside the paid version. Both entries serve Laguna S 2.1, an open-weight Mixture-of-Experts (MoE) model that Poolside built for agentic coding and released on Hugging Face under the OpenMDW-1.1 license. The listed difference between the two is access: Laguna S 2.1 Free carries a context window of 256K tokens, and the paid version carries a 1M-token window.
That makes Laguna S 2.1 Free the version to reach for first. Point an existing agent at it, run your own tasks, and judge the output before you commit to the paid version. Laguna S 2.1 scores 70.2% on Terminal-Bench 2.1 with thinking enabled, so a short evaluation on real repository work tells you quickly whether the model suits your harness and prompts.
A context window of 256K tokens covers a lot of practical work: focused bug fixes, single-service refactors, test repair, and terminal tasks that finish in a bounded number of steps. It runs short when an agent reads broadly across a large repository or keeps a session going for hours. Watch context usage in AI Gateway reporting and treat the first truncation as the signal to switch.
AI Gateway expanded capacity for both the free and paid versions after launch, so evaluation runs and batch experiments on Laguna S 2.1 Free have more room than they did at release. Capacity describes how much traffic the endpoint accepts, not how well the model performs.
Set the model to poolside/laguna-s-2.1-free in the AI SDK, Chat Completions API, Responses API, Messages API, or other API formats, from TypeScript or Python. To try Laguna S 2.1 Free inside a coding agent, run vercel ai-gateway coding-agents setup, then select poolside/laguna-s-2.1-free in the agent's model configuration. Moving to the paid version later means changing one model string.
What To Consider When Choosing a Provider
- Configuration: Pick a version by context budget, not by expected answer quality. Agent sessions grow quickly once terminal output, file contents, and tool results accumulate, so measure a representative run before you decide. If those runs stay inside 256K tokens, Laguna S 2.1 Free covers them. If they don't, move to the paid version at
laguna-s-2.1, which supports a 1M-token window. - Zero Data Retention: AI Gateway does not currently support Zero Data Retention for this model. See the documentation for models that support ZDR.
- Authentication: AI Gateway authenticates requests using an API key or OIDC token. You do not need to manage provider credentials directly.
When to Use Laguna S 2.1 Free
Best for
- First Evaluation Runs: Testing Laguna S 2.1 on your own repositories before you commit to it
- Coding Agent Prototypes: Early scaffolding where prompts stay well inside 256K tokens
- Side-by-Side Model Comparisons: Benchmarking against other models you already route through AI Gateway
- Bounded Terminal Tasks: Focused build, test, and repair work that finishes in a limited number of steps
- Demos and Workshops: Sessions that need a working agentic coding model on short notice
Consider alternatives when
- Long-Horizon Agent Sessions: Runs that outgrow 256K tokens belong on the paid version at
laguna-s-2.1 - Whole-Repository Prompts: Large codebases sent in a single request exceed the window this version allows
- Sustained Production Traffic: The paid version is the endpoint to standardize on once a workload ships
- Image or Audio Input: Laguna S 2.1 takes text only, so multimodal work needs a model such as Inkling
- Broad General Assistance: Agentic coding is the target, not open-domain chat or factual recall
Conclusion
Laguna S 2.1 Free exists so you can evaluate Laguna S 2.1 on real work before you standardize on it. Run your agent against it, measure how much context a typical session consumes, and move to laguna-s-2.1 when sessions outgrow 256K tokens. Open https://ai-sdk.dev/playground/poolside:poolside/laguna-s-2.1-vercel to try it interactively.