Skip to main content

1. Model Introduction

Laguna-XS.2 is an open-source hybrid sliding-window-attention MoE model from Poolside, built for agentic coding and long-horizon software engineering work. Key Features:
  • MoE: 33.4B total parameters, 3.0B active per token, 256 routed experts (top-8) plus 1 shared.
  • Long context: 131,072 tokens.
  • Agentic coding: Tuned for tool-using software engineering agents and long-horizon execution.
  • Hybrid reasoning: <think>...</think> segments toggled per request via chat_template_kwargs={"enable_thinking": ...}.
Available Quantizations:
VariantHugging Face path
BF16poolside/Laguna-XS.2
FP8poolside/Laguna-XS.2-FP8
NVFP4poolside/Laguna-XS.2-NVFP4
License: Apache 2.0 For details, see the Hugging Face model card and the Laguna deeper-dive blog post.

2. SGLang Installation

Laguna-XS.2 support is on main but not yet in a tagged release; install from the SGLang nightly wheel index, or pull a pre-built Docker image:
Command
For the full Docker setup and other installation methods, please refer to the official SGLang installation guide.

3. Model Deployment

3.1 Basic Configuration

Interactive Command Generator: Use the configuration selector below to generate a launch command for your hardware.

3.2 Configuration Tips

  • Trust remote code (--trust-remote-code): Laguna-XS.2 ships custom modeling/config code on the Hugging Face Hub, so this flag is required for the server to load the model.
  • Quantization: NVFP4 requires Blackwell (B200 / B300); BF16 and FP8 run on either H200 or B200. FP8’s first launch triggers a multi-session DeepGEMM JIT pre-compile (~10-20 min); pre-warm with python3 -m sglang.compile_deep_gemm --model poolside/Laguna-XS.2-FP8 to avoid that cost on every restart.
  • Reasoning parser (--reasoning-parser poolside_v1): Splits <think>...</think> segments into reasoning_content so content holds only the final answer. Disable only if you want the raw <think> tags in content.
  • Tool call parser (--tool-call-parser poolside_v1): Required for OpenAI-compatible tool-call streaming. Disable only for chat-only deployments.
  • DP attention: For higher-throughput deployments, enable the DP-Attention toggle — it emits --dp <N> --enable-dp-attention with --dp matching --tp (tune independently if needed).
  • Thinking default: Thinking is off by default at the model level. Opt in per request with extra_body={"chat_template_kwargs": {"enable_thinking": True}}.

4. Model Invocation

The samples below assume the server is reachable at http://localhost:30000/v1.

4.1 Basic Chat

Example
Output Example:
Output

4.2 Reasoning (Thinking Mode)

Laguna-XS.2 emits reasoning between <think>...</think> tags. The --reasoning-parser poolside_v1 flag separates the thinking text into reasoning_content so content holds only the final answer. Thinking is opt-in per request:
Example
Output Example:
Output
To disable thinking, omit extra_body (off by default) or pass chat_template_kwargs={"enable_thinking": False} explicitly.

4.3 Tool Calling

Example
Output Example:
Output
reasoning_content is None because thinking is off by default; content carries the brief assistant message that precedes the tool call. Add extra_body={"chat_template_kwargs": {"enable_thinking": True}} if you want interleaved reasoning before the tool call.

5. Benchmark

5.1 Accuracy Benchmark

Test Environment:
  • Hardware: NVIDIA H200 (4×H200)
  • Model: poolside/Laguna-XS.2 (BF16)
  • Tensor Parallelism: 4
  • SGLang Version: 0.5.12.dev20260509+g096ad02b0 (nightly wheel containing the #24204 merge commit; same code path as the original PR runs)
  • Reasoning Parser: poolside_v1
  • Tool Call Parser: poolside_v1
  • Sampling: temperature=0.6, max_tokens=16384, chat_template_kwargs={"enable_thinking": true}, n_repeats=1
  • Grader: NeMo-Skills math_verify (math) and eval_mcq (multichoice)
Results (from PR #24204):

5.2 Speed Benchmark

Test Environment:
  • Hardware: NVIDIA H200 (1×H200 for TP=1, 4×H200 for TP=4)
  • Model: poolside/Laguna-XS.2 (BF16)
  • SGLang Version: 0.5.12.dev20260509+g096ad02b0 (nightly wheel containing the #24204 merge commit; same code path as the original PR runs)
  • Workload: sglang.bench_serving --backend sglang --dataset-name random (defaults: --random-input-len 1024 --random-output-len 1024 --random-range-ratio 0.0)
  • Server flags identical to the accuracy runs above.

5.2.1 Latency Benchmark (10 prompts, concurrency = 1)

Command

5.2.2 Throughput Benchmark (1000 prompts, concurrency = 100)

Command
TP=4 delivers roughly 2.0× total-token throughput and ~1.7× lower mean TTFT compared to TP=1 on the cc=100 random workload.