Skip to main content

Deployment

For all methods and hardware platforms, see the official SGLang installation guide. The two paths below match the Python / Docker toggle in the command panel.
Command
Then run the Python output of the command panel below in that environment.
Pick your card + checkpoint precision to generate the launch command. The model runs single-GPU on every supported card — H200, RTX PRO 6000, RTX 5090 and DGX Spark — and ships one operating point.
--mamba-full-memory-ratio is the one sizing flag that matters for throughput on hybrid GDN models: the default (0.9) over-provisions the KV pool and silently clamps concurrency. Set your average request length in the Mamba ratio calculator below; everything else follows the panels, and the computed value is pinned into the command.
Every cell above — RTX 5090, RTX PRO 6000 and DGX Spark, across all five checkpoints and every Speculative Decoding / Serving Strategy / SSM dtype combination — is measured on v0.5.19. That is 202 cells, each one served and scored on the full 1319-question GSM8K (93.18-95.15%). The serving envelope behind the pins is ISL 8192 / OSL 1024 at concurrency 1; throughput and acceptance-length numbers were not re-taken in that sweep.

Mamba ratio calculator

Hybrid GDN models split post-weight memory into a worst-case-reserved GDN state pool (sets the concurrency ceiling) and a paged attention KV pool, divided by --mamba-full-memory-ratio. Every parameter below except L and the target concurrency is read live from the Deploy panel and Playground selection; the balanced value is the per-request cost ratio:
Formula
  • S — state slots per running request: extra_buffer=5 (default), extra_buffer_lazy=4, no_buffer=3, disabled radix cache =1. For the two extra_buffer strategies, SGLANG_OPT_MAMBA_SKIP_DECODE_LOCK=1 frees one slot, and extra_buffer frees one more with the overlap scheduler off; the calculator reads both knobs.
  • D — verify intermediate states under speculative decoding: --speculative-num-draft-tokens for EAGLE/MTP (4 at the recommended 3/1/4) and for DFLASH (8, DFlash2’s block size); --speculative-dspark-block-size + 1 for DSPARK, where the block size falls back to the draft checkpoint’s block_size when the flag is omitted (7 for RadixArk/Qwen3.8-27B-DSpark, so D = 8); 0 with speculation off or with --enable-linear-replayssm-spec, which keeps the verify intermediates on a fixed ring instead of per-request slots.
  • state_bytes — one state slot, from the fixed geometry (48 GDN layers x 48 heads x 128 x 128 at --mamba-ssm-dtype, plus bf16 conv state): 153.9 MB at fp32, 78.4 MB at bf16.
  • kv_bytes_per_token — 16 attention layers x GQA 4 x 256 x K+V: 32.8 KB at fp8, 65.5 KB at bf16.
  • L — average total request length in tokens: input + output.
--max-mamba-cache-size = target_concurrency x S is the equivalent explicit pin and overrides the ratio; the calculator emits it alongside. D is not a term here: the engine divides the state pool by S alone and sizes the speculative verify buffer separately, so folding D into the pin would over-provision the pool. After boot, verify with the max_running_requests line in the server log — it should not be capped below your target concurrency.

Playground

The Playground is where you experiment with SGLang features beyond the recipes above. The Deploy panel emits this model’s documented launch recipes; the Playground lets you turn on additional knobs on top of whichever cell the Deploy panel is currently showing.

1. Model Introduction

Qwen3.8-27B is a dense hybrid Gated Delta Networks (GDN) vision-language model: a 27B causal language model paired with a vision encoder, with native image and video understanding alongside text. SGLang serves it through the Qwen3-VL path, so the vision tower is live on the recipes below. The language model is 64 layers, laid out as 16 repeats of 3 × (Gated DeltaNet → FFN) followed by 1 × (Gated Attention → FFN) — 48 linear-attention layers to 16 full-attention ones. Gated DeltaNet runs 48 value heads and 16 QK heads at head_dim 128; Gated Attention is GQA 24/4 at head_dim 256 with a 64-dim rotary slice. Hidden size is 5120 over a 17,408-dim FFN, and the checkpoint ships an MTP head trained with multiple steps. Context is 262,144 tokens natively, extensible to 1,000,000. The serving-relevant architecture is identical to Qwen3.6-27B. Thinking mode is on by default and can be disabled per request; reasoning depth is tunable with reasoning_effort, and preserve_thinking retains reasoning context from earlier messages.
ModelQuantizationWeights
Qwen3.8-27BBF16Qwen/Qwen3.8-27B
Qwen3.8-27B-FP8FP8 (blockwise)Qwen/Qwen3.8-27B-FP8
Qwen3.8-27B-NVFP4 (FP4 head)NVFP4 W4A4 + FP8 projections, lm_head packed to FP4RadixArk/Qwen3.8-27B-NVFP4
Qwen3.8-27B-NVFP4 (BF16 head)Same body, lm_head left dense in BF16RadixArk/Qwen3.8-27B-NVFP4-BF16-LMHead
Qwen3.8-27B-NVFP4 (NVIDIA)NVIDIA’s ModelOpt export of the same W4A4 body, lm_head packed to FP4nvidia/Qwen3.8-27B-NVFP4
The two RadixArk NVFP4 exports differ only in the lm_head: one packs it to FP4, the other leaves it dense in BF16. The dense head is ~1.7 GB larger on disk and ~3.2 GB larger at runtime, so it is the harder of the two to fit — every recipe on this page was measured against it, and the FP4-head cells reuse those pins unchanged. NVIDIA’s own export is that same W4A4 body with that same FP4 head: identical quantized-layer map (FP8 attention and GDN projections, NVFP4 MLPs), identical tensor set, identical 21.9 GB on disk. On GB300, RTX PRO 6000 and DGX Spark its cells reuse the FP4-head pins unchanged, and both SM12x grids have been re-measured against this export on v0.5.19: all 16 overlay combinations per card serve and score 94.01-95.00% (RTX PRO 6000) and 94.16-95.07% (DGX Spark) on the full 1319-question GSM8K. The RTX 5090 is measured too — all 15 overlay combinations it offers serve and score 93.93-94.92% — and every winning launch command there is identical to the FP4-head export’s, which is the strongest form of the claim above. What the 32GB card does need is the draft-model rows pinning their own pools: those recipes pin --max-running-requests 1, but nothing caps the pools to match, so the KV pool sizes itself for 127,332 tokens against the 9,216 one 8192-in/1024-out request needs, and the engine’s default split then leaves the GDN state pool far short of the slots it needs once the draft model’s weights are counted against --mem-fraction-static. The DSPARK row therefore pins --max-total-tokens and a measured --mamba-full-memory-ratio, as do DFLASH2 and MTP on the dense-lm_head export at float32 state. Those pins override the calculator’s live value for the selections that carry them. The no-speculation row needs none of it and runs at the pins shown. The two RadixArk checkpoints declare kv_cache_quant_algo: FP8, so SGLang’s default --kv-cache-dtype auto already puts their KV pool in fp8_e4m3. The NVIDIA export ships no kv_cache_scheme, so auto would leave its pool in BF16 instead. Every recipe on this page pins --kv-cache-dtype fp8_e4m3 explicitly, so all three run the same fp8_e4m3 pool regardless; the difference only shows up if you switch the Playground’s KV Cache Precision row back to Auto.

2. Configuration Tips

  • SM120/SM121 (RTX PRO 6000 Blackwell, RTX 5090, DGX Spark): use --attention-backend flashinfer; trtllm_mha is SM100-only. MTP with the FlashInfer backend requires a FlashInfer build whose prefill plan accepts uniform_q_len (newer than 0.6.15.post1); otherwise run spec with --attention-backend triton. On DGX Spark the 128GB is unified memory shared with the host CPU, so all three checkpoints fit, and its cells reuse the RTX PRO 6000 recipe at --mem-fraction-static 0.80 rather than a separate operating point. The one lower pin is the unified pool pricing the host’s memory too: 0.85 of 128GB leaves ~8GB for the OS — exactly DGX OS earlyoom’s SIGTERM threshold — and the first long prefill or boot-time graph capture dips under it and gets the scheduler killed with exit code -15 and no traceback (journalctl -u earlyoom shows the kill). At 0.85, 15 of the 48 cells were killed that way, and which cells is margin noise; at 0.80 every cell served on every attempt. Validated on SM121 / aarch64: all 80 configurations (5 checkpoints x Speculative Decoding x Serving Strategy x Mamba SSM Dtype, DFLASH2 included) served on GB10 on v0.5.19 at ISL 8192 / OSL 1024, concurrency 1, and each scored the full 1319-question GSM8K (93.18-95.15%); the float32 and bfloat16 halves ran on two separate GB10 boxes. No throughput or acceptance-length numbers were re-taken. The sweep exercises the FlashInfer plan / uniform_q_len path above, which raised no arity error on that build. Three host quirks when reproducing on GB10: docker GPU access is CDI-only (--device nvidia.com/gpu=all, as no nvidia runtime is registered); nvidia-smi reports Not Supported for memory because it is unified with the CPU — gate a relaunch on MemAvailable in /proc/meminfo instead; and the BF16 checkpoint takes ~6.5 minutes just to load its 18 shards from NVMe, so budget ~10 minutes to READY before calling a boot hung.
  • H200 (SM90): BF16 and FP8 only — the card has no FP4 tensor cores, so an NVFP4 checkpoint’s MLP would fall back to the Marlin W4A16 weight-only path, and all three NVFP4 cells are greyed out. The H200 recipes use 32768-token prefill chunks (SM90 prefill is fast enough that a big chunk barely stalls decode, unlike the SM120 guidance below), and the FlashInfer GDN prefill backend engages by default under them. --attention-backend fa3 is a valid alternative, measured slightly faster at bs=1.
  • MTP: --speculative-algorithm EAGLE --speculative-num-steps 3 --speculative-eagle-topk 1 --speculative-num-draft-tokens 4 uses the in-checkpoint MTP head. (This recipe was originally documented with NEXTN, an alias of EAGLE — same algorithm.)
  • DSpark: the trained draft model is a separate checkpoint — add --speculative-algorithm DSPARK --speculative-draft-model-path RadixArk/Qwen3.8-27B-DSpark (the Playground’s Speculative Decoding card emits this pair). DSpark does not take --speculative-num-draft-tokens: its verify window is --speculative-dspark-block-size (gamma) + 1, and gamma is auto-inferred from the draft checkpoint when the flag is omitted (7 for this checkpoint, so D = 8). That D is a term in the balanced ratio — r = (S + D) x token_equiv / L, where token_equiv is the state slot expressed in KV tokens, state_bytes / kv_bytes_per_token (4698 at fp32 state / 2394 at bf16, over fp8 KV) — so DSpark needs a materially higher --mamba-full-memory-ratio than no-spec at the same S, and pinning a different gamma changes the ratio with it. MTP is the opposite case: with --enable-linear-replayssm-spec its draft intermediates move onto a fixed ring, so D = 0 and the ratio returns to the no-spec value. The calculator applies both rules.
  • DFlash2: a trained block-diffusion draft in a separate checkpoint — add --speculative-algorithm DFLASH --speculative-draft-model-path incoai/Qwen3.8-27B-DFlash2 --speculative-num-draft-tokens 8 (8 is the draft’s block size, and it is the D term in the ratio, same value as DSpark’s). Runs on Ascend NPUs as well (#35629): the selector verify falls back to argmax there, matching what EAGLE and first-generation DFlash drafts (such as z-lab/Qwen3-8B-DFlash-b16) already do on NPU. Therefore, NPU currently guarantees lossless verification only for greedy requests; use temperature=0 and top_k=1. Non-greedy requests log a warning; both draft proposal and target verification fall back to greedy, so their requested sampling distribution is not preserved. The selector projects candidates through the target lm_head, including quantized heads, so it runs on the NVFP4 checkpoint (whose head is NVFP4-packed; the BF16 and FP8 checkpoints keep a dense head). The Ascend comparison in #35629 used an A3 Series device with BF16 target weights, --tp-size 2 --attention-backend ascend --mamba-ssm-dtype bfloat16 --mamba-radix-cache-strategy extra_buffer, and disabled RadixCache for both baseline and DFlash2 to exclude cache warm-up and prefix reuse. The DFlash2 run added the three flags shown above. That comparison’s accuracy used zero-shot GSM8K with greedy sampling, max_new_tokens=2048, 128 examples, and concurrency levels 1, 2, 4, 8 and 16 — a different protocol from this page’s own sweep below. Validation: every SM12x cell on this page is measured end to end on v0.5.19 — 202 cells over the five checkpoints, four speculative options, two serving tiers and two GDN state dtypes, full 1319-question GSM8K on each, 93.18-95.15%. The RTX PRO 6000 and DGX Spark recipes need no changes. On the 32GB RTX 5090 the panel applies the measured pins automatically: DFlash2 at --mem-fraction-static 0.91 with --chunked-prefill-size 1024 — at 0.91 the pools fit but a 2048-token chunk’s activations do not — DSpark at 0.88 (bfloat16), 0.91 (float32) and 0.92 on the dense-lm_head export, all three with their pools pinned and the last two also cutting the prefill chunk to 1024 and 512, EAGLE at 0.93 (bfloat16) and 0.94 (float32), and no-speculation at 0.90. Whether float32 is available with a draft model depends on the lm_head: on the BF16-head export it is greyed out for both DSpark and DFlash2, since the dense head’s ~3.2 GB leave no fp32 state pool that also clears prefill graph capture. The FP4-head export frees that headroom back — DSpark serves at 0.91 and DFlash2 High-Throughput at 0.895 with --mamba-full-memory-ratio 10 overriding the balanced value — and only DFlash2 Low-Latency stays out of reach, where five fp32 slots and a full request’s KV never coexist. bfloat16 remains the faster choice regardless: DFlash2 posts 4.92 ms median TPOT at an accept length of 4.29, the best result on this card.
  • Hardware fit: FP8 weights ~28.5GB (not serviceable beyond bs≤2 on 32GB cards); NVFP4 weights ~16.5GB (recommended for RTX 5090-class GPUs).
  • --mamba-radix-cache-strategy extra_buffer_lazy lowers the state cost per request from 5 slots to 4 at no accuracy cost. On small-VRAM cards (RTX 5090 32GB) the state pool bounds concurrency long before KV does — prefer lowering S (lazy strategy, or --disable-radix-cache for S=1); the calculator re-derives the ratio for the new S. The balanced ratio itself is VRAM-independent.
  • --mamba-ssm-dtype: the GDN state slot is 153.9 MB at float32 (the checkpoint’s declared precision) and 78.4 MB at bfloat16, so bf16 roughly halves the state pool and hands the difference to KV — measured on an RTX 5090 with no speculation, 97,280 KV tokens at bf16 against 68,588 at fp32. On 32GB cards it also decides whether a config fits at all: EAGLE needs --mem-fraction-static 0.94 at fp32 but 0.92 at bf16. Speed is not a one-way trade — with speculative decoding fp32 sometimes wins (NVFP4 + EAGLE: 152.9 vs 144.5 tok/s/user) and sometimes loses (FP8 + EAGLE: 106.3 vs 116.1); measure both for your quantization. Treat bfloat16 as an accuracy gate and validate it for your workload. On SM120 both precisions run the Triton linear-attn prefill path — the FlashInfer GDN prefill fast path gates on SM100, where its validated domain is in fact a bf16 state pool — so no dtype forces an extra flag here. One interaction to know: --enable-linear-replayssm-spec auto-selects fp32 state when --mamba-ssm-dtype is unset, and an explicit non-fp32 value logs a state-drift warning at boot. The SSM dtype row always emits the flag explicitly, so the bf16 + EAGLE cells run with that warning — accounted for in their validation.
  • --chunked-prefill-size 2048: decode steps stall behind each prefill chunk on hybrid GDN models, and 8192-token chunks stall them ~600ms at a time. 2048 keeps decode inter-token latency smooth under mixed load and also improves single-wave TTFT.

3. Agent Harnesses

Agent harnesses drive the model through the OpenAI-compatible endpoint — or, for Claude Code, through SGLang’s Anthropic-compatible one — so any of them works once three things line up. The parsers ship in the command. Every recipe above carries --reasoning-parser qwen3 --tool-call-parser qwen3_coder, because without them a harness receives tool calls as raw text instead of structured tool_calls. The Parsers card in the Playground is therefore an opt-out — both chips start on, and turning one off strips its flag. qwen3_coder is the right tool-call parser for this checkpoint: its chat template instructs the model to reply with an inner <function=…> / <parameter=…> block nested in <tool_call></tool_call>, which is exactly what that parser decodes. The Hermes parser (--tool-call-parser hermes) reads a different payload — bare JSON inside <tool_call> — so pointing a Hermes-format harness at this model without switching the flag yields tool calls that never parse. --reasoning-parser qwen3 matches the template’s enable_thinking toggle, which defaults to on. Endpoint and model id. The base URL is http://<host>:30000/v1. The model string a harness sends must equal the server’s --model-path — the OpenAI /v1/models name defaults to it — unless you override it with --served-model-name, which is usually worth doing to keep harness configs short. SGLang also serves an Anthropic-compatible /v1/messages, which is what §3.3 uses. It converts each request to the OpenAI shape, hands it to the same chat-serving path, and converts the response back — so the parser flags above apply there identically. Auth. --api-key is unset by default, so the server accepts unauthenticated requests. Harnesses that insist on a key can send any placeholder; set --api-key on the server if the endpoint is reachable beyond localhost.

3.1 OpenCode

OpenCode reaches a self-hosted endpoint through a provider entry in opencode.json.
Store the credential first — pick Other, give the provider an id, and enter any placeholder when the server has no --api-key:
Command
Then declare the provider in opencode.json:
Config
npm selects the transport — @ai-sdk/openai-compatible is the one for a plain OpenAI-shaped endpoint. apiKey is optional and takes a "{env:VAR_NAME}" reference rather than a literal. The models keys are the ids sent on the wire, so they must match the served model name. Confirm with /models.

3.2 Pi

Pi (@earendil-works/pi-coding-agent) registers providers from an extension rather than a config file.
Extension
api: "openai-completions" is what selects the OpenAI-compatible transport, and apiKey takes a $ENV_VAR reference rather than a literal. contextWindow is the checkpoint’s native 262,144; set maxTokens to whatever output cap you want per turn. Confirm registration with pi --list-models.

3.3 Claude Code

Claude Code speaks the Anthropic API, so it points at SGLang’s /v1/messages rather than the OpenAI endpoint.
Anthropic documents that routing Claude Code to non-Claude models through a gateway is not supported. The wiring below works because SGLang implements the Anthropic message format, but it sits outside what Claude Code is tested against — expect newer Claude Code features to degrade or fail.
ANTHROPIC_BASE_URL is the server origin — Claude Code appends /v1/messages itself, so leave the /v1 suffix off:
Command
The two credential variables travel in different headers: ANTHROPIC_AUTH_TOKEN goes out as Authorization: Bearer, ANTHROPIC_API_KEY as x-api-key. Either satisfies a server started without --api-key; with --api-key set, pick the variable matching the header your server reads. A credential variable also takes precedence over a saved claude.ai login for that session.The same pair can live in a settings file instead, which persists across shells and wins over a shell export:
Config
Run /status in Claude Code to confirm which base URL and credential source the session picked up.

3.4 Hermes Agent

Hermes Agent (Nous Research, MIT) selects a self-hosted endpoint through its setup wizard or its config file.
Command
Equivalently, in ~/.hermes/config.yaml:
Config
For several endpoints at once, declare them under providers: and switch with /model custom:<name> mid-session:
Config