> ## Documentation Index
> Fetch the complete documentation index at: https://lmsysorg-dsv4-1.mintlify.site/llms.txt
> Use this file to discover all available pages before exploring further.

# Performance Optimization

> Choose performance levers for SGLang Diffusion by latency, throughput, memory, and quality tradeoffs.

Use this page as the starting point for SGLang Diffusion performance work. It separates performance levers into two decision classes:

* **Output-preserving / lossless-style:** system settings that should preserve model behavior while changing residency, parallelism, kernels, or scheduling.
* **Quality-tradeoff / lossy or approximate:** techniques that can change the denoising path, numerical representation, or generated output.

The docs use "output-preserving" instead of promising bit-exact "lossless" because different kernels, GPU types, or precision paths can still introduce small numerical differences. The decision boundary is whether the optimization intentionally trades quality or output equivalence for speed.

## Start Here

1. Pick a serving or generation mode from [Deployment and Performance Modes](/docs/sglang-diffusion/deployment_cookbook). `--performance-mode auto` is the default; use `speed` when the model fits in GPU memory and latency matters most, `memory` when GPU memory is the bottleneck, and `manual` when every performance flag should be explicit.
2. Choose the right attention backend from [Attention Backends](/docs/sglang-diffusion/attention_backends).
3. Use [Sequence Parallelism](/docs/sglang-diffusion/ring_sp_performance) only when the model and video shape benefit from sequence splitting.
4. Use [Inference Batching](/docs/sglang-diffusion/dynamic_batching) for concurrent compatible requests during serving.
5. Use [Profiling](/docs/sglang-diffusion/profiling) before changing several levers at once.

## Choose a request quality tier

`--quality` is cumulative: a broader tier never drops an optimization from a
stricter tier.

| Tier                 | Optimization boundary                                                                                                                                 |
| -------------------- | ----------------------------------------------------------------------------------------------------------------------------------------------------- |
| `lossless` (default) | The selected deployment's reference execution plus all unconditional bit-exact replacements                                                           |
| `extra-high`         | Everything in `lossless`, plus only request-gated DiT/VAE kernel fusions; the tier does not itself enable sparse, caching, or other approximate paths |
| `high`               | Everything in `extra-high`, plus model-owned high-only paths such as an audited Cache-DiT policy or lower-precision VAE decode                        |

Use `extra-high` when you want to isolate fusion wins from approximate
acceleration. A tier may be a no-op when the active model has no eligible path.
Separately configured quantization, attention, or caching options still apply.
See [Fused Kernels](/docs/sglang-diffusion/fused_kernels) for the current
request-gated families and their numerical contracts.

## Output-Preserving / Lossless-Style Levers

These settings should preserve model behavior while changing residency, parallelism, kernels, or scheduling. They are the first choices for production tuning.

<table style={{width: "100%", borderCollapse: "collapse", tableLayout: "fixed"}}>
  <colgroup>
    <col style={{width: "24%"}} />

    <col style={{width: "30%"}} />

    <col style={{width: "46%"}} />
  </colgroup>

  <thead>
    <tr style={{borderBottom: "2px solid #d55816"}}>
      <th style={{textAlign: "left", padding: "10px 12px", fontWeight: 700}}>Lever</th>
      <th style={{textAlign: "left", padding: "10px 12px", fontWeight: 700}}>Use when</th>
      <th style={{textAlign: "left", padding: "10px 12px", fontWeight: 700}}>Docs</th>
    </tr>
  </thead>

  <tbody>
    <tr>
      <td style={{padding: "9px 12px", fontWeight: 500}}><code>--performance-mode</code></td>
      <td style={{padding: "9px 12px"}}>You want a safe preset for speed or memory without overriding explicit flags.</td>
      <td style={{padding: "9px 12px"}}><a href="/docs/sglang-diffusion/deployment_cookbook">Deployment and Performance Modes</a></td>
    </tr>

    <tr>
      <td style={{padding: "9px 12px", fontWeight: 500}}>Breakable CUDA graph</td>
      <td style={{padding: "9px 12px"}}>A supported pipeline serves a fixed set of shapes and eager execution is launch-bound.</td>
      <td style={{padding: "9px 12px"}}><a href="/docs/sglang-diffusion/api/cli">CLI reference</a></td>
    </tr>

    <tr>
      <td style={{padding: "9px 12px", fontWeight: 500}}>Offload, FSDP, CFG parallelism</td>
      <td style={{padding: "9px 12px"}}>GPU memory, multi-GPU residency, or CFG branch splitting is the main bottleneck.</td>
      <td style={{padding: "9px 12px"}}><a href="/docs/sglang-diffusion/deployment_cookbook">Deployment and Performance Modes</a></td>
    </tr>

    <tr>
      <td style={{padding: "9px 12px", fontWeight: 500}}>Sequence parallelism</td>
      <td style={{padding: "9px 12px"}}>Long image/video sequences need sequence-level parallelism.</td>
      <td style={{padding: "9px 12px"}}><a href="/docs/sglang-diffusion/ring_sp_performance">Sequence Parallelism</a></td>
    </tr>

    <tr>
      <td style={{padding: "9px 12px", fontWeight: 500}}><code>--encoder-parallel</code></td>
      <td style={{padding: "9px 12px"}}>Text/image encoding is a visible share of the request and the DiT replica sits idle during it.</td>
      <td style={{padding: "9px 12px"}}><a href="/docs/sglang-diffusion/encoder_parallel">Encoder Parallelism</a></td>
    </tr>

    <tr>
      <td style={{padding: "9px 12px", fontWeight: 500}}>Attention backend</td>
      <td style={{padding: "9px 12px"}}>Kernel choice dominates DiT latency or memory.</td>
      <td style={{padding: "9px 12px"}}><a href="/docs/sglang-diffusion/attention_backends">Attention Backends</a></td>
    </tr>

    <tr>
      <td style={{padding: "9px 12px", fontWeight: 500}}>Fused kernels</td>
      <td style={{padding: "9px 12px"}}>You want to know which elementwise chains are already fused, or to opt into the request-gated set.</td>
      <td style={{padding: "9px 12px"}}><a href="/docs/sglang-diffusion/fused_kernels">Fused Kernels</a></td>
    </tr>

    <tr>
      <td style={{padding: "9px 12px", fontWeight: 500}}>Dynamic batching</td>
      <td style={{padding: "9px 12px"}}>Serving many compatible requests concurrently.</td>
      <td style={{padding: "9px 12px"}}><a href="/docs/sglang-diffusion/dynamic_batching">Inference Batching</a></td>
    </tr>
  </tbody>
</table>

## Quality-Tradeoff / Lossy Or Approximate Levers

These techniques can change the denoising path, numerical representation, or generated output. They are useful after you have a baseline and an acceptance criterion for quality.

<table style={{width: "100%", borderCollapse: "collapse", tableLayout: "fixed"}}>
  <colgroup>
    <col style={{width: "24%"}} />

    <col style={{width: "30%"}} />

    <col style={{width: "46%"}} />
  </colgroup>

  <thead>
    <tr style={{borderBottom: "2px solid #d55816"}}>
      <th style={{textAlign: "left", padding: "10px 12px", fontWeight: 700}}>Lever</th>
      <th style={{textAlign: "left", padding: "10px 12px", fontWeight: 700}}>Tradeoff</th>
      <th style={{textAlign: "left", padding: "10px 12px", fontWeight: 700}}>Docs</th>
    </tr>
  </thead>

  <tbody>
    <tr>
      <td style={{padding: "9px 12px", fontWeight: 500}}>Cache-DiT</td>
      <td style={{padding: "9px 12px"}}>Skips selected DiT block or step computation based on cache decisions.</td>
      <td style={{padding: "9px 12px"}}><a href="/docs/sglang-diffusion/cache_dit">Cache-DiT</a></td>
    </tr>

    <tr>
      <td style={{padding: "9px 12px", fontWeight: 500}}>TeaCache</td>
      <td style={{padding: "9px 12px"}}>Reuses residuals when consecutive denoising steps are similar enough.</td>
      <td style={{padding: "9px 12px"}}><a href="/docs/sglang-diffusion/teacache">TeaCache</a></td>
    </tr>

    <tr>
      <td style={{padding: "9px 12px", fontWeight: 500}}>Progressive resolution</td>
      <td style={{padding: "9px 12px"}}>Runs early denoising at lower latent resolution for supported pipelines.</td>
      <td style={{padding: "9px 12px"}}><a href="/docs/sglang-diffusion/progressive_resolution">Progressive Resolution Generation</a></td>
    </tr>

    <tr>
      <td style={{padding: "9px 12px", fontWeight: 500}}>Quantization</td>
      <td style={{padding: "9px 12px"}}>Uses lower-precision transformer weights or activations.</td>
      <td style={{padding: "9px 12px"}}><a href="/docs/sglang-diffusion/quantization">Quantization</a></td>
    </tr>
  </tbody>
</table>

## Practical Order

1. Establish a baseline with the target model, resolution, frame count, step count, and GPU type.
2. Select `--performance-mode` and explicit residency or parallelism flags.
3. Compare `quality=lossless` with `quality=extra-high` to isolate the request-gated fusion set.
4. Compare breakable CUDA graph against eager execution for supported fixed-shape pipelines. Pass every served resolution to `--warmup-resolutions` and confirm capture in the server log. Models with request-gated DiT fusions cannot combine those fusions with a graph captured from the lossless branches.
5. Tune attention backend and batching for the deployment pattern.
6. Profile if the bottleneck is unclear.
7. Add `quality=high`, caching, progressive resolution, or quantization only after comparing output quality against your acceptance target.

## Per-model tuning starting points

Warmup and breakable CUDA graph (BCG) solve different problems. `--warmup-mode request` runs a warmup copy derived from the first request to prime one-time compilation and caches; it does not remove recurring Python launches from each denoising step. BCG captures supported DiT segments and can reduce that recurring launch overhead for captured shapes. Use the table below as a first experiment, then keep a lever only when profiling confirms that it addresses the active bottleneck.

| Observed bottleneck or constraint                                                     | Example models                                                                                                                             | First experiment                                              | What to verify                                                                                                                                                                                                                               |
| ------------------------------------------------------------------------------------- | ------------------------------------------------------------------------------------------------------------------------------------------ | ------------------------------------------------------------- | -------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------- |
| The first execution of a shape pays substantial compilation or cache setup            | ERNIE-Image-Turbo, GLM-Image, FastWan / TurboWan family, FLUX.2-klein-4B, LingBot-World, LongLive, Cosmos3, SANA, LongCat-Image-Edit-Turbo | Compare `--warmup-mode request` with `off`                    | Separate warmup time, first-real-request latency, and warmed steady-state latency. Request-based warmup is primarily a benchmark aid and still consumes time before the first real request completes.                                        |
| A supported fixed-shape pipeline shows recurring host-launch gaps between GPU kernels | GLM-Image, Z-Image-Turbo, Qwen-Image-2512, Ideogram-4, LongCat-Image, LTX-2 / LTX-2.3, SANA-1.5 1.6B, JoyEcho                              | Compare eager execution with `--enable-breakable-cuda-graph`  | Confirm graph capture and replay in the server log, then compare steady-state latency and GPU memory. Pass additional production shapes through `--warmup-resolutions`; without it, BCG captures only the model's default warmup resolution. |
| The model does not fit while both DiT stages are resident                             | Wan2.2-T2V-A14B, Wan2.2-I2V-A14B, LingBot-Video-MoE                                                                                        | Start with `--dit-layerwise-offload`                          | Verify peak GPU memory first, then measure the transfer overhead. Add warmup separately if first-run compilation is also material.                                                                                                           |
| Dynamic masks, metadata, or image conditions make graph replay ineffective            | Qwen-Image-Edit family, FireRed-Image-Edit                                                                                                 | Keep eager execution as the baseline                          | Profile before attempting BCG. Dynamic per-step host work can outweigh graph replay savings, and unsupported pipelines fall back to eager execution.                                                                                         |
| Kernels or collectives already dominate a many-step workload                          | Qwen-Image, FLUX.1-dev, FLUX.2-dev, LTX-2, Wan2.1 14B, HunyuanVideo, MOVA                                                                  | Warm the baseline, then profile before enabling another lever | A small host gap limits the benefit available from BCG. Prioritize kernels, attention, parallelism, or residency according to the trace.                                                                                                     |
| A one-shot or very few-step workload cannot amortize warmup                           | FastVideo-FastH3 (4-step)                                                                                                                  | Start with `--warmup-mode off`                                | Compare end-to-end latency including warmup. Use request warmup only when intentionally separating cold-start setup from the measured request.                                                                                               |

The example assignments come from single-GPU NVIDIA B300/GB300 profiles and are directional rather than an exhaustive compatibility list. A model can match more than one row as resolution, frame count, step count, parallelism, or GPU type changes; the measured bottleneck takes precedence over the model name.

BCG is enabled only for model and pipeline configurations accepted by the runtime support check. It captures the default warmup shape automatically. Use `--warmup-resolutions` for additional served resolutions; for video and variable prompt lengths, set `--warmup-num-frames` and `--bcg-text-buckets` to cover the intended workload. Requests that do not match a captured signature fall back to eager execution.

Warmup and BCG do not intentionally trade output quality for speed, but that is not a guarantee of byte-identical output. Different execution paths can introduce numerical differences, and some pipelines customize the scheduler or schedule used by a synthetic warmup request. Before production rollout, compare output hashes when bitwise stability is required, otherwise run the project's quality acceptance check.

Performance results are configuration-dependent. Record the exact checkpoint revision, GPU, precision, resolution, frame count, step count, parallelism, command line, and whether the measurement includes warmup. Re-profile with [Profiling](/docs/sglang-diffusion/profiling) on the target workload rather than transferring a percentage from another model or shape.

## Diagnostics

[Profiling](/docs/sglang-diffusion/profiling) is not an optimization technique by itself. It belongs in the performance workflow because it tells you which stage, kernel, or denoising step is worth optimizing before you change multiple levers.

## References

* [Deployment and Performance Modes](/docs/sglang-diffusion/deployment_cookbook)
* [Attention Backends](/docs/sglang-diffusion/attention_backends)
* [Fused Kernels](/docs/sglang-diffusion/fused_kernels)
* [Sequence Parallelism](/docs/sglang-diffusion/ring_sp_performance)
* [Caching Strategies](/docs/sglang-diffusion/caching-acceleration)
* [Profiling](/docs/sglang-diffusion/profiling)
