Skip to main content

1. Model Introduction

Krea-2 is Krea’s photorealistic text-to-image family, built as a single-stream MMDiT with a Qwen3-VL text encoder and Qwen-Image VAE. Both public variants use the same native SGLang pipeline and differ mainly in their sampling target. Choose Turbo for interactive generation: it is distilled to 8 steps with guidance_scale=1.0. Choose Raw when maximum fidelity matters more than latency: it uses roughly 52 steps with classifier-free guidance. Neither checkpoint is an image-editing model; use the Qwen-Image-Edit or FLUX.2 path when an input image must be preserved or transformed.

2. SGLang-diffusion Installation

SGLang-diffusion offers multiple installation methods. You can choose the most suitable installation method based on your hardware platform and requirements. Please refer to the official SGLang-diffusion installation guide for installation instructions.

3. Model Deployment

This section covers deploying Krea-2-Turbo for fast, high-quality image generation.

3.1 Basic Configuration

Krea-2-Turbo generates high-quality images in only 8 inference steps. Launch the server with:
Command
The step count and guidance scale are request-time settings (see API Usage); Krea-2-Turbo defaults to 8 steps with guidance_scale = 1.0.

3.2 Configuration Tips

See Performance Optimization for acceleration features and their runtime requirements.
  • --num-gpus: Number of GPUs to use.
  • Multi-GPU (tensor and/or sequence parallelism): see Section 3.3.

3.3 Multi-GPU: tensor and sequence parallelism

Krea-2 supports two multi-GPU axes that can be combined; --num-gpus must equal tp_size × ulysses_degree.
  • Tensor parallelism (--tp-size N) shards the DiT weights across GPUs, lowering per-GPU VRAM. Krea-2’s attention heads (48 query / 12 KV) and text heads (20) are divisible by a tp size of 1, 2, or 4.
  • Sequence parallelism / Ulysses (--ulysses-degree N) shards the image-token sequence across GPUs while keeping the text prefix replicated. It does not shard weights (per-GPU VRAM is unchanged), but its output is bitwise-identical to single-GPU. It currently requires a single prompt per request (ragged/padded multi-prompt batches under SP are not supported — use --tp-size for those).
Command
Measured on 2× H200 (Krea-2-Turbo, 8 steps, 1024×1024): --tp-size 2 and --ulysses-degree 2 each give ~1.7× denoise speedup over single-GPU; the hybrid TP=2 × SP=2 reaches ~2.8× on 4 GPUs. Choosing: on memory-constrained GPUs prefer --tp-size (it shards the ~24 GB DiT, e.g. ~38 GB → ~27 GB per GPU on 2 GPUs); on large-VRAM GPUs sequence parallelism is marginally faster and numerically exact, and the two compose for the highest throughput.

4. API Usage

For complete API documentation, please refer to the official API usage guide.

4.1 Generate an Image

Generate an image with the OpenAI-compatible images API:
Example
You can also generate a single image from the command line:
Command

4.2 Advanced Usage

4.2.1 Cache-DiT Acceleration

SGLang integrates Cache-DiT, a caching acceleration engine for Diffusion Transformers (DiT), to speed up inference with minimal quality loss. Enable it by setting SGLANG_CACHE_DIT_ENABLED=true. For more details, see the SGLang Cache-DiT documentation. Cache-DiT works for both Krea-2 variants with no extra configuration: SGLang tracks each request’s classifier-free-guidance mode, so Krea-2-Turbo (no CFG, guidance_scale = 1.0) and Krea-2-Raw (CFG, guidance_scale ≈ 4.5) both cache correctly and automatically. Basic Usage
Command
Measured per-image denoise speedup with the default cache settings (NVIDIA H200, 1024x1024, seed 0): Caching has the most headroom on Raw’s longer schedule; the 8-step distilled Turbo has only a few cacheable steps after warmup. Advanced Usage
  • DBCache Parameters: DBCache controls block-level caching behavior:
ParameterEnv VariableDefaultDescription
FnSGLANG_CACHE_DIT_FN1Number of first blocks to always compute
BnSGLANG_CACHE_DIT_BN0Number of last blocks to always compute
WSGLANG_CACHE_DIT_WARMUP4Warmup steps before caching starts
RSGLANG_CACHE_DIT_RDT0.24Residual difference threshold
MCSGLANG_CACHE_DIT_MC3Maximum continuous cached steps
  • TaylorSeer Configuration: TaylorSeer improves caching accuracy using Taylor expansion (best suited to the longer Raw schedule; not recommended for the 8-step Turbo):
ParameterEnv VariableDefaultDescription
EnableSGLANG_CACHE_DIT_TAYLORSEERfalseEnable TaylorSeer calibrator
OrderSGLANG_CACHE_DIT_TS_ORDER1Taylor expansion order (1 or 2)
Combined Configuration Example (Krea-2-Raw, default cache settings shown explicitly):
Command

4.2.2 Memory and Component Residency

Krea-2’s DiT is ~24 GB in bf16 (the bulk of the model). On memory-constrained GPUs you can keep less of it resident:
  • --component-residency dit=layerwise-offload: stream the DiT’s transformer blocks layer-by-layer with async host-to-device prefetch overlap, so only a small working set stays on the GPU. This is the primary way to fit Krea-2 on a single consumer / 32 GB-class card, at a modest latency cost. Tune the memory/latency trade-off with --dit-offload-prefetch-size (0.0 prefetches one layer for the lowest memory; larger values prefetch more layers — faster but more memory).
  • --component-residency dit=component-offload: keep the complete DiT on CPU between denoising uses. This and layerwise offload are distinct modes; do not combine them for the same component.
  • --component-residency text_encoder=component-offload: offload the Qwen3-VL text encoder while it is idle during denoising.
  • --component-residency vae=component-offload: offload the VAE between uses.
  • --pin-cpu-memory: pin host memory for offload. Add only as a temporary workaround if you hit CUDA error: invalid argument.
The legacy --dit-layerwise-offload, --dit-cpu-offload, --text-encoder-cpu-offload, and --vae-cpu-offload forms remain accepted. If both legacy DiT offload flags are enabled, layerwise offload is the effective DiT mode. On large-VRAM GPUs (e.g. H200), keep everything resident (offloads off) for the fastest latency.

5. Benchmark

Test Environment:
  • Hardware: NVIDIA H200 GPU (1x)
  • Model: krea/Krea-2-Turbo (8 inference steps)
  • sglang diffusion version: 0.5.13
Server Command (used for both benchmarks below):
Command

5.1 Generate an image

Benchmark Command:
Command
Result:
Output

5.2 Generate images with high concurrency

Benchmark Command:
Command
Result:
Output