1. Model Introduction
LongLive 2.0 is NVIDIA’s 4-step text/image-to-video model distilled from Wan2.2-TI2V-5B. Its main strength is extending few-step causal generation across prompt changes, so a single request can produce multi-shot sequences without paying a full diffusion schedule for every shot. Choose it for low-step long or multi-shot generation rather than maximum one-shot fidelity. The SGLang path uses a Diffusers conversion of the official weights, and scene continuity still depends on prompt-block and sink settings; validate transitions on the target storyboard. The model weights use the NVIDIA Open Model License. See the paper and GitHub repository for training details.2. SGLang-diffusion Installation
Please refer to the official SGLang-diffusion installation guide for installation instructions.3. Deployment
Command
auto mode, GPUs with at least 60 GiB available keep the DiT, text encoder,
and VAE resident. This uses about 44 GiB on H200 for the 832x480 preset and
avoids moving encoder and decoder layers from host memory on every request.
Smaller GPUs retain the layerwise-offload defaults.
If the GPU runs out of memory, move the text encoder, VAE, and DiT to CPU between stages:
Command
Rabinovich/LongLive-2.0-5B-Diffusers is the Diffusers-format conversion of the official Efficient-Large-Model/LongLive-2.0-5B weights.
4. Generation
4.1 Single prompt
Generate one clip without starting a server:Command
4.2 Multi-shot long video
Multi-shot prompts are sampling parameters, so pass them through the Python API:Python
chunks_per_shot causal blocks before the next prompt is used. The multi-shot defaults mirror the original LongLive prompt-block settings.
4.3 Key parameters
These are SGLang request parameters. Original LongLive configs use latent-framenum_output_frames; SGLang exposes output-video num_frames.
num_frames: 61 in the examples. This maps to 16 latent frames, while the original release config defaults to 128 latent frames.num_inference_steps: 4, matching originalsampling_steps.guidance_scale: 1.0, matching the original inference config.height/width: 704 / 1280 by default, matching original latent H/W 44 / 80 with 16x spatial compression.shot_prompts,chunks_per_shot,scene_cut_prefix,multi_shot_sink, andmulti_shot_rope_offset: SGLang request fields for the original prompt-block and multi-shot behavior.
4.4 Image-to-video
Pass a first frame with--image-path to condition the clip on an image:
Command
5. Notes
num_framesmust map to a whole number of causal blocks. The latent frame count is(num_frames - 1) / 4 + 1and must be divisible by 8. For example, 61, 125, and 189 frames give 16, 32, and 48 latent frames.- SGLang supports T2V sizes 1280x704, 704x1280, 832x480, and 480x832.
- I2V request images follow the Wan TI2V preprocessing path in SGLang. This is different from the original LongLive dataset resize path.
- For multi-shot runs, set
num_framesto matchlen(shot_prompts) * chunks_per_shot * 8latent frames, that isnum_frames = (len(shot_prompts) * chunks_per_shot * 8 - 1) * 4 + 1.
