sglang-diffusion-benchmark-profile

Use when benchmarking denoise latency or profiling a diffusion bottleneck in SGLang.

Install
npx skills add 'https://github.com/sgl-project/sglang/tree/main/python/sglang/multimodal_gen/.claude/skills/sglang-diffusion-benchmark-profile'
Download bundle ↓
main · a9fb1c3Scanned 2026-09-17

Contributors

GitHub-linked commit authors for this SKILL.md at the saved revision. Co-authors and history before file renames are not included.

File history ↗
View on GitHub
← Back to SKILL.md

name: benchmark-and-profile-reference description: Reference commands and workflow for denoise benchmarks, perf dumps, and torch.profiler analysis in SGLang Diffusion.

SGLang Diffusion Benchmark and Profile Guide

Primary Metric: Denoise Latency

  • Denoise latency is the total DiT forward-pass time across all inference steps.
  • It is the dominant cost for diffusion inference and the main optimization target.
  • End-to-end latency and peak memory are secondary sanity checks.

Correctness First: Faster but incorrect output is not an improvement. Always compare generated images or videos against a reference baseline before and after any change.

Scope

This guide intentionally stops at:

  • checked-in denoise benchmarks
  • structured perf dumps
  • torch.profiler trace capture
  • hotspot ranking
  • mapping hotspots to known fast paths

If the hotspot survives this checklist, package the perf dump, profiler trace, exact command, and shape/topology notes for the appropriate kernel, Nsight, or framework-specific optimization workflow. Do not grow this skill back into a general Nsight or kernel-authoring guide.

Prerequisites

ENV_PY=python/sglang/multimodal_gen/.claude/skills/sglang-diffusion-benchmark-profile/scripts/diffusion_skill_env.py
BENCH_PY=python/sglang/multimodal_gen/.claude/skills/sglang-diffusion-benchmark-profile/scripts/bench_diffusion_denoise.py
ROOT=$(python3 "$ENV_PY" print-root)
cd "$ROOT"
python3 "$ENV_PY" check-write-access >/dev/null

export HF_TOKEN=<your_hf_token>  # required for gated repos such as black-forest-labs/FLUX.*
export FLASHINFER_DISABLE_VERSION_CHECK=1
# Required for correctly attributed stage-level denoise/decode timings. The
# checked-in benchmark helper sets this by default unless you explicitly set 0.
export SGLANG_DIFFUSION_SYNC_STAGE_PROFILING=1
# Leave CUDA_VISIBLE_DEVICES unset to let the preset helper select the number
# of idle GPUs it requires. For manual runs, set --count to that command's
# exact --num-gpus value.

ASSET_DIR=$(python3 "$ENV_PY" print-assets-dir --mkdir)
BENCH_DIR=$(python3 "$ENV_PY" print-output-dir --kind benchmarks --mkdir)
PROFILE_DIR=$(python3 "$ENV_PY" print-output-dir --kind profiles --mkdir)
CONFIG_DIR="${BENCH_DIR}/generated_configs"
mkdir -p "${CONFIG_DIR}"
export PROFILE_DIR

check() {
  local label="$1"
  shift
  "$@" &>/dev/null && echo "[OK]  $label" || echo "[MISS] $label"
}

check "sglang" python3 -c "import sglang"
check "torch+CUDA" python3 -c "import torch; assert torch.cuda.is_available()"
check "torch.profiler" python3 -c "import torch.profiler"

Native Backend Gate

Every benchmark and profile result in this guide must come from the native SGLang diffusion backend.

If the command log contains any of:

  • Falling back to diffusers backend
  • Using diffusers backend
  • Loaded diffusers pipeline

then stop immediately:

  • do not record the perf dump or trace as valid benchmark evidence
  • do not compare it against other runs
  • do not continue to hotspot ranking or kernel optimization
  • first fix backend selection so the model stays on the native SGLang diffusion path

The checked-in benchmark helper pins --backend=sglang so native presets fail fast instead of silently falling back through --backend=auto. Do the same for manual native profiling commands unless you are intentionally collecting a diffusers baseline.

Environment notes:

  • all commands below assume you are inside the configured diffusion container shell
  • export HF_TOKEN before any gated Hugging Face model run
  • export FLASHINFER_DISABLE_VERSION_CHECK=1 before any benchmark or profiler run
  • keep SGLANG_DIFFUSION_SYNC_STAGE_PROFILING=1 for stage-level comparisons; without it, asynchronous GPU work can be charged to a later stage
  • re-run print-idle-gpus before each perf command if GPU availability may have changed
  • keep benchmark commands within 4 GPUs or fewer

Download input images required by some presets:

wget -O "${ASSET_DIR}/cat.png" \
  https://huggingface.co/datasets/huggingface/documentation-images/resolve/main/diffusers/cat.png
wget -O "${ASSET_DIR}/mova_single_person.jpg" \
  https://github.com/OpenMOSS/MOVA/raw/main/assets/single_person.jpg

Benchmark Presets

Treat "$BENCH_PY" as the source of truth for preset order.

Nightly diffusion comparison is server/API based (sglang serve plus requests). This skill stays on sglang generate for local benchmarking and profiling, but the nightly-aligned presets in bench_diffusion_denoise.py mirror scripts/ci/utils/diffusion/comparison_configs.json on model, task, prompt, reference image, size, frames, seed, GPU count, and SGLang serve args. If comparison_configs.json omits sampling params such as steps or guidance, the nightly-aligned sglang generate preset omits them too and relies on the same runtime defaults. When in doubt, re-check that JSON before trusting this reference.

List the current preset order:

PYTHONPATH=python python3 "$BENCH_PY" --list-models

Check that the nightly presets still match the Nvidia nightly comparison config:

PYTHONPATH=python python3 "$BENCH_PY" --validate-nightly-alignment

Run one preset and save a perf dump:

PYTHONPATH=python python3 "$BENCH_PY" \
  --model ltx2 \
  --label baseline \
  --output-dir "${BENCH_DIR}"

The helper defaults to eager. Add --torch-compile only for a labeled compile control. --no-torch-compile remains accepted for compatibility but is no longer required.

Run one explicit quality or BCG comparator with --quality {lossless,extra-high,high} and --breakable-cuda-graph. BCG and torch.compile are intentionally mutually exclusive in this helper. An extra-high/high+BCG command is only a compatibility probe: it is invalid if request-scoped DiT fusions mount after the lossless warmup graphs were captured. When a preset has explicit width and height, the helper declares that same --warmup-resolutions value automatically. Video presets with an explicit frame count also declare the matching --warmup-num-frames:

PYTHONPATH=python python3 "$BENCH_PY" \
  --model longcat-image \
  --quality extra-high \
  --breakable-cuda-graph \
  --label bcg-extra-high \
  --output-dir "${BENCH_DIR}"

For optimization discovery, use the full repeated matrix. It runs Eager/BCG/BCG/Eager at lossless, then the same sequence at extra-high and high, while holding one GPU set and one isolated checkpoint cache. The extra-high/high+BCG cells test whether the combination is actually supported; do not average them when the runtime rejects the combination or the helper detects a late quality-fusion mount. The helper hashes every generated image, video, audio, or 3D mesh artifact. It first requires the two Eager rows at each quality to agree, then rejects any BCG row whose hash differs from that Eager reference. Cleanup occurs only after all twelve runs, including on failure or interruption:

MODEL_CACHE_ROOT=/path/to/task-owned/model-caches
PYTHONPATH=python python3 "$BENCH_PY" \
  --model longcat-image \
  --quality-bcg-matrix \
  --label h200 \
  --output-dir "${BENCH_DIR}" \
  --model-cache-root "${MODEL_CACHE_ROOT}" \
  --cleanup-model-cache

Before starting, confirm the chosen GPU set has no foreign process and remains unchanged through every run boundary. The helper rejects a BCG row unless its log contains [Diffusion BCG] captured and contains none of: support-gate disable, capture failure, serving signature MISSED, a message that no graph will be captured, or a request-scoped quality-gated DiT fusion mounted after capture. Do not average rejected rows with valid results.

BCG signatures include more than width and height. The helper maps an explicit video request frame count to --warmup-num-frames, while --warmup-resolutions declares WxH. Other temporal or conditioning inputs can still differ from the captured signature. The helper marks such a row invalid; fix the model's BCG warmup/padding contract before claiming a speedup.

The helper sets SGLANG_DIFFUSION_SYNC_STAGE_PROFILING=1 for accurate stage attribution. Set it to 0 explicitly only when collecting an e2e-only run and do not compare its per-stage values with synchronized results.

For downloaded checkpoints, isolate and clean the model cache after the preset finishes. Cleanup also runs after an error or interruption, and appends a JSONL record with pre/post byte and weight-file counts:

MODEL_CACHE_ROOT=/path/to/task-owned/model-caches
PYTHONPATH=python python3 "$BENCH_PY" \
  --model longcat-image \
  --label baseline \
  --output-dir "${BENCH_DIR}" \
  --model-cache-root "${MODEL_CACHE_ROOT}" \
  --cleanup-model-cache

The helper refuses to reuse an existing per-run cache directory and never redirects SGLANG_CACHE_DIR, so compiled kernel caches remain separate. Never point this option at a shared Hugging Face or ModelScope cache.

When a machine already exposes a read-only Hugging Face cache, seed the task-owned cache with a copy-on-write directory overlay instead of copying its checkpoints. Immutable blobs and snapshot payloads remain symlinks, while metadata directories stay writable so a partial seed can download missing files into the task cache. The option may be repeated. Cleanup removes only the task-owned overlay and new downloads; it never follows links or modifies the seed cache:

PYTHONPATH=python python3 "$BENCH_PY" \
  --model longcat-image \
  --quality-bcg-matrix \
  --label h100 \
  --output-dir "${BENCH_DIR}" \
  --model-cache-root "${MODEL_CACHE_ROOT}" \
  --seed-model-cache-root /path/to/read-only/huggingface \
  --cleanup-model-cache

Each seed path must be either a Hugging Face home containing hub/ or the hub directory itself. Do not seed from a task cache that is being cleaned.

Run the LTX-2.3 one-stage skill preset:

PYTHONPATH=python python3 "$BENCH_PY" \
  --model ltx23-one-stage \
  --label baseline \
  --output-dir "${BENCH_DIR}"

Run the nightly-aligned LTX-2.3 TI2V two-stage preset:

PYTHONPATH=python python3 "$BENCH_PY" \
  --model ltx23-ti2v-two-stage \
  --label baseline \
  --output-dir "${BENCH_DIR}"

Run the LTX-2.3 two-stage skill preset:

PYTHONPATH=python python3 "$BENCH_PY" \
  --model ltx23-two-stage \
  --label baseline \
  --output-dir "${BENCH_DIR}"

Run the current-source MiniMax-H3 T2VA preset. The helper forces eager mode for this model even when --torch-compile is requested:

export CUDA_VISIBLE_DEVICES=$(python3 "$ENV_PY" print-idle-gpus --count 4)
PYTHONPATH=python python3 "$BENCH_PY" \
  --model minimax-h3-t2va \
  --label baseline \
  --output-dir "${BENCH_DIR}"

Run the full preset sweep only when you have enough GPU time for both the nightly-aligned cases and the source-tracked extras:

PYTHONPATH=python python3 "$BENCH_PY" \
  --all \
  --label prXXXX \
  --output-dir "${BENCH_DIR}"

Nightly-aligned presets come first, followed by current-source extras from the registry / GPU test cases, then broader skill-only stress presets.

Use the preset categories this way:

  • Nightly-aligned: exact mirrors of scripts/ci/utils/diffusion/comparison_configs.json; use these when the goal is apples-to-apples comparison with CI / nightly coverage.
  • Current-source extras: models or request shapes with explicit support evidence in the current registry, GPU cases, compatibility matrix, pipeline files, or unit tests, but without a nightly comparison case yet.
  • Skill-only stress / coverage presets: extra profiling scenarios kept by this skill to stress a topology, high-resolution path, multi-GPU mode, or model-specific stage. These may be older than the latest registry additions, so re-check the active source tree before treating them as support-matrix commitments.
PresetModelNightlyNotes
fluxblack-forest-labs/FLUX.1-devYes: flux1_dev_t2i_1024Prompt, 1024x1024, seed 42, 2 GPUs, TP size 2, resident DiT; no explicit steps/guidance override
flux2black-forest-labs/FLUX.2-devYes: flux2_dev_t2i_1024Prompt, 1024x1024, seed 42, 2 GPUs, TP size 2, resident DiT; no explicit steps/guidance override
qwenQwen/Qwen-Image-2512Yes: qwen_image_2512_t2i_1024Prompt, 1024x1024, seed 42, 2 GPUs, TP size 2; no explicit steps/guidance override
qwen-editQwen/Qwen-Image-Edit-2511Yes: qwen_image_edit_2511Uses the nightly cat image and edit prompt, 2 GPUs, TP size 2
zimageTongyi-MAI/Z-Image-TurboYes: zimage_turbo_t2i_1024Prompt, 1024x1024, seed 42, 2 GPUs, TP size 2; no explicit steps/guidance override
wan-t2vWan-AI/Wan2.2-T2V-A14B-DiffusersYes: wan22_t2v_a14b_720p1280x720, 81 frames, 4 GPUs, CFG parallel, Ulysses degree 2, text encoder CPU offload and pinned CPU memory
wan-ti2vWan-AI/Wan2.2-TI2V-5B-DiffusersYes: wan22_ti2v_5b_720pNightly cat image and motion prompt, 1280x720, 81 frames, seed 42
ltx23-ti2v-two-stageLightricks/LTX-2.3Yes: ltx2.3_twostage_ti2v_2gpusNightly cat image, motion prompt, LTX2TwoStagePipeline, 2 GPUs, --cfg-parallel-size 2, 768x512, 121 frames, seed 42
ideogram4-fp8ideogram-ai/ideogram-4-fp8Yes: ideogram4_fp8_t2i_2gpuPrompt, 1024x1024, seed 42, 2 GPUs, TP size 2, FlashAttention backend; sampling preset owns steps/guidance
cosmos3-super-t2vnvidia/Cosmos3-SuperYes: cosmos3_super_t2v_2gpuPrompt, 1280x720, 81 frames, seed 42, 2 GPUs, TP size 2, guardrails disabled for benchmark isolation
cosmos3-super-t2v-cfg2tp2nvidia/Cosmos3-SuperNoExplicit four-GPU TP2 x CFG2 throughput comparator. On H200 it was 48.00% faster end to end than TP2, but the topology changed the deterministic output (SSIM 0.914244, PSNR 29.469771 dB), so do not treat it as lossless-equivalent or select it automatically.
wan-i2vWan-AI/Wan2.2-I2V-A14B-DiffusersYes: wan22_i2v_a14b_720pNightly cat image and motion prompt, 1280x720, 81 frames, 4 GPUs, CFG parallel, Ulysses degree 2, text encoder CPU offload and pinned CPU memory
minimax-h3-t2vaMiniMaxAI/MiniMax-H3Yes: minimax_h3_t2va_5sH3 FL2VA-partition T2VA baseline: 1344x768 resolved canvas, 5 seconds / 124 frames at 24 fps, 50 joint video-audio steps, 4 GPUs, TP2 + Ulysses2, eager BF16/FP32. The helper writes H3's request contract to a generated config.
fasth3-t2va-vsaFastVideo/FastVideo-FastH3-4-step-Preview-v1-VSA-DataFreeNoFastH3 4-step distilled T2VA on the trained VSA-H3 backend: 1344x768, 10 seconds / 243 frames, five sigma points = four DiT forwards, 4 GPUs, Ulysses 4, eager, 2-step warmup request. Compare against --attention-backend fa on the same weights for the dense-fallback gap.
longcat-imagemeituan-longcat/LongCat-ImageNoEager DiT baseline at 1024x1024, 50 steps, guidance 4.5; prompt rewrite is disabled so Qwen2.5-VL does not contaminate the DiT A/B.
longcat-image-editmeituan-longcat/LongCat-Image-EditNoNative edit baseline using the public SGLang edit fixture. Its 1536x1024 source resolves to 1264x848 under the checkpoint's roughly-one-megapixel aspect-ratio rule, and the BCG comparator captures that exact serving canvas; prompt rewrite is disabled to isolate the DiT.
longcat-image-edit-turbomeituan-longcat/LongCat-Image-Edit-TurboNoMatching distilled edit baseline using the same public fixture, prompt, and 1264x848 BCG canvas. Its registered sampling class owns the eight-step, guidance-1 schedule.
qwen-edit-baseQwen/Qwen-Image-EditNoCovers the original native QwenImageEditPipelineConfig, which is distinct from the 2509/2511 edit-plus paths; public SGLang edit fixture, 1024x1024.
qwen-image-layeredQwen/Qwen-Image-LayeredNoNative layered-image path using the same public reference image and four-frame request as the GPU server case, at the registered 640x640 canvas.
stable-diffusion-3.5-mediumstabilityai/stable-diffusion-3.5-medium-diffusersNoRepresentative native StableDiffusion3PipelineConfig path at 1024x1024. The repository is gated, so export HF_TOKEN; an unauthenticated run is a recorded access blocker, not model evidence.
sana-videoEfficient-Large-Model/SANA-Video_2B_480p_diffusersNoCI-sized T2V baseline: 832x480, 17 frames, 8 steps, guidance 6.0. The BCG comparator declares the same 17-frame warmup shape. Compare all three tiers; extra-high and high enable the BF16-input first linear-attention GEMM while retaining FP32 output and the FP32 second GEMM.
sana-wm-bidirectionalEfficient-Large-Model/SANA-WM_bidirectionalNoDense two-stage TI2V baseline at the native 1280x704 shape, 49 frames, 16 fps, 20 steps, guidance 4.5, and a 48-frame forward/left action program. Uses the shared cat fixture.
sana-wm-streamingEfficient-Large-Model/SANA-WM_streamingNoMatching offline chunk-causal two-stage baseline with the streaming DiT and chunked refiner enabled; uses the same shape, fixture, seed, and camera action for comparison.
lingbot-video-moerobbyant/lingbot-video-moe-30b-a3bNoOne-GPU eager baseline using the CI structured-JSON caption, 384x640, 17 frames, 12 steps, and text-encoder CPU offload.
lingbot-worldrobbyant/lingbot-world-fast-diffusersNoOne-H200 offline single-chunk profile for the registered causal DMD path: 832x480x9, four steps, guidance 1.0, the shared image fixture, and forward-camera actions for all nine frames. Keep stateful websocket latency as a separate metric.
lingbot-world-v2robbyant/lingbot-world-v2-14b-causal-fast-diffusersNoMatching controlled single-chunk profile for the separately registered v2 checkpoint. The fixed shape, action program, and schedule make v1/v2 hotspot comparisons reproducible without presenting one-chunk e2e as stateful realtime latency.
fastwan21-t2v-1.3bFastVideo/FastWan2.1-T2V-1.3B-DiffusersNoOne-GPU 832x480, 61-frame, 3-step DMD baseline. The preset pins manual mode with a resident DiT so lossless/extra-high/high comparisons do not measure an offload-policy change.
wan21-t2v-1.3bWan-AI/Wan2.1-T2V-1.3B-DiffusersNoRegistered one-GPU 832x480, 81-frame Wan2.1 baseline at 50 steps and guidance 3.0. Keep it separate from FastWan and TurboWan because the longer schedule changes the end-to-end weight of VAE optimizations.
wan21-t2v-14bWan-AI/Wan2.1-T2V-14B-DiffusersNoCookbook-aligned four-GPU CFG/Ulysses baseline at 832x480, 81 frames, 50 steps, and guidance 5.0. Text encoding stays CPU-offloaded as in the documented deployment command.
wan21-i2v-14b-480pWan-AI/Wan2.1-I2V-14B-480P-DiffusersNoFour-GPU CFG/Ulysses image-conditioned baseline at 832x480, 81 frames, 50 steps, and guidance 5.0. Uses the shared cat fixture and its motion prompt.
wan21-i2v-14b-720pWan-AI/Wan2.1-I2V-14B-720P-DiffusersNoFour-GPU CFG/Ulysses image-conditioned baseline at 1280x720, 81 frames, 50 steps, and guidance 5.0. Keep it separate from 480P because it is a distinct checkpoint and attention shape.
wan21-fun-inp-1.3bweizhou03/Wan2.1-Fun-1.3B-InP-DiffusersNoRegistered one-GPU Wan2.1 Fun image-conditioned path at 832x480, 81 frames, 50 steps, and guidance 6.0. Uses the shared cat fixture and motion prompt.
krea2-turbokrea/Krea-2-TurboNoRecent T2I checkpoint at 1024x1024, 8 steps, guidance 1.0.
krea2-rawkrea/Krea-2-RawNoRecent T2I checkpoint at 1024x1024, 50 steps, guidance 4.5; keep separate from Turbo because CFG and the longer schedule change the hotspot mix.
ideogram4-fastfal/ideogram-v4-fastNoRecent distilled T2I checkpoint at 1024x1024; the registered sampling class owns its step and guidance defaults.
ideogram4-instantfal/ideogram-v4-instantNoRecent distilled T2I checkpoint at 1024x1024; the registered sampling class owns its step and guidance defaults.
longlive2-t2vRabinovich/LongLive-2.0-5B-DiffusersNoCI-aligned 832x480, 61-frame causal DMD T2V baseline at 4 steps and guidance 1.0.
longlive2-i2vRabinovich/LongLive-2.0-5B-DiffusersNoCI-aligned 960x928, 61-frame causal DMD I2V baseline using the cat image.
fast-hunyuanFastVideo/FastHunyuan-diffusersNoValidated one-H200 832x480, 61-frame FastHunyuan baseline using its registered 6-step schedule.
turbowan21-t2v-1.3bIPostYellow/TurboWan2.1-T2V-1.3B-DiffusersNoRegistered one-GPU TurboWan path at 832x480, 81 frames, and 4 steps.
turbowan21-t2v-14b-480pIPostYellow/TurboWan2.1-T2V-14B-DiffusersNoOne-H200 TurboWan 14B path at 832x480, 81 frames, and its 4-step DMD schedule.
turbowan21-t2v-14b-720pIPostYellow/TurboWan2.1-T2V-14B-720P-DiffusersNoOne-H200 high-resolution TurboWan 14B path at 1280x720, 81 frames, and its 4-step DMD schedule. Keep it separate because it is a distinct checkpoint.
turbowan22-i2v-a14bIPostYellow/TurboWan2.2-I2V-A14B-DiffusersNoFour-GPU CFG/Ulysses image-conditioned baseline at 1280x720, 81 frames, and its 4-step DMD schedule. Uses the shared cat fixture and keeps both high- and low-noise guidance at 3.5.
helios-midBestWishYsh/Helios-MidNoCI-sized 640x384, 33-frame pyramid-SR baseline using Helios-Mid's 20-step schedule.
helios-distilledBestWishYsh/Helios-DistilledNoCI-sized 640x384, 33-frame DMD baseline at 10 steps and guidance 1.0.
joy-echojdopensource/JoyAI-EchoNoCI-aligned two-GPU Ulysses baseline at 640x384, 33 frames, 8 steps, with the cross-request memory bank disabled for isolated single-request timing.
cosmos3-edge-t2invidia/Cosmos3-EdgeNoOne-GPU eager T2I baseline at Edge's native 640x640 shape, 35 steps, guidance 7.0.
cosmos3-edge-t2vnvidia/Cosmos3-EdgeNoOne-GPU eager T2V baseline at Edge's native 832x480 video shape, 81 frames, 35 steps, and guidance 5.0.
cosmos3-edge-i2vnvidia/Cosmos3-EdgeNoMatching one-GPU I2V baseline with the shared cat fixture; keep it separate because image conditioning adds the VAE encode and latent-mask paths.
cosmos3-super-i2vnvidia/Cosmos3-Super-Image2VideoNoRegistered specialized I2V checkpoint with the shared cat fixture; 1280x720, 81 frames, 35 steps, guidance 6.0, flow shift 10.0, seed 42, 2 GPUs, TP size 2, and guardrails disabled for benchmark isolation.
cosmos3-super-t2i-distillednvidia/Cosmos3-Super-Text2Image-4StepNoFour-GPU eager distilled T2I baseline. The checkpoint owns its fixed sigma schedule; the preset does not override the step count.
ltx25Lightricks/LTX-2.5-DiffusersNoOne-stage distilled eager baseline at 960x544, 121 frames, 8 steps, guidance 1.0.
ltx25-diffusion-decoderLightricks/LTX-2.5-DiffusersNoSame fixed DiT workload with --use-diffusion-decoder; attribute decoder time separately and confirm NATTEN na3d is active.
ltx2Lightricks/LTX-2NoCurrent-source two-stage LTX-2 preset with 2 GPUs, CFG parallel, 768x512, 121 frames
qwen-imageQwen/Qwen-ImageNoCurrent-source extra covering the base Qwen-Image native path, separate from the nightly Qwen-Image-2512 case
qwen-edit-2509Qwen/Qwen-Image-Edit-2509NoCurrent-source extra for the pre-2511 edit-plus path; uses the cat image, 1024x1024
zimage-baseTongyi-MAI/Z-ImageNoCurrent-source extra for non-turbo Z-Image; keep it separate from zimage / Z-Image-Turbo
flux2-kleinblack-forest-labs/FLUX.2-klein-4BNoCurrent-source extra for the distilled FLUX.2 Klein path; gated repo, 1024x1024, DiT layerwise offload disabled
flux2-klein-baseblack-forest-labs/FLUX.2-klein-base-4BNoCurrent-source extra for the undistilled FLUX.2 Klein Base path; gated repo, 1024x1024, DiT layerwise offload disabled
cosmos3-nano-t2invidia/Cosmos3-NanoNoCurrent-source extra for the single-frame Cosmos3 image path; sets SGLANG_DISABLE_COSMOS3_GUARDRAILS=1 in the helper environment
cosmos3-nano-t2vnvidia/Cosmos3-NanoNoCurrent-source extra for a short Cosmos3 video path; sets SGLANG_DISABLE_COSMOS3_GUARDRAILS=1 in the helper environment
ernie-image-turbobaidu/ERNIE-Image-TurboNoCurrent-source extra for ERNIE-Image Turbo
glm-imagezai-org/GLM-ImageNoCurrent-source extra for GLM-Image
sana-1.5-1.6bEfficient-Large-Model/SANA1.5_1.6B_1024px_diffusersNoCurrent-source extra for a SANA native image path
fastwan22-ti2v-5bFastVideo/FastWan2.2-TI2V-5B-FullAttn-DiffusersNoCurrent-source extra matching the FastWan2.2 TI2V registered path
wan22-t2v-nvfp4nvidia/Wan2.2-T2V-A14B-Diffusers-NVFP4NoBlackwell-only one-GPU ModelOpt NVFP4 T2V baseline at 832x480 and 81 frames. Manual mode keeps the DiT resident so the trace measures FP4 kernels instead of layerwise transfer.
ltx23-hq-two-stageLightricks/LTX-2.3NoCurrent-source extra for LTX2TwoStageHQPipeline with --ltx2-two-stage-device-mode=original; high-resolution and VRAM-heavy
ltx23-one-stageLightricks/LTX-2.3NoSkill-only extra preset for the native LTX-2.3 one-stage baseline; 2 GPUs, 768x512, 121 frames, fps 24, 30 steps, guidance 3.0, seed 1234
ltx23-two-stageLightricks/LTX-2.3NoSkill-only high-resolution stress preset for the native LTX-2.3 two-stage path; uses LTX2TwoStagePipeline, 2 GPUs, 1536x1024, 121 frames, fps 24, 30 steps, guidance 3.0, seed 1234
ltx23-two-stage-cfg-parallelLightricks/LTX-2.3NoSkill-only high-resolution CFG-parallel stress preset matching ltx23-two-stage plus --cfg-parallel-size 2
hunyuanvideohunyuanvideo-community/HunyuanVideoNoSkill-only native T2V preset at a model-supported 960x544 resolution, 65 requested frames, and 30 steps. Sequence-parallel runs may increase the frame count to satisfy their topology; record the resolved shape from the runtime log.
mova-360pOpenMOSS-Team/MOVA-360pNoTwo-GPU Ulysses I2VA baseline at 640x352 and 193 frames. Uses the upstream single-person fixture and a two-step profiling schedule.
mova-720pOpenMOSS-Team/MOVA-720pNoFour-GPU Ulysses I2VA baseline at 1280x720 and 193 frames. Uses the same upstream single-person fixture and two-step profiling schedule.
heliosBestWishYsh/Helios-BaseNoSkill-only extra preset
joyai-editjdopensource/JoyAI-Image-Edit-DiffusersNoSkill-only JoyAI image-edit preset; uses the cat image, 1024x1024, 40 steps, guidance 4.0, 2-GPU CFG parallel
firered-edit-1.0FireRedTeam/FireRed-Image-Edit-1.0NoSkill-only FireRed 1.0 image-edit preset; QwenImageEditPlus native path; uses 2-GPU CFG parallel
firered-edit-1.1FireRedTeam/FireRed-Image-Edit-1.1NoSkill-only FireRed 1.1 image-edit preset; QwenImageEditPlus native path; uses 2-GPU CFG parallel
hunyuan3d-shapetencent/Hunyuan3D-2NoSkill-only Hunyuan3D shape-generation preset; primary metric is Hunyuan3DShapeDenoisingStage

Pi0.5 is registered as an action-policy pipeline, not an image/video sglang generate pipeline, so it must not be inserted into this preset table or timed with visual-output hashes. Use its checked-in real-model lane instead:

SGLANG_RUN_PI05_E2E=1 \
SGLANG_PI05_E2E_NUM_GPUS=1 \
SGLANG_PI05_E2E_PERF_DUMP=/path/to/pi05-perf.json \
PYTHONPATH=python python3 -m pytest -s \
  python/sglang/multimodal_gen/test/single_test_file/test_pi05_e2e.py

The action lane uses three deterministic 224x224 camera inputs, deterministic noise, two denoise steps by default, repeatability/prefix-cache checks, and a three-request median. Treat action_denoise_ms as its primary metric. Isolate and clean its model cache with the same task-owned-cache discipline as visual models; BCG/quality comparisons are not applicable to this API.

For Wan2.2 video models, remember the difference between nightly alignment and best latency tuning:

  • the nightly-aligned 4-GPU commands intentionally keep --enable-cfg-parallel --ulysses-degree=2 so CFG and ring behavior stay covered
  • do not assume that is the fastest topology
  • for pure latency tuning, benchmark pure Ulysses too, for example --ulysses-degree=4 --ring-degree=1 on 4 GPUs, and on 8 GPUs compare pure --ulysses-degree=8 against --enable-cfg-parallel --ulysses-degree=4

For MiniMax-H3, keep the native contract intact:

  • use the root model ID and select fl2va or ref2va with --model-variant; do not point at a checkpoint subdirectory
  • use eager BF16/FP32 for consistency ground truth; current H3 torch.compile changes numerical output
  • keep BCG off in the validated recipe. The support gate alone is not enough: prompt-dependent packed-sequence host boundaries can differ between warmup and serving and cause a signature miss. Any experimental fix must prove real segment replay, byte-identical media, and an e2e win without excessive graph memory
  • use Ulysses, not Ring, for H3's packed multi-segment attention; CFG parallel is invalid because the released pipeline has one denoising branch
  • keep the released overlapping tiled video-VAE decode. H3 rejects spatial, spatial_shard, and patch decode modes after output mismatches

Manual command example: MiniMax-H3 T2VA

Create ${CONFIG_DIR}/minimax-h3-t2va.json with the model-specific request fields below. The generic width, height, and frame flags are intentionally absent because H3 resolves all three from target:

{
  "task": "t2va",
  "conditions": [],
  "target": {
    "short_edge": 768,
    "aspect_ratio": "16:9",
    "duration_seconds": 5.0
  },
  "num_inference_steps": 50,
  "flow_shift": 12.0,
  "audio_flow_shift": 3.0
}

Then run the same lossless 4-GPU H100 topology and 5-second shape used by the source-tracked preset:

sglang generate \
  --backend=sglang \
  --model-path=MiniMaxAI/MiniMax-H3 \
  --model-variant=fl2va \
  --config="${CONFIG_DIR}/minimax-h3-t2va.json" \
  --prompt="At night, while their owner sleeps in a bedroom, three cats march in loudly playing tiny brass instruments, then abruptly file out." \
  --seed=1101 --num-gpus=4 --tp-size=2 --ulysses-degree=2 \
  --performance-mode=speed --enable-torch-compile=false \
  --save-output --warmup-mode request \
  --perf-dump-path="${BENCH_DIR}/minimax-h3-t2va-baseline.json"

The benchmark helper creates this config automatically. For ModelScope, set SGLANG_USE_MODELSCOPE=true, replace the root model ID with MiniMax/MiniMax-H3, and keep the selected variant unchanged. When --output-dir is provided, the helper places the generated config under that directory's generated_configs/ subdirectory so the run is self-contained.

For a serving benchmark, use the driver maintained by the H3 cookbook after launching the corresponding sglang serve command:

python3 -m sglang.multimodal_gen.benchmarks.bench_serving \
  --host 127.0.0.1 --port 30010 \
  --model MiniMaxAI/MiniMax-H3 \
  --dataset vbench --task text-to-video \
  --num-prompts 1 --max-concurrency 1 \
  --warmup-requests 1 --warmup-inference-steps 50 \
  --extra-body '{"task":"t2va","conditions":[],"target":{"short_edge":768,"aspect_ratio":"16:9","duration_seconds":5.0},"seconds":5,"flow_shift":12.0,"audio_flow_shift":3.0}'

H3 correctness is joint video/audio correctness. Use eager BF16/FP32 as the only ground truth and keep prompt, seed, target, step count, shifts, partition, and topology fixed. For a lossless kernel/runtime change:

  • compare decoded frames after frame-count and timestamp alignment; report at least frame-wise PSNR/SSIM plus the worst frame, not only an average
  • extract the 32 kHz stereo audio stream and compare channel order, sample count, waveform error, and a time-aligned log-mel or spectral metric
  • verify the MP4 contract remains H.264 video at 24 fps plus one AAC stereo audio stream
  • run the relevant kernel/unit exactness test when replacing an existing H3 BF16 fast path. Do not hide a failed exact test behind a permissive end-to-end perceptual threshold

There is no source-wide universal perceptual threshold for arbitrary H3 changes. Record the acceptance bounds before optimization and tighten them for changes that claim to preserve eager math. Approximate Cache-DiT or FP8 runs must be labeled separately and validated for both output modalities.

Manual command example: LTX-2 Two-Stage

sglang generate \
  --model-path=Lightricks/LTX-2 \
  --pipeline-class-name=LTX2TwoStagePipeline \
  --prompt="A cat and a dog baking a cake together in a kitchen." \
  --width=768 --height=512 \
  --num-frames=121 \
  --seed=42 --num-gpus=2 --enable-cfg-parallel \
  --save-output --enable-torch-compile --warmup-mode request

LTX2TwoStagePipeline is a native path. The spatial upsampler and distilled LoRA are auto-resolved from the same model snapshot unless you override them.

Manual command example: LTX-2.3 TI2V Two-Stage

sglang generate \
  --model-path=Lightricks/LTX-2.3 \
  --pipeline-class-name=LTX2TwoStagePipeline \
  --prompt="The cat starts walking slowly towards the camera." \
  --image-path="${ASSET_DIR}/cat.png" \
  --width=768 --height=512 \
  --num-frames=121 \
  --seed=42 --num-gpus=2 --cfg-parallel-size=2 \
  --save-output --enable-torch-compile --warmup-mode request

This matches the nightly comparison case ltx2.3_twostage_ti2v_2gpus.

Manual command example: LTX-2.3 One-Stage

sglang generate \
  --model-path=Lightricks/LTX-2.3 \
  --prompt="A beautiful sunset over the ocean" \
  --negative-prompt="shaky, glitchy, low quality, worst quality, deformed, distorted, disfigured, motion smear, motion artifacts, fused fingers, bad anatomy, weird hand, ugly, transition, static." \
  --width=768 --height=512 \
  --num-frames=121 --fps=24 \
  --num-inference-steps=30 --guidance-scale=3.0 \
  --seed=1234 --num-gpus=2 \
  --save-output --enable-torch-compile --warmup-mode request

Use this when you want the native LTX2Pipeline baseline for LTX-2.3 at the validated one-stage resolution.

Manual command example: LTX-2.3 Two-Stage High-Resolution Stress

sglang generate \
  --model-path=Lightricks/LTX-2.3 \
  --pipeline-class-name=LTX2TwoStagePipeline \
  --prompt="A beautiful sunset over the ocean" \
  --negative-prompt="shaky, glitchy, low quality, worst quality, deformed, distorted, disfigured, motion smear, motion artifacts, fused fingers, bad anatomy, weird hand, ugly, transition, static." \
  --width=1536 --height=1024 \
  --num-frames=121 --fps=24 \
  --num-inference-steps=30 --guidance-scale=3.0 \
  --seed=1234 --num-gpus=2 \
  --save-output --enable-torch-compile --warmup-mode request

This matches the skill-only ltx23-two-stage preset. Use it as a high-resolution stress target, not as a nightly comparison case.

Manual command example: JoyAI Image Edit

sglang generate \
  --backend=sglang \
  --model-path=jdopensource/JoyAI-Image-Edit-Diffusers \
  --prompt="Make the cat wear a red hat" \
  --image-path="${ASSET_DIR}/cat.png" \
  --width=1024 --height=1024 \
  --num-inference-steps=40 --guidance-scale=4.0 \
  --num-gpus=2 --enable-cfg-parallel --ulysses-degree=1 \
  --dit-layerwise-offload false --dit-cpu-offload false \
  --save-output --enable-torch-compile --warmup-mode request

Manual command example: FireRed Image Edit

sglang generate \
  --backend=sglang \
  --model-path=FireRedTeam/FireRed-Image-Edit-1.1 \
  --prompt="Make the cat wear a red hat" \
  --image-path="${ASSET_DIR}/cat.png" \
  --width=1024 --height=1024 \
  --num-inference-steps=40 --guidance-scale=4.0 \
  --num-gpus=2 --enable-cfg-parallel --ulysses-degree=1 \
  --dit-layerwise-offload false --dit-cpu-offload false \
  --save-output --enable-torch-compile --warmup-mode request

Use FireRedTeam/FireRed-Image-Edit-1.0 in the same command when comparing the 1.0 checkpoint. Both FireRed presets use the native QwenImageEditPlusPipeline path. On H100, 2-GPU CFG parallel reduced 40-step denoise latency versus the otherwise matching 2-GPU Ulysses command: FireRed 1.0 from 13419.15 ms to 10955.90 ms, and FireRed 1.1 from 13414.72 ms to 10934.21 ms.

Manual command example: Hunyuan3D Shape

OUTPUT_DIR=$(python3 "$ENV_PY" print-output-dir --kind benchmarks --mkdir)
CONFIG_DIR="${OUTPUT_DIR}/generated_configs"
mkdir -p "${CONFIG_DIR}"
printf '{"paint_enable": false}\n' > "${CONFIG_DIR}/hunyuan3d-shape.json"

sglang generate \
  --backend=sglang \
  --model-path=tencent/Hunyuan3D-2 \
  --prompt="generate 3d mesh" \
  --image-path="${ASSET_DIR}/cat.png" \
  --config="${CONFIG_DIR}/hunyuan3d-shape.json" \
  --num-inference-steps=50 --guidance-scale=5.0 \
  --dit-layerwise-offload false --dit-cpu-offload false \
  --save-output --enable-torch-compile --warmup-mode request

For Hunyuan3D, compare the denoise stage separately from mesh export and paint stages. The benchmark helper reports Hunyuan3DShapeDenoisingStage as the primary denoise metric.

Manual command example: Wan2.2-I2V-A14B 720P

# Select four idle GPUs first:
# export CUDA_VISIBLE_DEVICES=$(python3 "$ENV_PY" print-idle-gpus --count 4)
sglang generate \
  --model-path=Wan-AI/Wan2.2-I2V-A14B-Diffusers \
  --prompt="The cat starts walking slowly towards the camera." \
  --image-path="${ASSET_DIR}/cat.png" \
  --width=1280 --height=720 --num-frames=81 \
  --seed=42 --save-output \
  --num-gpus=4 --enable-cfg-parallel --ulysses-degree=2 \
  --text-encoder-cpu-offload --pin-cpu-memory \
  --warmup-mode request --enable-torch-compile

Wan2.2-I2V-A14B uses the 720p max-area config by default, and explicit --width/--height overrides control the target area while preserving the reference-image aspect ratio.

Perf Dump Workflow

For every benchmark run, write a perf dump JSON:

sglang generate ... --warmup-mode request --perf-dump-path "${BENCH_DIR}/<result>.json"

Before/after comparison:

python3 python/sglang/multimodal_gen/benchmarks/compare_perf.py \
  "${BENCH_DIR}/baseline.json" \
  "${BENCH_DIR}/new.json"

Always keep:

  • denoise latency
  • end-to-end latency
  • peak GPU memory
  • exact command line, model shape, dtype, request quality, GPU topology, and whether synchronized stage profiling was enabled

Never keep a perf dump produced after a diffusers-backend fallback. Also reject a zero-exit run if either the requested perf dump or generated media is absent: some generation failures are reported through the response payload without a nonzero process exit.

For quality=lossless, compare saved artifact hashes and require byte equality for a claimed lossless fast path or BCG change. For quality=extra-high and quality=high, keep the lossless artifact as ground truth and report both aggregate and worst-frame SSIM/PSNR. Repository defaults are SSIM 0.95 / PSNR 28 dB for images and SSIM 0.92 / PSNR 24 dB for videos; checked-in model/hardware consistency metadata may override them. Always inspect the image or a start/middle/end video contact sheet in addition to scalar metrics.

Use denoise timing to locate the opportunity, but gate a performance PR on repeated saved-request end-to-end time. The project threshold for this sweep is at least 1.5% mean e2e improvement on same-GPU ABBA runs. Attach one representative baseline/candidate profile plus before/after images or videos to the PR description.

Stage durations are host wall times around asynchronous GPU launches unless SGLANG_DIFFUSION_SYNC_STAGE_PROFILING=1. Without the sync, queued denoise work can leak into the next blocking stage and inflate DecodingStage by 2-3x. Use synchronized dumps for denoise/decode attribution and keep the setting identical in every before/after pair.

torch.profiler Workflow

1. Establish the baseline

PYTHONPATH=python python3 "$BENCH_PY" \
  --model flux \
  --label baseline \
  --output-dir "${BENCH_DIR}"

Keep model shape, seed, and GPU topology fixed for every comparison. Save one reference image or video before changing code. If the active task requires torch.compile off, add --no-torch-compile here too.

MiniMax-H3 always requires eager mode for consistency ground truth. The minimax-h3-t2va helper preset enforces it, and manual H3 profile commands must pass --enable-torch-compile=false.

For H3, one --profile-all-stages trace separates text/condition encoding, MiniMaxH3DenoisingStage, and the aggregate MiniMaxH3DecodingStage. The decoding stage contains both video decode and rank-0 audio decode. If decoding is hot, add temporary record_function or NVTX scopes around video_vae.decode_base and _decode_audio in the H3 decoding stage, then re-run the same all-stage profile. Do not attribute aggregate decoding time to one VAE without those inner scopes.

2. Capture a representative trace

By default SGLang profiles the denoising stage. The default sampling window is 5 profiled timesteps after warmup.

SGLANG_DIFFUSION_TORCH_PROFILER_DIR="${PROFILE_DIR}/torch" \
sglang generate \
  --model-path=black-forest-labs/FLUX.1-dev \
  --prompt="A futuristic cyberpunk city at night" \
  --width=1024 --height=1024 --num-inference-steps=50 \
  --seed=42 --enable-torch-compile --warmup-mode request \
  --profile

Use --profile-all-stages only when you really need text encoder, VAE, or other non-denoise stages too.

The generated trace path is printed in the console and also lands under SGLANG_DIFFUSION_TORCH_PROFILER_DIR. The diffusion profiler falls back to SGLANG_TORCH_PROFILER_DIR and then ./logs when the diffusion-specific env var is unset. Open the trace in Perfetto if you want a timeline view:

3. Rank the hot CUDA kernels

Use this parser for a quick top-k table without opening a browser:

import collections
import glob
import gzip
import json
import os

log_dir = (
    os.environ.get("SGLANG_DIFFUSION_TORCH_PROFILER_DIR")
    or os.environ.get("SGLANG_TORCH_PROFILER_DIR")
    or "./logs"
)
trace_path = sorted(
    glob.glob(f"{log_dir}/*.trace.json.gz"),
    key=os.path.getmtime,
    reverse=True,
)[0]

with gzip.open(trace_path, "rb") as f:
    data = json.loads(f.read())

cuda_ops = collections.defaultdict(lambda: {"total_us": 0, "count": 0})
for event in data.get("traceEvents", []):
    if event.get("cat") in ("kernel", "gpu_memcpy") and "dur" in event:
        cuda_ops[event.get("name", "unknown")]["total_us"] += event["dur"]
        cuda_ops[event.get("name", "unknown")]["count"] += 1

print(f"{'Kernel':<90} {'Total(ms)':>10} {'Count':>6}")
for name, stat in sorted(cuda_ops.items(), key=lambda item: -item[1]["total_us"])[:30]:
    print(f"{name:<90} {stat['total_us'] / 1000:>10.3f} {stat['count']:>6}")

If you need better attribution, add record_function(...) scopes around DiT attention, norm, modulation, MLP, or communication boundaries and re-run.

4. Classify the hotspot with existing-fast-paths.md

Do not jump from a hot kernel straight into new code. First classify it against the known mainline families.

What the trace showsFirst interpretation
fused_inplace_qknorm_rope missing, but separate qk norm plus rope show upCheck whether the fused diffusion QK norm + RoPE path should have engaged
to_q -> to_k -> to_v on NVFP4 or Nunchaku FLUX-family checkpointsTreat as a packed-QKV fast-path miss or checkpoint-format mismatch
rmsnorm_scale or rmsnorm_tanh_residual missing on Z-ImageCheck the bf16-native Triton eligibility guards before proposing a new fusion
FLUX.1, GLM-Image, or SANA shows separate LayerNorm plus adaLN elementwise kernelsCheck the bit-exact modulate_scale_shift and fused_layernorm_modulate guards/self-test before proposing another norm fusion
quality=extra-high or quality=high shows the same FLUX/GLM DiT or FLUX-family/Wan VAE chain as losslessCheck whether the request-scoped quality gate mounted and whether every site passed its all-or-nothing compatibility checks
LTX-2 split RoPE appears as a long PyTorch elementwise chainCheck the apply_ltx2_split_rotary_emb Triton path and its shape guards
Wan decode is dominated by causal cat + pad + contiguous, feature-cache copies, or repeat_interleave + permute + addCheck the bit-exact Wan causal-cache and DupUp3D data-movement kernels before writing a new decoder kernel
masked attention spends time packing/unpacking Q/K/VCheck whether fused varlen USP pack/scatter should have engaged
all_to_all, ring attention, or async A2A dominateClassify against Ulysses, USP, or turbo-layer overlap first
Fixed-resolution image/video traces show many small launch gapsCheck supported breakable CUDA graph capture, declared warmup resolutions, and text buckets before adding a new graph mechanism
H3 shows separate indexed gather + scale/shift, QK norm + RoPE, or three Q/K/V Ulysses relayoutsCheck H3's indexed-modulation, fused QK-norm+RoPE, packed Ulysses-QKV, and USP relayout guards before writing a new kernel
H3 TP traces show one AdaLN collective per blockCheck the batched TP AdaLN projection/all-gather path in minimax_h3.py before attempting communication overlap
split fc1 -> gelu -> quant -> fc2.lora_down on Nunchaku FLUXTreat as a missing fused GELU MLP path
attention kernels dominateConfirm backend, topology, and shape guards before proposing a new kernel

If the hot path is already covered by a mainline optimization family, fix the enablement, shape guard, backend choice, or checkpoint mapping first.

5. Hand off only real kernel work

Only after the hotspot survives the fast-path checklist:

  1. save a baseline perf dump
  2. save a representative torch.profiler trace
  3. note the exact model, shape, dtype, and GPU topology
  4. hand the work to the appropriate kernel, Nsight, or framework-specific optimization workflow

This skill intentionally stops here. It tells you whether you are looking at:

  • a missing existing optimization
  • a configuration or backend problem
  • or a real kernel opportunity worth handing off

Minimal Merge Checklist

  • fixed-shape baseline perf dump saved
  • fixed-shape new perf dump saved
  • quality/BCG applicability matrix attempted on one GPU set
  • BCG rows show capture and no disable/failure/signature-miss/late-quality-fusion marker
  • request shape, seed, steps, guidance, topology, residency, and synchronized stage profiling match
  • compare_perf.py table generated
  • one representative torch.profiler trace saved
  • hotspot classified against existing-fast-paths.md
  • lossless artifact hash is exact; extra-high/high aggregate and worst-frame SSIM/PSNR pass the checked-in threshold
  • reference image or start/middle/end video contact sheet checked visually
  • any PR claim has repeated saved-request e2e improvement >= 1.5%
  • task-owned checkpoint cache cleaned and ledger shows zero residual weight files
  • any remaining kernel work handed off with perf/profile evidence attached
Referenced from SKILL.md