Inference Frontier

Qwen-Image 2.1×

Resolution
1024×1024
Batch
1
Steps
40
Precision
BF16
Guidance
1
Attention
torch_sdpa
Offload
text encoder layerwise
Protocol
v0.5
Prompts
20 public
Updated
2026-10-01
0.020.040.060.080.100.120.14002468101214E2E latency (s)Quality loss (LPIPS)↙ betterDPCache K=20BaselineLPIPS ≤ .050

Drag the dashed line (or focus its grip and use the arrow keys) to set the quality limit. The fastest recipe within it is selected; recipes above it are greyed out.

Fastest within LPIPS ≤ .050
DPCache K=207.1s · 1.9× faster · LPIPS .011
LatencyLPIPSSpeedupRecipeEngineVerified
2.0s.1296.7×FP8 W8A8 per-channel + SageAttention2 + fused text-encoder and VAE kernels + DPCache K=12sglang-diffusion 5ead00ae6✓
4.5s.0803.0×DPCache K=12sglang-diffusion 5ead00ae6✓
5.2s.0912.6×Cache-DiT stocksglang-diffusion 5ead00ae6✓
5.7s.1032.4×FP8 W8A8 per-channel + SageAttention2 + fused text-encoder and VAE kernelssglang-diffusion 5ead00ae6✓
5.8s.0982.3×FP8 W8A8 per-channel + SageAttention2sglang-diffusion 5ead00ae6✓
6.8s.0812.0×FP8 W8A8 per-channelsglang-diffusion 5ead00ae6✓
7.1s.0111.9×DPCache K=20sglang-diffusion 5ead00ae6✓
7.8s.0171.7×Cache-DiT conservativesglang-diffusion 5ead00ae6✓
12.5s.0571.1×SageAttention2sglang-diffusion 5ead00ae6✓
13.6s—1.0×Baselinesglang-diffusion 5ead00ae6✓

Add a recipe

Found a better optimization combination? Submit a reproducible recipe through GitHub.

Submit via GitHub ↗

Submissions are evaluated under the same model, hardware, workload, and quality protocol. The submission unit is a complete recipe.