Drag the dashed line (or focus its grip and use the arrow keys) to set the quality limit. The fastest recipe within it is selected; recipes above it are greyed out.
| Latency | LPIPS | Speedup | Recipe | Engine | Verified |
|---|---|---|---|---|---|
| 2.6s | .193 | 6.9× | FP8 W8A8 + SageAttention2 + fused concat + FP8 quantize + DPCache K=16 | sglang-diffusion 5ead00ae6 | ✓ |
| 3.1s | .120 | 5.8× | FP8 W8A8 (double-stream blocks BF16) + SageAttention2 + fused concat + FP8 quantize + DPCache K=16 | sglang-diffusion 5ead00ae6 | ✓ |
| 8.0s | .188 | 2.2× | FP8 W8A8 + SageAttention2 | sglang-diffusion 5ead00ae6 | ✓ |
| 17.8s | — | 1.0× | Baseline | sglang-diffusion 5ead00ae6 | ✓ |
Found a better optimization combination? Submit a reproducible recipe through GitHub.
Submit via GitHub ↗Submissions are evaluated under the same model, hardware, workload, and quality protocol. The submission unit is a complete recipe.