For one model on one GPU, every optimization recipe is measured against the baseline. The frontier is the recipe that solves
for the quality loss ε you choose — LPIPS for image and video, word-error-rate increase for speech.
seconds per image · LPIPS vs baseline
| LPIPS | Latency | Speedup | Recipe |
|---|---|---|---|
| .011 | 7.1s | 1.9× | DPCache K=20 |
| .080 | 4.5s | 3.0× | DPCache K=12 |
| .129 | 2.0s | 6.7× | FP8 W8A8 per-channel + SageAttention2 + fused text-encoder and VAE kernels + DPCache K=12 |
| LPIPS | Latency | Speedup | Recipe |
|---|---|---|---|
| .120 | 3.1s | 5.8× | FP8 W8A8 (double-stream blocks BF16) + SageAttention2 + fused concat + FP8 quantize + DPCache K=16 |
| .193 | 2.6s | 6.9× | FP8 W8A8 + SageAttention2 + fused concat + FP8 quantize + DPCache K=16 |
One pull request, one recipe config; maintainers measure and publish it. GitHub ↗