Inference Frontier

For one model on one GPU, every optimization recipe is measured against the baseline. The frontier is the recipe that solves

minimizelatency(recipe)subject toloss(recipe, baseline) ≤ ε

for the quality loss ε you choose — LPIPS for image and video, word-error-rate increase for speech.

Image models

seconds per image · LPIPS vs baseline

baseline 13.6s
LPIPSLatencySpeedupRecipe
.0117.1s1.9×DPCache K=20
.0804.5s3.0×DPCache K=12
.1292.0s6.7×FP8 W8A8 per-channel + SageAttention2 + fused text-encoder and VAE kernels + DPCache K=12
View benchmark
baseline 17.8s
LPIPSLatencySpeedupRecipe
.1203.1s5.8×FP8 W8A8 (double-stream blocks BF16) + SageAttention2 + fused concat + FP8 quantize + DPCache K=16
.1932.6s6.9×FP8 W8A8 + SageAttention2 + fused concat + FP8 quantize + DPCache K=16
View benchmark

How it's measured

  • Baseline: the engine's native BF16 run at the page's step count. Every recipe is compared to its output.
  • Quality: a loss against the baseline output for the same prompt and seed, over a fixed prompt set — LPIPS for image, LPIPS for video, ΔWER for speech.
  • Speed: latency of a single request, median after warmup — seconds per image or clip, milliseconds to first audio for speech.
  • Verified: re-measured by the maintainers on a private held-out prompt set.

Submit a recipe

One pull request, one recipe config; maintainers measure and publish it. GitHub ↗