Inference Frontier

FLUX.2 [klein] 4B×

Resolution
1024×1024
Batch
1
Steps
50
Precision
BF16
Guidance
4
Attention
torch_sdpa
Offload
none
Protocol
v0.1
Prompts
20 public
Updated
2026-10-01
0.020.040.060.080.100.120.140.160.180.200024681012141618E2E latency (s)Quality loss (LPIPS)↙ betterBaselineLPIPS ≤ .050

Drag the dashed line (or focus its grip and use the arrow keys) to set the quality limit. The fastest recipe within it is selected; recipes above it are greyed out.

Fastest within LPIPS ≤ .050
Baseline17.8s · 1.0× faster · LPIPS —
LatencyLPIPSSpeedupRecipeEngineVerified
2.6s.1936.9×FP8 W8A8 + SageAttention2 + fused concat + FP8 quantize + DPCache K=16sglang-diffusion 5ead00ae6✓
3.1s.1205.8×FP8 W8A8 (double-stream blocks BF16) + SageAttention2 + fused concat + FP8 quantize + DPCache K=16sglang-diffusion 5ead00ae6✓
8.0s.1882.2×FP8 W8A8 + SageAttention2sglang-diffusion 5ead00ae6✓
17.8s—1.0×Baselinesglang-diffusion 5ead00ae6✓

Add a recipe

Found a better optimization combination? Submit a reproducible recipe through GitHub.

Submit via GitHub ↗

Submissions are evaluated under the same model, hardware, workload, and quality protocol. The submission unit is a complete recipe.