Key Info
Qwen-Image-2.1 was benchmarked locally on an Apple M5 Max using MLX in bf16 (via a custom MFLUX pipeline), generating a 1600×672 image in about 26s at 10 steps and 97s at 50 steps. The run averaged roughly 1.78s per diffusion step plus ~8s of fixed overhead, with peak memory around 31GB.
Highlights
- At 1600×672 (2.39:1) with a fixed seed, 10/20/30/40/50 steps took about 26/44/61/79/97 seconds respectively.
- Per-step cost was ~1.78s, with ~8s one-time overhead for loading, prompt encoding, and VAE decode.
- Peak memory was ~31GB in bf16, giving a practical data point for running Qwen-Image-2.1 on high-end Apple Silicon setups.