When it comes to picking a base model, I trust anonymous blind-vote rankings way more these days, I barely look at the benchmark numbers each company posts themselves anymore.
The logic is simple: voters can’t see which company made which image, two images side by side, pure gut call, pick one. Run it all through Elo at the end, put GPT-Image, FLUX, Midjourney, DALL-E all together, and who’s stronger and who’s weaker becomes crystal clear, no need to take anyone’s word for it. Official PR images are all cherry-picked, you look at them and still don’t feel confident, and for our product selection work we’ve already started using this leaderboard as a reference.
There’s one flaw though, voting taste skews toward crowd-pleasing styles, photorealistic portraits get an unfair advantage. If you’re doing design-type work, you still need to run your own verification round after checking the leaderboard.