Anonymous blind-vote text-to-image rankings are way more trustworthy than official benchmarks

When it comes to picking a base model, I trust anonymous blind-vote rankings way more these days, I barely look at the benchmark numbers each company posts themselves anymore.

The logic is simple: voters can’t see which company made which image, two images side by side, pure gut call, pick one. Run it all through Elo at the end, put GPT-Image, FLUX, Midjourney, DALL-E all together, and who’s stronger and who’s weaker becomes crystal clear, no need to take anyone’s word for it. Official PR images are all cherry-picked, you look at them and still don’t feel confident, and for our product selection work we’ve already started using this leaderboard as a reference.

There’s one flaw though, voting taste skews toward crowd-pleasing styles, photorealistic portraits get an unfair advantage. If you’re doing design-type work, you still need to run your own verification round after checking the leaderboard.

Bump

1 Like

Applying the Elo system to image evaluation is kind of interesting, just worried about vote manipulation

The photorealistic-portrait-advantage point really hits it, for illustration work what we picked out is a completely different story

I’m not too worried about vote manipulation, once the sample size piles up it’s hard to push it one-sided