From a blurry mess in the early days to now, text-to-image has leveled up fast these past few years

Anyone else remember what text-to-image output looked like when you first got into this? Wrong finger counts, mangled faces, messy composition — most of it was blurry, only good as a rough sketch for inspiration. If you actually had to deliver something, you still had to do it by hand.

The iteration speed these past few years has been insane, visibly so. Now a random generation can look pretty realistic, details hold up even zoomed in, and the style range is absurdly wide — realistic, anime, material rendering, it can pretty much nail all of it. My own gut feeling is the whole evaluation standard shifted a level: it used to be “can it get it right,” now it’s “can it look good.”

Video is basically following the same logic, just with the time dimension thrown in too. For people in this line of work, the excitement is real, and so is the sense of crisis.

The wrong finger count thing is way too real, early on it was all gacha-pulling

1 Like

The bar’s risen so fast, I can’t even look at my drafts from last year anymore

Still got that early batch of images on my hard drive, dug them up and honestly can’t stand to look at them :sob:

I’ve still got that early batch on my hard drive too. Back then I actually thought it was good enough to hand in :man_facepalming: