That Tsinghua Duan Yueqi team paper for CVPR 2026 says text-to-image needs to shift from "tweaking parameters" to "doing control."

Leiphone’s GAIR just posted a paper column about Tsinghua Duan Yueqi’s team’s work, accepted at CVPR 2026. The core idea is pretty interesting: the methodology for text-to-image needs to level up from tuning parameters to controlling the output. I think this direction has been the pain point the industry has been grappling with for the past couple of years—people are way past just “generating a nice-looking image.”

They want the stuff in the image to follow their instructions: composition, pose, lighting, individual elements all need to be dialed in precisely. A lot of the old approaches were essentially fiddling with training and sampling, which is an indirect way to influence the result. Moving toward “control” means treating controllability as a first-class citizen in the design.

I didn’t read the full paper, so I won’t pretend to interpret the specifics—I’ll wait until someone digs into it before we talk more. But I agree with the direction of this shift. Whoever can nail controllability will be the one who truly gets into the production pipeline.

Marking this, will come back for the original text later.

Controllability has always been the bottleneck for real-world use. It doesn’t matter if the generated image looks good—what really kills you is when you try to tweak one small part and the whole thing falls apart.

Got an arxiv link? Wanna see how they actually do the control.

This is exactly what we need—same character, different poses and angles, and it still stays consistent.

Top

“From tweaking parameters to controlling the output” — that’s pretty accurate. These days I spend most of my time just gacha-style re-rolling over and over.

Every year at CVPR there’s a bunch of people claiming they’re gonna upgrade the methodology, but whether it actually makes it into a real product is another story.

The problem where tweaking one part makes the whole image fall apart is a total dealbreaker. If you can’t get controllability right, no matter how pretty the output looks, you can’t actually deliver it to the client.

“Fix one part, the whole image falls apart” — too real. Most of the time spent on delivery is just locking in the parts that already look good. If the controllability was better, it’d save half the hassle.