Yo, so if you use multi-agent generation for Chinese architecture images, won't the aesthetic get all over the place?

Using multi-agent systems to generate Chinese architecture, what I’m worried about is exactly that the aesthetic will fall apart. You split the image across different nodes, each handling its own piece—one for dougong brackets, one for roof ridges, one for courtyard layout.

Individually, each part looks pretty detailed, but when you stitch them together, it’s easy to lose the overall vibe. Chinese architecture is all about that unified restraint and negative space. Once the agents start doing their own thing, it can easily turn into just piling on details.

From what I’ve tried, you need to lock down the tone with a strong style anchor upfront, and then have every subsequent task branch out from that anchor—otherwise it falls apart. How do you guys keep consistency in systems like this?

Yeah, breaking down nodes does tend to fall apart easily. I tried it once with a dougong bracket set—the ridge was a total mess.

Yeah, the strong constraint anchor approach makes sense. You can’t really piece together that kind of vibe or spirit—it’s not something you can just assemble.

The hardest thing to pass down to downstream nodes is leaving blank space—the machine just fills it all up.

We start by generating one overall mood image as a reference, and every node has to stick close to it.