Using multi-agent systems to generate Chinese architecture, what I’m worried about is exactly that the aesthetic will fall apart. You split the image across different nodes, each handling its own piece—one for dougong brackets, one for roof ridges, one for courtyard layout.
Individually, each part looks pretty detailed, but when you stitch them together, it’s easy to lose the overall vibe. Chinese architecture is all about that unified restraint and negative space. Once the agents start doing their own thing, it can easily turn into just piling on details.
From what I’ve tried, you need to lock down the tone with a strong style anchor upfront, and then have every subsequent task branch out from that anchor—otherwise it falls apart. How do you guys keep consistency in systems like this?