Qwen just updated its 3rd-gen image generation base model, two key points: complex instruction understanding and commercial-grade detail rendering, both aimed squarely at real deployment — even with lots of constraints it still breaks things down cleanly, output goes straight into production material. A few days ago they also just released a preview of their flagship language model with over 20 trillion parameters, pushing image and language together, multimodal all moving in sync. Let’s see how it handles complex prompts in batch runs.
If they really nail complex instruction understanding, anyone writing long prompts is set free.