A lot of people think SD1.5 has hit its ceiling, but I don’t see it that way—there’s still plenty to squeeze out of the 1.5 base. I recently went through the full pipeline following FollowFox’s Cosmopolitan approach, so here’s my log.
For the dataset, I used synthetic images generated by Midjourney, picking a subset from the original table that was over 7GB. Ended up with over 11,000 image-text pairs, setting aside 15% as a validation set. The trainer was still EveryDream2, but this time I deliberately trained at 768 resolution alongside 512, not just 512.
The 768 run kept randomly crashing on Runpod, wasted a ton of time retrying. Eventually, for 512 I used D-Adaptation, and for 768 I switched back to the reliable AdamW. The stopping criterion was when validation loss started trending upward consistently—around 70 epochs for 512, 60 for 768.
After that came model merging—mixing vodka bases with each other, then blending in third-party models. Old faces like realistic vision and dreamshaper were mixed at a 60/20/20 ratio. The final pick was that 768 hybrid version, which is more flexible than the old BloodyMary. If you’re interested, it’s available on Civitai.