Ugh, I fell back into the SD1.5 fine-tuning rabbit hole again. Let me walk through the whole training pipeline.

A lot of people think SD1.5 has hit its ceiling, but I don’t see it that way—there’s still plenty to squeeze out of the 1.5 base. I recently went through the full pipeline following FollowFox’s Cosmopolitan approach, so here’s my log.

For the dataset, I used synthetic images generated by Midjourney, picking a subset from the original table that was over 7GB. Ended up with over 11,000 image-text pairs, setting aside 15% as a validation set. The trainer was still EveryDream2, but this time I deliberately trained at 768 resolution alongside 512, not just 512.

The 768 run kept randomly crashing on Runpod, wasted a ton of time retrying. Eventually, for 512 I used D-Adaptation, and for 768 I switched back to the reliable AdamW. The stopping criterion was when validation loss started trending upward consistently—around 70 epochs for 512, 60 for 768.

After that came model merging—mixing vodka bases with each other, then blending in third-party models. Old faces like realistic vision and dreamshaper were mixed at a 60/20/20 ratio. The final pick was that 768 hybrid version, which is more flexible than the old BloodyMary. If you’re interested, it’s available on Civitai.

Yeah, training together at 768 is definitely the key. Pure 512 has way too obvious of a ceiling.

Yeah, I’ve had that random disconnection issue with Runpod too. It’s enough to make your blood pressure max out.

Mixing models is kinda voodoo but it actually works, no idea why.

Stop training as soon as the loss starts going up. That’s a way more reliable criterion than eyeballing the images.

1.5’s not dead yet, lightweight models are still where it’s at.

Bookmarked.

Runpod dropping instances in the middle of the night is brutal, lost a training run halfway through, now I save a checkpoint every few steps before I dare sleep.

2 Likes