Just helped the team put together a training rig and spent two weeks messing around with Kohya_ss. Here’s a quick rundown on what each VRAM tier can handle.
For SD1.5 LoRA, it’s pretty forgiving—6G can barely scrape by, but 8G is where it gets comfy. Pretty much any modern card will do. When you move to SDXL, things change: 12G is the bare minimum, and 16G gives you enough headroom to crank up resolution and batch size. Flux.1 is a real VRAM hog—16G is the starting point, and you’ll need gradient checkpointing just to keep it alive. 24G is where it actually feels smooth.
On my 4090, training an SDXL LoRA for 1500 steps takes about 10-15 minutes. Flux? That’s 2-3 times longer. If you’re on a tight budget, the 4060Ti 16G is the best bang for your buck. Don’t cheap out on the 8G version—OOM errors will drive you crazy.
One thing newbies always overlook: cache your latents before training. Do that network cache step, and the overall speed difference is huge. Also, don’t put your dataset on a mechanical hard drive—Kohya reads images constantly, and a slow drive will waste a ton of time over thousands of steps.