Yo, anyone know if TensorRT FP4 speeds up image gen? Will the 50-series GPUs be lit for that?

TensorRT with FP4 to speed up image generation - how much of a beast are the 50-series cards? I tested it on my rig. After enabling TensorRT, the same Flux model’s generation speed jumped significantly, and with the 50-series’ native FP4 support, VRAM usage dropped a lot too. The difference is most obvious when you’re generating in bulk.

But this isn’t free performance - TensorRT requires you to compile the model into an engine first, and if you change the resolution or batch size, you might need to recompile. That hassle adds up. Plus, not all nodes and samplers support this acceleration; for unsupported ones, you’re stuck with the regular path.

If you’re in a production scenario where you generate in bulk at fixed resolutions, compile once and enjoy the speed for a long time - totally worth it. But if you’re constantly switching models and sizes, just messing around, that speed boost might not be worth the repeated recompilation headache. Depends on how you use it.

Batch generating at fixed resolution with TensorRT is so satisfying, one compile lasts forever.

Every time you switch models or change sizes, you have to redo the whole thing. That kind of hassle is enough to scare off anyone who’s constantly experimenting.

FP4 VRAM drops noticeably, the biggest difference is in batch runs on the 50-series.

Yeah I’m the type who changes sizes every day. When you do the math, the speed boost isn’t worth having to recompile.

Oh yeah, I’m all over this one.