TensorRT with FP4 to speed up image generation - how much of a beast are the 50-series cards? I tested it on my rig. After enabling TensorRT, the same Flux model’s generation speed jumped significantly, and with the 50-series’ native FP4 support, VRAM usage dropped a lot too. The difference is most obvious when you’re generating in bulk.
But this isn’t free performance - TensorRT requires you to compile the model into an engine first, and if you change the resolution or batch size, you might need to recompile. That hassle adds up. Plus, not all nodes and samplers support this acceleration; for unsupported ones, you’re stuck with the regular path.
If you’re in a production scenario where you generate in bulk at fixed resolutions, compile once and enjoy the speed for a long time - totally worth it. But if you’re constantly switching models and sizes, just messing around, that speed boost might not be worth the repeated recompilation headache. Depends on how you use it.