Fei-Fei Li's team dropped a multimodal world model — one image fills in the whole 3D world

Let’s get this out of the way first: robot training grounds are what this thing is actually aimed at. Gaming is just riding along.

The news itself is that Fei-Fei Li’s team released a multimodal world model — feed it one image and it fills in the full 3D world, and you can also use it to build training grounds for robots. Both are really the same direction: growing interactive space out of 2D input.

Game devs are drooling over the first half of that. Building greybox scenes eats up so much manpower — if you could go straight from a concept image to walkable space, early iteration speed would be on a whole different level.

But how it handles occluded backs, and whether the generated geometry is actually editable — none of that got explained. Don’t get too excited yet.

Damn

I tried a tool that lays out greybox from a concept image — poly count exploded and you basically couldn’t edit it anymore.

“World model” has gotten a bit overused the past couple years, better wait for the actual API before getting excited.

1 Like

Running this for real is probably not light on VRAM :sweat_smile:

First

Would save a ton of time for graybox scenes, as long as the output is actually editable.