It's mostly "did you know you can attach an LLM to a PCG system?", but there is one idea here you don't see much of: an image model performs the composition (which image models are really good at), and then you extract the objects into 3d via things like SAM3D before placing them in the world, which is pretty interesting.
The rest is standard stuff you'll find in your favorite PCG system/game engine. Still, think most people don't realize how good LLMs are at 3D these days, especially Fable 5 (which this project predates).
Use any image generation tool (or any image) and have a multimodal LLM like Gemini extract individual parts of the image “isolate with a transparent background” - then you can use those images for image-to-3d in those tools.
Don’t forget to use low or “smart” poly features otherwise you get too many vertices to the point you can’t performantly raycast etc.
But yeah - it’s there.
I've found it's actually best to do maximal quality generation and decimate/meshopt to desired budget.
If you're raycasting against your raw art you're doing it wrong anyway, none of the big players do this precisely for the reason you mentioned. And beyond physics meshes, if you're doing anything where performance is a problem you probably want multiple LODs so you have to do this work anyway.
These days you can just ask your agent to build you a LOD pipeline and forget about it rather than sacrificing the art; Claude knows Blender.
Decimating ends up ruining the look of the model to get it down to a game ready size.
Even a simple house model, or humanoid, using Decimate in Blender there is no way that will be even recognizable.
(I agree with you about Blender though, I always hated Blender's simplify)
Some 3D AI providers still leak like that, but Meshy came up with their own binary format that splits the stream up into separate files that are assembled with a response it only gets if you have Premium.
Thought it was cool technology - and wonder if browser games in general are doing stuff like that to protect 3D assets from being easily ripped from the Network tab.
I get that you run this, and then edit it. But the generated villages just aren’t interesting in my opinion. It’s probably great for tencent’s market where you’re mass producing gacha style games, but I don’t think worlds made in this fashion will scratch the open world itch like the best open world games out there.
I wouldn't say we don't need artists or writers, but I think most people will be surprised how quickly fully AI driven pipelines will advance here.
My sneaking suspicion is that there are a lot more cases than get reported, because authors are disincentivized to disclose AI usage, and readers by and large feel icky about reading AI work for reasons unrelated to quality.
I can totally see a weird future where a lot of what people like is AI generated, and everyone lies to each other about it. So it's hard to do research on this phenomenon.
In the summer example in the hero images, the building placement + small pockets of water on the left looks odd and low attention to detail. A similar poor quality result as if an uncaring human used a scatter brush...
Curious if the examples are cherry-picked and by how much, or if this is one-shoted
Does it matter? Why?
In modern polished games, the art team spends a lot of time tweaking things to make sure it looks great and coherent. Most people are not willing to have their agent run for a month to do this work. I don't know if it's been tried yet but I would be really interested to see the result.
Usually a combination of distance and frustum culling (breaking the model down into several that only load when you need to see it) with 4k or 8k textures can make a low poly game ready asset look very high quality