back
136 comments
I can't wait for this to become usable. I love VR but the content generation is just sooooo labour intensive. Help creating 3D models would help so much and be the #1 enabler for the metaverse IMO.
VR is especially unforgiving of "fake" detailing, you need as much detail as possible in the actual geometry to really sell it. That's the opposite how these models currently work, they output goopy low-res geometry and approximate most of the detailing with textures, which would be immediately register as fake with stereoscopic depth perception.
There are a few services that do this already, but they are all somewhat lacking, hopefully Meta's paper / solution brings some significant improvements in this space.

The existing ones:

- Meshy https://www.meshy.ai/ one of the first movers in this space, though it's quality isn't that great

- Rodin https://hyperhuman.deemos.com/rodin newer but folks are saying this is better

- Luma Labs has a 3D generator https://lumalabs.ai/genie but doesn't seem that popular

Regardless of cost, the meta verse is a dead end concept - at least until we get Ghost in the Shell-style computer-brain interfaces. Second Life probably was peak popularity for this kind of meta verse.
I've been bullish[1] on this as a major aspect of generative AI for a while now, so it's great to see this paper published.

3D has an extremely steep learning curve once you try to do anything non-trivial, especially in terms of asset creation for VR etc. but my real interest is where this leads in terms of real-world items. One of the major hurdles is that in the real-world we aren't as forgiving as we are in VR/games. I'm not entirely surprised to see that most of the outputs are "artistic" ones, but I'm really interested to see where this ends up when we can give AI combined inputs from text/photos/LIDAR etc and have it make the model for a physical item that can be 3D printed.

[1] https://www.technicalchops.com/articles/ai-inputs-and-output...

I have to be a soggy blanket person, but there are some pretty strong reasons why 3D generative AI is going to have a much shallower adoption cycle than 2D. (I founded a company that was likely the first generative AI company for 3D assets)

1. 3D is actually a broad collection of formats and not a single thing or representation. This is because of the deep relation between surface topology and render performance. 2. 3D is much more challenging to use in any workflow. Its much more physical, and the ergonomics of 3DOF makes it naturally hard to place in as many places as 2D 3. 3D is much more expensive to produce per unit value, in many ways. This is why, for example, almost every indie web comic artist draws in 2D instead of 3D. In an ai first world it might be less “work” but will still be leaps and bounds more expensive.

In my opinion, the media that have the most appeal for genAI are basically (in order)

- images - videos - music - general audio - 2D animation - tightly scoped 3D experiences such as avatars - games - general 3D models.

My conclusion from being in this space was that there’s likely a world where 3D-style videos generated from pixels are more poised to take off than 3D as a data type.

>I've been bullish[1] on this as a major aspect of generative AI for a while now, so it's great to see this paper published.

Me too. My first thought when seeing 2D AI generated images was that 3D would be a logical next step. Compared to pixels, there's so much additional data to work with when training these models I just assumed that 3D would be an easier problem to solve than 2D image generation. You have 3D point data, edges, edge loops, bone systems etc and a lot of the existing data is pretty well labeled too.

I'm excited to see the main deterrents to Indy gave dev: art and sound, get AI'd away. A single developer could use an off-the-shelf game engine and some AI generated assets (perhaps combining with whatever they can buy cheap) to develop some fun games.

Still, getting from still models to something that animates is necessary:(

Autodesk has been building practical 3d models for years with generative design. I have to imagine it's only getting better with these recent advances, but I'm not immersed in the space.
I tried all the recent wave of text/image to 3D model services, some touting 100 MM+ valuations and tens of millions raised and found them all to produce unusable garbage.
SOTA text-to-image 5 years ago was complete garbage. Most people would think the same. Look how good it got now.

You have to look at this as stepping stone research.

I have too, and you’re quite right. Also the various 2D-to-3D face generators are mostly awful. I’ve done a deep dive on that and nearly all of them seem to only create slight perturbations on some base model, regardless of the input.
We tried them too. My wife is a 3D artist, but we needed a lot of assets that frankly weren't that important. The plan was to use the output as a starting point and improve as needed manually.

The problem is that the output you get is just baked meshes. If the object connects together or has a few pieces you'll have to essentially undo some of that work. Similar problems with textures as the AI doesn't work normally like other artists do.

All of this is also on top of the output being basically garbage. Input photos ultimately fail in ways that would require so much work to fix it invalidates the concept. By the time you start to get something approaching decent output you've put in more work or money than just having someone make it to begin with while essentially also losing all control over the art pipeline.

The field evolves quickly. The Meta 3D Gen paper states that Rodin Gen-1 [1,2,3] has a clean topology. As a non professional, the wireframes from some examples indeed look nice.

[1]: https://github.com/CLAY-3D/OpenCLAY

[2]: https://hyperhuman.deemos.com/

[3]: https://huggingface.co/spaces/DEEMOSTECH/Rodin

The gap from demos/papers to reality is huge. ML has a bad replication crisis.
Haven’t tried all, but yeah, pretty bad so far
This is crazy impressive, and the fact they have the whole thing running with a PBR texturing pipeline is really cool.

That being said, I wonder if the use of signed distance fields (SDFs) results in bad topology.

I saw a paper earlier this week that was recently released that seems to build "game-ready" topology --- stuff that might actually be riggable for animation. https://github.com/buaacyw/MeshAnything

The obvious major caveat with MeshAnything is that it only scales up to outputs with about 800 polygons, so even if their claims about the quality of their topology hold up it's not actually good for much as it stands. For reference a modern AAA game character model can easily exceed 100,000 polygons, and models made to be rendered offline can be an order of magnitude bigger still.
Looks fine, but you can tell the topology isn’t good based on the lack of wireframes.
They seem to admit as much in Table 1 which indicates this model is not capable of "clean topology". Somewhat annoyingly, they do not discuss topology anywhere else in the paper (at least, I could not find the word "topology" via Ctrl+F).
Credit where it's due, unlike most of these papers they do at least show some of their models sans textures on page 11, so you can see how undefined the actual geometry is (e.g. none of the characters have eyes until they are painted on).
Afaik, there's no topology - it outputs signed distance fields, not meshes.
That doesn’t matter for things like 3D printing and CNC machining. Additionally, there are other mesh fixer AI tools. This is going to be gold for me.
Such a silly argument. Fixing topology is a nearly solved problem in geometry processing. (Or just start with a good topology and 'paste' a texture onto it like they develop techniques for here.)
I think this is another precursor step in recreating our reality digitally. As long as you're able to react to the persons' state, with enough metrics you're able to recreate environments and scenarios within a 'safe environment' for people to push through and learn to cope with the scenarios they don't feel safe to address in the 'real' world.

When the person then emerges from this virtual world, it'll be like an egg hatching into a new birth, having learned the lessons in their virtual cocoon.

If you don't like this idea, it's an interesting thought experiment regardless as we can't verify, we're not already in a form of this.

Interesting, but what are the practical uses of 3d assets beyond gaming, where does it create a real advantage over what we already use as visual information and user interfaces? I cannot see VR replacing the interactions we have. It requires cumbersome, expensive hardware, it floods the users with additional mostly useless information (image, sound, 3d itself) they have to process, it's slow and expensive to create and maintain, in short: it's inefficient compared to established tech, which will always run circles around the lame try of imitating real world interactions in a 3d virtual space. The potential availability of very expensively (in terms of computing power) generated assets doesn't change that. It's still hard to do right, and even if done right, it seems like only a gimmick hardly anyone can stomach for more then a couple of hours at best. It's information overload to most people, and they have better alternatives.
> Interesting, but what are the practical uses of 3d assets beyond gaming

Probably many areas that we already use 3D assets/texturing for. Maybe objects to fill out an architectural render, CG in movies/TV shows, 3D printing, or just as an inspiration/mock-up to build off of. I'd imagine this generator is less useful for product design/manufacturing at the moment due to lack precise constraints - but maybe once we get the equivalent of ControlNets.

If weights are released, it may also serve as a nice foundation model, or synthetic data generator, for other 3D tasks (including non-generative tasks like defect detection), in the same way Stable Diffusion and Segment Anything have for 2D tasks.

> I cannot see VR replacing the interactions we have. It requires cumbersome, expensive hardware

Currently sure, but it's been a reasonably safe bet that hardware will get smaller and cheaper. Something like the Bigscreen Beyond already has a fairly small form factor.

But, I feel you're basing judgement of a 3D generator on one currently-niche potential use of 3D assets, that being VR/AR user interfaces (and in particular ones intended to replace a phone rather than, for instance, the interactive interfaces within VR games/experiences).

> The potential availability of very expensively (in terms of computing power) generated assets doesn't change that

Even just comparing computing power and not the human labour required, this is probably going to be an extremely cheap way to generate assets. The paper reports 30 seconds for AssetGen, then a further 20 seconds for TextureGen - both being feed-forward generators. They don't mention which GPU, but previous similar models have ran in a couple of minutes on consumer GPUs.

Seeems like simple enough 3D-to-3D will be possible soon!

I'll use it to upscale 8x all meshes and textures in the original Mafia and Unreal Tournament, write a good bye letter to my family and disappear.

I think the kids will understand when they grow up.

In the comparison between the models only Rodin seems to produce clean topology, hopefully in the future we will see a model with the strength of both, hopefully from Meta as Rodin is a commercial model.
Ya would be cool if we had something open that competed with rodin, but just like elevenlabs for voice, seems closed is gonna be ahead for a while
Would love for an artist to provide some input, but I imagine this could be really good if it generates models that you can edit or start from later .

Or, just throw a PS1 filter on top and make some retro games

Sure.

The paper doesn't show topology, UVs or the output texture, so we're left to assume the models look something like what you'd find when using photogrammetry: triangulated blobs with highly segmented UV islands and very large textures. Fine for background elements in a 3D render, but unsuitable for use in a game engine or real-time pipeline.

In my job I've sometimes been given 3D scans and asked to include them in a game. They require extensive cleanup to become usable, if you care about visual quality and performance at all.

Unless the topology is good it may not be worth it.
> for an artist to provide some input

Sure, the results are excellent.

> Or, just throw a PS1 filter on top and make some retro games

There's so many creative ways to use these workflows. Consider how much people achieved with NES graphics. The biggest obstacles are tools and marketplaces.

Can this potentially support :

- Image Input to 3D model Output

- 3D model(format) as Input

Question: What is the current state of the art commercially available product in that niche?

This a pipeline for text to 3D.

But it's using for 3D gen, a model that is more flexible:

https://assetgen.github.io/

It can be conditioned on text or image.

can somebody please please integrate SAM with 3d primitive RAGging? This is the golden chalice solution as a 3d modeler, having one of those "blobs" generated by Luma and likes aren't very useful
I’m puzzled by the poor texture quality in these. The colours are just bad - it looks like the textures are blown out (the detail at the bright end clip into white) and much too contrasty ( the turkey does that transition from red to white via a band of yellow). I wonder why that is - was the training data just done on the cheap?
Probably this is the best way to build the Metaverse. Publish all the research, let people build products over it and soon we'll in need for a place and platform to make use of all the instant assets in virtual spaces.

Well played, Meta.

Can this be used for image to 3D generation? What is the SOTA in this area these days?
Is there a way to try this yet?
Meta 3D Gen represents a significant step forward in the realm of 3D content generation, particularly for VR applications. The ability to generate detailed 3D models from text inputs could drastically reduce the labor-intensive process of content creation, making it more accessible and scalable. However, as some commenters have pointed out, the current technology still faces challenges, especially in producing high-quality, detailed geometry that holds up under the scrutiny of VR’s stereoscopic depth perception. The integration of PBR texturing is a promising feature, but the real test will be in how well these models can be refined and utilized in practical applications. It’s an exciting development, but there’s still a long way to go before it can fully meet the needs of VR developers and artists.
For starters, I'd love to just see a rock-solid neural network replacement for screened poisson surface reconstruction. (I have seen MeshAnything and I don't think that's the end-game.)
Are those guys still banging on about that Metaverse? That's taken a decided back seat to all the AI innovation in the past 18 months.
Why these pages want to bother visitors with popups? Just use "only essential" as default.
Is this going toward 3D games entirely "hallucinated"? That would be amazing.
Is this just a paper or can I run the program and generate some stuff?
Not sure how adding Gen AI is going to make VR any better? I wanted to type "it's like throwing good money after bad", but that's not quite right. Both are black holes where VC money is turned into papers and demos.
Work like this is the only way to revive the now defunct Metaverse, I was wondering whether Meta would fund research such as this that could lower the financial barrier to entry for Metaverse participants