After generating millions upon millions of images, artificial intelligences are preparing to gain movement and dimension. The role of 3D modeling in the industry is gigantic, and if a new platform can further streamline the creation process, it could spark a real revolution. With that in mind, we discovered Point-E, a system designed by OpenAI that generates three-dimensional models combining text and point clouds.
From astronaut kittens to fantastical scenes full of details, generative image models have led to an explosion of activity on the Web. Some of the public has focused on debating the legal and ethical situation of certain developments, but one thing is certain: the advancement of this technology is not going to stop.
What's next on the list? Besides video, another priority is 3D modeling. Generally, professional tools have a fairly complex learning curve, but if the concept of "text-to-3D" becomes a reality, it wouldn't be unreasonable to talk about a democratization of three-dimensional design. The people at OpenAI are heading in that direction after confirming the opening of their Point-E project, an equivalent to DALL-E for 3D objects.
Point-E: From Text to 3D Objects with AI
According to the official paper, one of the main advantages of Point-E is its speed. While Google's DreamFusion requires a significant amount of time and graphics cards to generate models, Point-E takes less than two minutes on an Nvidia V100 unit. Jim Fan of Nvidia explains that Point-E is 600 times faster than DreamFusion, although it falls short in quality.
This is due to the generation method in Point-E. The system uses point clouds to represent the three-dimensional object. This limitation reduces the precision of certain features such as texture or final shape, but the objects are much easier to reproduce. Point-E's trick to compensate for this is converting the point cloud to meshes (read as "meshes"). In relaxed terms, we can speak of three models: one from text to image, another from image to 3D point cloud, and the third responsible for the mesh..
OpenAI trained this system with several million 3D objects, including metadata. Now, Point-E tends to make mistakes, especially when interpreting the image generated by the first model, but there is a lot of work ahead, which will inevitably lead to more optimized models.
Official site: Click here
Access the study (PDF): Click here