Voice-Powered 3D Object Generation for the Metaverse

Pierre Dave Victor Katuhe, An‐Chao Tsai · 2023

This research proposes a method for generating or modifying 3D objects using a human's voice as input for a machine learning model. The input voice is translated into text using a Google API and processed by a BERT-based text encoder to comprehend each phrase and generate the desired 3D object. The method is designed to make it easier for people with no design knowledge to create 3D objects and contribute to the development of the Metaverse, a virtual reality beyond reality. The shape generation model has been trained on a dataset of 15,038 shapes with 75,344 natural language descriptions, and the system includes a shape auto-encoder to generate different styles of 3D shapes. The proposed method is the first to use speech recognition for the generation of 3D objects in the Metaverse and aims to generate higher quality and faster results than existing methods using text for 3D object generation.

Read the paper · More papers on PaperTik