Patch-Based 3D CNN-Transformer Approach for Non-Uniformly Compressed Voxel Classification

Felix Koba, Aleksandr Marin, Stevan Čakić · 2025

Three-dimensional (3D) voxel models have wide applications in fields such as robotics, medical imaging, autonomous navigation, and augmented reality. To address challenges of spatial sparsity and computational efficiency, this research proposes a voxel-based 3D convolutional neural network (3D CNN) integrated with a Transformer encoder for object classification on the ModelNet10 dataset. The 3D models are non-uniformly compressed into a 16×16×16 voxel grid, with compression factors incorporated as inputs to reduce sparsity. An in-house augmentation tool further enhances the dataset size and diversity by performing real-time voxel editing and labeling. The proposed method achieves an accuracy of 92.15% while utilizing only 184k trainable parameters, demonstrating an efficient and lightweight approach to object classification.

Read the paper · More papers on PaperTik