Patch-Based 3D CNN-Transformer Approach for Non-Uniformly Compressed Voxel Classification
Felix Koba, Aleksandr Marin, Stevan Čakić · 2025
Three-dimensional (3D) voxel models have wide applications in fields such as robotics, medical imaging, autonomous navigation, and augmented reality. To address challenges of spatial sparsity and computational efficiency, this research proposes a voxel-based 3D convolutional neural network (3D CNN) integrated with a Transformer encoder for object classification on the ModelNet10 dataset. The 3D models are non-uniformly compressed into a 16×16×16 voxel grid, with compression factors incorporated as inputs to reduce sparsity. An in-house augmentation tool further enhances the dataset size and diversity by performing real-time voxel editing and labeling. The proposed method achieves an accuracy of 92.15% while utilizing only 184k trainable parameters, demonstrating an efficient and lightweight approach to object classification.