Equivariant networks for 3D reasoning from images

David Klee · 2024

Equivariant neural networks generalize predictions across symmetric transformations ofthe input. Such networks could be applied to many computer problems that exhibit 3D rotation or translation symmetry to improve accuracy and sample-efficiency. However, equivariant networks require that the input can be meaningfully transformed by the symmetry group, which prevents their application to 2D image inputs that cannot be transformed by out-of-plane rotations. This thesis explores network architectures that learn 3D equivariant features from 2D inputs. In particular, we introduce methods that map 2D image features onto the discrete Icosahedral group and continuous SO(3) group to accurately predict object orientation with significantly less data than non-equivariant methods. We provide a mechanism for predicting complex, multi-modal distributions over 3D rotations, and formalize the space of rotation-equivariant mappings from 2D to 3D representations. The proposed methods achieve state of the art performance on the challenging PASCAL3D+ pose estimation benchmark. The final work in the thesis investigates techniques to improve equivariant robot learning algorithms when the input symmetry does not align with the task symmetry, such as with a freely placed camera. These techniques increase success rates consistently across tasks, with both RGB and RGBD image observations.--Author's abstract

Read the paper · More papers on PaperTik