Ablation Study to Clarify the Mechanism of Object Segmentation in Multi-Object Representation Learning
Takayuki Komatsu, Yoshiyuki Ohmura, Yasuo Kuniyoshi · 2024
The goal of multi-object representation learning is to represent a visual input that contains multiple objects. Multi-object representation learning methods have adopted simultaneous unsupervised learning of segmenting an input image into individual objects and encoding these objects into each latent vector. Previous methods have combined many techniques, such as image reconstruction, regularization of latent vectors, and other auxiliary loss functions. Therefore, it is not clear what is the essential and simple mechanism that contributes to the appropriate behavior of multi-object representation learning. In this study, we focused on elucidating the mechanism of object segmentation and conducted the ablation study on the loss functions in the multi-object representation learning method. We employed MONet [1] as the target of the ablation study and evaluated the object segmentation performance when each loss function in MONet is removed or replaced. Our results showed that the Variational Autoencoder regularization loss [2], which is used for single object representation learning, did not affect segmentation performance and the other losses did affect it. Based on this result, we hypothesized that it is important that there exists a winner-take-all mechanism among multiple latent vectors. We confirmed this hypothesis by evaluating a new loss function with the same mechanism as the hypothesis.