Grounded Language Acquisition from Object and Action Imagery
James Kubricht, Zhaoyuan Yang, Jianwei Qiu, Peter Tu · 2024
Deep learning approaches to natural language processing have made great strides in recent years. While transformer models have demonstrated impressive knowledge and reasoning capabilities, it is unclear how produced symbols are grounded in data from the world. In this paper, we explore the development of a private language for visual data representation by training emergent language (EL) encoders/decoders in both i) a traditional referential game environment and ii) a contrastive learning environment utilizing a within-class matching training paradigm. An additional classification layer-utilizing neural machine translation and random forest classification-was used to transform symbolic representations (sequences of integer symbols) to class labels. These methods were applied in two experiments focusing on object recognition and action recognition. For object recognition, a set of sketches produced by human participants from real imagery was used and for action recognition, 2D trajectory images were generated from 3D motion capture systems. In order to interpret the symbols produced for data in each experiment, a Gradient-weighted Class Activation Mapping (GradCAM) method was used to identify pixel regions indicating semantic features which contribute evidence towards symbols in learned languages. Results indicate that: i) symbols used to represent images appear to shift focus between different semantic components in an image and ii) this shift occurs gradually over the course of a sentence.