Unsupervised Linking of Visual Features to Textual Descriptions in Long Manipulation Activities

Eren Erdal Aksoy, Ekaterina Ovchinnikova, Adil Orhan, Yezhou Yang, Tamim Asfour · IEEE Robotics and Automation Letters · 2017

We present a novel unsupervised framework, which links continuous visual features and symbolic textual descriptions of manipulation activity videos. First, we extract the semantic representation of visually observed manipulations by applying a bottom-up approach to the continuous image streams. We then employ a rule-based reasoning to link visual and linguistic inputs. The proposed framework allows robots 1) to autonomously parse, classify, and label sequentially and/or concurrently performed atomic manipulations (e.g., “cutting” or “ stirring”), 2) to simultaneously categorize and identify manipulated objects without using any standard feature-based recognition techniques, and 3) to generate textual descriptions for long activities, e.g., “breakfast preparation.” We evaluated the framework using a dataset of 120 atomic manipulations and 20 long activities.

Read the paper · More papers on PaperTik