When Text and Speech are Not Enough: A Multimodal Dataset of Collaboration in a Situated Task
Ibrahim Khebour, Richard A. Brutti, Indrani Dey, Rachel Dickler, Kelsey Sikes, Kenneth H. Lai, Mariah Bradford, Brittany E. Cates, Paige Hansen, Changsoo Jung, Brett Wisniewski, Corbyn Terpstra, Leanne Hirshfield, Sadhana Puntambekar, Nathaniel Blanchard, James D. Pustejovsky, Nikhil Krishnaswamy · Journal of Open Humanities Data · 2024
To adequately model information exchanged in real human-human interactions, considering speech or text alone leaves out many critical modalities. The channels contributing to the “making of sense” in human-human interactions include but are not limited to gesture, speech, user-interaction modeling, gaze, joint attention, and involvement/engagement, all of which need to be adequately modeled to automatically extract correct and meaningful information. In this paper, we present a multimodal dataset of a novel situated and shared collaborative task, with the above channels annotated to encode these different aspects of the situated and embodied involvement of the participants in the joint activity.