Evaluation of a Coreference Resolution Model on Mention Patterns with Common Ground
Jaap P. Kruijt · Frontiers in artificial intelligence and applications · 2022
Efficient communication between humans and AI requires the AI to understand references that are made during the conversation.For this, the AI needs to establish common ground with the human in order to make correct judgments and take the right actions [1].Common ground is the shared information that speakers rely on during a conversation, which is built up over time as the speakers share more interactions [2].In our research, we focus on the role of common ground in resolving third-person references in humanrobot social interaction.In social dialogue, these references are often vague and contextdependent.Especially in a conversation between two well-acquainted individuals, the references to people they both know can become highly ambiguous.Through shared interactions, these references become increasingly efficient (i.e.shorter).This comes at the cost of intelligibility for outsiders who do not share the common ground [3].For machines that have no understanding of the common ground then, these references are difficult to interpret and relate to other references.In future work, we aim to approach this problem by building a reference resolution model which utilises a knowledge-rich approach and builds up common ground with a human in an interactive setting, where the robot and the human can coordinate to form the common ground together.In preparation for this, here we first investigate the limitations of existing reference resolution models in social interaction scenarios by evaluating to what extent they utilise common ground in resolving vague references in social dialogue.Our expectation is that these models do not fare well with references that require long-distance common ground knowledge, but that providing them with the relevant background knowledge will improve performance.Machine learning-based coreference resolution models can achieve impressive performance (e.g.[4,5]).However, most datasets used in coreference resolution tasks consist of snippets of formal text, e.g. from news articles.These datasets are not useful for