GiCCS: A German in-Context Conversational Similarity Benchmark

Shima Asaadi, Zahra Kolagar, Alina Liebel, Alessandra Zarcone · Zenodo (CERN European Organization for Nuclear Research) · 2022

We introduce GiCCS, a first conversational STS evaluation benchmark for German. We collected the similarity annotations for GiCCS using best-worst scaling and presenting the target items in context, in order to obtain highly-reliable context-dependent similarity scores. In our paper, we present benchmarking experiments for evaluating LMs on capturing the similarity of utterances. Results suggest that pretraining LMs on conversational data and providing conversational context can be useful for capturing similarity of utterances in dialogues. GiCCS will be publicly available to encourage benchmarking of conversational LMs.

Read the paper · More papers on PaperTik