ForumSum: A Multi-Speaker Conversation Summarization Dataset
Misha Khalman, Yao Zhao, Mohammad Saleh · 2021
Abstractive summarization quality had large improvements since recent language pretraining techniques.However, currently there is a lack of datasets for the growing needs of conversation summarization applications.Thus we collected ForumSum 1 , a diverse and highquality conversation summarization dataset with human written summaries.The conversations in ForumSum dataset are collected from a wide variety of internet forums.To make the dataset easily expandable, we also release the process of dataset creation.Our experiments show that models trained on ForumSum have better zero-shot and few-shot transferability to other datasets than the existing large chat summarization dataset SAMSum.We also show that using a conversational corpus for pre-training improves the quality of the chat summarization model.