Identifying the gist of conversational text: automatic keyword extraction and summarization
Yang Liu, Fei Liu · 2011
With the rapid development of communication technologies and mass storage techniques, many conversational texts have quickly emerged as significant information sources, such as emails, forums, meeting conversation transcripts, chat logs, microblogs, etc. The ability to identify the gist of these conversational texts enables us to quickly browse through the huge amount of data and obtain the essential information. On the other hand, the conversational text style also poses great challenges to the traditional language processing techniques, including redundancies, disfluencies, ill-formed sentence structure, high word error rates, and so forth. In this work, we focus on keyword extraction and summarization on meeting transcripts, and also explore summarizing the Twitter posts (tweets) as another domain of conversational text. We propose to extract keywords using a novel supervised framework that incorporates various knowledge sources: beyond the traditional widely used features (e.g., TF-IDF, position information), we introduce additional rich features including term specificity information, decision-making sentence related features, speaker and prominence based features, and features extracted from system generated summaries. We propose a feedback strategy to reinforce the impact of summary sentences on selecting effective keywords. We conduct analysis to evaluate feature effectiveness using different feature selection processes, and define various measurements to characterize the quality of summaries that can benefit the keyword extraction task. We also evaluate system performance using both human transcripts and different automatic speech recognizer (ASR) output (1-best and n-best), and show promising improved keyword extraction results using n-best ASR output over 1-best hypothesis. For extractive meeting summarization, we explore multiple meeting-specific characteristics. We propose to use topic labels and speaker-dependent characteristics (such as verboseness, gender, native language, role in the meeting) to improve extractive meeting summarization performance. These properties were incorporated in both unsupervised Maximum Marginal Relevance (MMR) approach and the supervised framework. We observe consistent improvements using our proposed approaches, on both human transcripts and ASR output, and using different evaluation metrics including ROUGE, Pyramid, and a DA-level F-measure score. Beyond extractive summarization, we propose to perform sentence compression on the extractive summary to improve its readability and make it more like an abstractive summary. Various automatic compression algorithms are investigated, including the integer linear programming (ILP) based approach with filler phrase detection, a noisy-channel approach using Markovization formulation of grammar rules, as well as the conditional random fields (CRF) based approach. The automatically compressed utterances are compared against both human compression and the abstractive summaries. We also evaluate the impact of using compressed utterances on summarization, and propose a fully automatic summarizer that generates compressed meeting summaries by combing the utterance compression module with an extractive summarization system. We perform exploratory summarization studies on another domain of conversational text – the Twitter posts, to help users quickly browse through any available topics. As an important first step, we propose a novel letter transformation approach to convert the nonstandard tokens in the tweets into standard English words. Different from the prior work, our approach requires neither pre-categorization nor human supervision. The approach models the generation process from the dictionary words to nonstandard tokens under a sequence labeling framework. We also explore summarizing the Twitter topics using the concept-based global optimization approach, and investigate the effect of both noisy nonstandard tokens andlinked web contents on the summarization performance.