Extraction and Visualization of Trend Information from Newspaper Articles and Blogs

Hidetsugu Nanba, Nao Okuda, Manabu Okumura · 2007

Trend information is a summarization of temporal statistical data, such as changes in product prices and sales. We propose a method for extracting trend information from multiple newspaper articles and blogs, and visualizing the information as graphs. As target texts for extraction of trend information, the MuST (Multimodal Summarization for Trend Information) workshop focuses on newspaper articles. In addition to newspapers, we focus on blogs, because useful information for analysing trend information is often written in blogs, such as the reasons for increases/decreases of statistics and the impact of increases/decreases of statistics on society. To extract trend information, we extract temporal expressions and statistical values, and we devised methods for both operations. To investigate the effectiveness of our methods, we conducted some experiments. We obtained a recall of 6.3 % and precision of 31.3 % for newspaper articles, and a recall of 44.8 % and precision of 60.3 % for blogs. From the error analysis, we found that most errors in newspaper articles were caused by misconversion of temporal expressions such as “�� ” (the same year) or “�� ” (the previous month), into “YYYY-MM-DD ” form, although temporal expressions were detected correctly. In contrast to newspaper articles, there are few temporal expressions in blogs for which resolution is required, such as “�� ” (the same day) or “�� ” (the previous month). As a result, recall and precision for blogs are higher than those for newspaper articles.

Read the paper · More papers on PaperTik