Opinion Mapping: Information Visualization Approaches for Comparative Sentiment Analysis
William H. Hsu, Praveen Koduru · 2012
In this position paper, we discuss the problem of extracting information about chronic diseases from the large volume of text written in health blogs, mailing lists, forums, and other electronic venues, then making this information accessible via structured queries, while analyzing it to map out patterns among the opinions and demographics of users. Information retrieval systems exist for spatially-referenced demographic data about diseases such as diabetes and their therapies, but as in the case of communicable diseases, the databases that contain such data are manually populated. For example, in information portals such as HealthMap.org, which are searchable by location and disease, the data are user-reported and collaboratively maintained, but not automatically extracted from text. Furthermore, there is as yet no automated means of relating sentiments expressed by users in their text postings to their semistructured profile data. This is because the primary sources for this kind of information have been statistical surveys such as opinion polls, where text responses are often human-interpreted and demographic analysis is done post hoc, rather than as part of an information retrieval and extraction task. These limitations indicate a present need for text summarization techniques that integrate quantitative information extraction – which captures symptoms, diseases, and complications of diseases – with opinion summarization.