Gathering Public Concerns from Web Towards Building Corpus of Japanese Regional Concerns

Shun Shiramatsu, Norifumi Hirata, Robin M. E. Swezey, Hiroyuki Sano, Tadachika Ozono, Toramatsu Shintani · 2012

Importance of concern assessment has been increased in Japanese regional communities. We have developed an e-Participation web platform based on a Linked Open Data set called SOCIA (Social Opinions and Concerns for Ideal Argumentation). To sophisticate text mining technologies for supporting concern assessment, building a corpus of public concerns is an urgent task. There are two issues to utilize the dataset SOCIA as a corpus: (1) it is required to manage reliability of annotation and (2) to filter out noisy text not relevant to public concerns. To address these research issues, (1) we incorporate schema for describing meta-context information of annotation, that is, who is annotator, whether the annotator is a human or a software agent, and how reliable the annotation is. Furthermore, (2) we investigate the difference between features of concerns and that of non-concerns in Japanese microblog posts (i.e., tweets). Through the investigation, we address sample selection bias by formulating a novel metric for ranking features, i.e., bias-penalized information gain (BPIG).

Read the paper · More papers on PaperTik