A Framework and Tool for Collaborative Extraction of Reliable Information
Graham Neubig, Shinsuke Mori, Masahiro Mizukami · 2013
This research proposes a framework for ef-ficient information extraction and filtering in situations where 1) extreme reliability is important, 2) the amount of information to be combed through is massive, and 3) we can expect a relatively large number of human workers to be available. In particu-lar, we are motivated by needs in times of crisis, and assume that in order to ensure the high level of reliability required, it will be necessary to have at least one human worker confirm all extracted information. Given this setting, we propose a method to improve the efficiency of manual veri-fication by deciding which information to present to workers using machine learn-ing techniques. Even given this efficient search framework, the amount of informa-tion on the internet is still too much for one user to handle, so we additionally create a web-based framework that allows for col-laborative work, and an algorithm that al-lows for this framework to work on large data in real-time. We perform an eval-uation using data from Twitter after the Great East Japan Earthquake, and com-pare efficiency using both traditional key-word search and the proposed learning-based method. 1