Sophisticated Text Mining System for Extracting and Visualizing Numerical and Named Entity Information from a Large Number of Documents
Masaki Murata, Tamotsu Shirado, Kentaro Torisawa, Masakazu Iwatate, Koji Ichii, Qing Ma, Toshiyuki Kanamaru · 2008
We have developed a system that can semiautomat-ically extract numerical and named entity sets from a large number of Japanese documents and can create various kinds of tables and graphs. In our experi-ments, our system semiautomatically created approx-imately 300 kinds of graphs and tables at precisions of 0.2–0.8 with only 2 h of manual preparation from a 2-year stack of newspapers articles. Note that these newspaper articles contained a large quantity of data, and all of them could not be read or checked manually in such a short amount of time. From this perspective, we concluded that our system is useful and convenient for extracting information from a large number of doc-uments. We have constructed a demonstration system. In this paper, we briefly describe the demonstration system.