Iterative Learning of Relation Patterns for Market Analysis with UIMA

Sebastian Blohm, Jürgen Umbrich, Philipp Cimiano, York Sure · PUB – Publications at Bielefeld University (Bielefeld University) · 2007

Monitoring news as well as corporate and community web sites is a important task in market analysis. Yet, it requires much manual processing and is inherently incomplete as only a small fraction of the data produced on the Web can be processed by human analysts. We present here our ongoing efforts to employ UIMA annotations and search on UIMA-annotated documents for supporting market analysis. Information is extracted based on relation patterns that serve as queries. The results are then visualized to allow detailed market analysis. The current status of the reported work is that we are already implemented most of the required infrastructure (UIMA, Pronto, Omni nd) and started some of the outlined experiments. The Unstructured Information Management Architecture (UIMA) is a framework that enables a component-based analysis of unstructured documents(Ferrucci and Lally, 2004) providing a typed data structure for the derived structured information as well as means for de ning and controlling process ows. We have developed Pronto, a system for automatic learning of text patterns for extracting instances of semantic relations(Blohm and Cimiano, 2006). The system essentially uses the Web as a corpus accessing it via a Web search engine. Patterns are learned by abstraction over the textual occurrences of relation instances in the corpus. In the experiments we employ Pronto to produce UIMA annotations and additionally feed the UIMA annotation structure into the learning process to be able to learn structural patterns rather than plain text patterns. They will enable our system to annotated mentions of salient relations for market analysis. The output will allow retrieving facts about market events using structured queries.

Read the paper · More papers on PaperTik