Learning to Rank Adaptively for Scalable Information Extraction

Pablo Barrio, Gonçalo Simões, Héléna Galhardas, Luis Gravano · 2015

Information extraction systems extract structured data from natural language text, to support richer querying and anal-ysis of the data than would be possible over the unstruc-tured text. Unfortunately, information extraction is a com-putationally expensive task, so exhaustively processing all documents of a large collection might be prohibitive. Such exhaustive processing is generally unnecessary, though, be-cause many times only a small set of documents in a collec-tion is useful for a given information extraction task. There-fore, by identifying these useful documents, and not process-ing the rest, we could substantially improve the efficiency and scalability of an extraction task. Existing approaches for identifying such documents often miss useful documents and also lead to the processing of useless documents unnec-essarily, which in turn negatively impacts the quality and efficiency of the extraction process. To address these limita-tions of the state-of-the-art techniques, we propose a prin-cipled, learning-based approach for ranking documents ac-cording to their potential usefulness for an extraction task. Our low-overhead, online learning-to-rank methods exploit the information collected during extraction, as we process new documents and the fine-grained characteristics of the useful documents are revealed. Then, these methods decide when the ranking model should be updated, hence signifi-cantly improving the document ranking quality over time. Our experiments show that our approach achieves higher ac-curacy than the state-of-the-art alternatives. Importantly, our approach is lightweight and efficient, and hence is a sub-stantial step towards scalable information extraction. 1.

Read the paper · More papers on PaperTik