Annotating the World Wide Web

Boris Katz, Jimmy Lin · 2001

The Problem: Although vast amounts of information are available electronically today, no effective mechanism exists to provide humans with convenient access to that information. Motivation: Keyword search engines are popular because they provide results—often, too many results! If we could understand, at least partially, the meaning of documents, rather than just recognizing the words, wecould answer queries with much higher precision. Natural language is the most convenient and most intuitive method of information access, and people should beabletoretrieve information using a system capable of understanding and answering natural language questions. Previous Work: In [2], we proposed making use of natural language annotations to annotate the World Wide Web. Approach: Technology is not up to the task of analyzing the semantic content of unrestricted data—complex text, tables, images, sound files, etc. However, our experiments with the START system [3] show how this problem could be solved for a relatively small knowledge base using our annotation technology. Annotations are short, simple sentences and phrases which computers can analyze. We associate them with data which is opaque to analysis. Our question answering system can then analyze questions, match them against already analyzed and stored annotations, and retrieve opaque dataassociated with matching annotations. For example, we can associate an individual annotation, such as the sentence “John Adams discovered Neptune using mathematics, ” with a detailed paragraph

Read the paper · More papers on PaperTik