Structural text search and comparison using automatically extracted schema

Michael N. Gubanov, Philip A. Bernstein · International Workshop on the Web and Databases · 2006

An enormous amount of unstructured information is present on the web, in product manuals, e-mails, text documents, and other information sources. However, there is not enough support to automatically infer su‐cient structure from these data sources to be able to pose queries comparable in power to SQL. We present a prototype of new text database management system capable to automatically infer schema from text using natural language processing. It leverages extracted schema by supporting powerful structural search and fuzzy join operator between extracted entities.

Read the paper · More papers on PaperTik