Exploiting structure for intelligent Web search
Udo Kruschwitz · 2005
Together with the rapidly growing amount of online data, we register an immense need for intelligent search engines that access a restricted amount of data as found in intranets or other limited domains. These sorts of search engine must go beyond simple keyword indexing/matching, but they also have to be easily adaptable to new domains without huge costs. The paper presents a mechanism that addresses both of these points: first of all, the internal document structure is being used to extract concepts which impose a directory-like structure on the documents, similar to those found in classified directories. Furthermore, this is done in an efficient way which is largely language independent and does not make assumptions about the document structure.