Towards a semantic book search engine

Shah Khusro, Irfan Ullah · 2016

Traditional Information Retrieval (IR) methods were initially used for searching and ranking web pages on the Web. These methods were progressively modified to exploit the peculiarities of the Web including the use of the hyperlinked structure of the Web for relevance ranking. These Web IR techniques, however, are also being applied for searching and ranking to other forms of text collections which are not inherently web documents. Books (especially in PDF form) are by nature different from web pages because they lack an explicit hypertextual structure and therefore cannot be accurately and precisely searched and ranked using traditional approaches. Books contain a highly structured content with implicit logical connections among different parts of the same book as well as to related content in other books. These book structural semantics and logical connections could be discovered and used to establish a web of books where the logical concepts, images, figures, tables, and other parts are linked with each other thus resulting in a semantic graph, which could then be exploited by a semantic book search engine for more precise and accurate indexing, searching, ranking and recommendations. Based on this hypothesis, the paper outlines a high-level architecture for one of the possible implementations of a semantic book search engine, identifies all the potential areas of research for future researchers, and reports on our work in progress in the form of the proposed model for the purpose. The proposed architecture, if implemented in its true sense, has the potential to better serve the needs of all the stakeholders including authors, publishers, readers, and librarians.

Read the paper · More papers on PaperTik