The Application Study Based on Lucene Full-Text Retrieval
WU Dai-wen · Microcomputer applications · 2011
In the paper,it implements the second index in PDF document by Lucene API and PDFBox API.In order to locate the search Keyword more accurately,this paper designs and implements a new algorithm for the second index.It contains the information about the keywords' page number,coordinates,context and so on.Which can be made used of locating the retrieval results in the specific page of the book and marking the specific positions of the keywords.Thus,the effect of the second retrieval in PDF document is as similar as that in Baidu document library.