Improvement on Lucene text analyzer
WU Dai-wen · Information technology newsletter · 2011
This paper aims at the shortcomings of Lucene only analyzing and indexing HTML and TXT documents.It extracts the text from XLS,PPT,Doc and PDF documents by using open source tools such as POI and PDFBox.Then it indexes these extracted text with Lucene and encapsulate these information to Lucene Document object.So Lucene can analyze and index Doc,XLS,PPT and PDF documents.The experimental results show that it can greatly enhance retrieval adaptability by improving Lucene text parser.