Summarization based approach for Old Sinhala Text Archival Search and Preservation

K. A. M. P. Rathnasena, K. M. S. J. Kumarasinghe, D. T. P. Paranavitharana, D. V. A. U. Dayarathne, L. Ranathunga · 2018

Old books are to be preserved and protected for the future needs. Preservation of these archives is crucial. The preservation and conservation of ancient and old antiques can be done using digitization so that they can be preserved for many years. The screw errors, noises and poor printing mechanisms make it challenge to recognition. Correcting the misspelled Sinhala words is also a challenge because Sinhala is a complex language. This paper elaborates an extensive approach derived through machine vision and natural language processing to preserve old text content as digitally searchable content. The scanned images of old books are taken and preprocess them to remove the noises. The Segmentation is done to ease the recognition of characters. After Optical Character Recognition, Sinhala spell correction is done to correct the misspelled words. The system provides separate summaries in book wise and chapter wise to get an abstract idea of books and chapters. Summary creation for Sinhala language is a challenge as Sinhala is a structured language. The System has mitigated most of these challenges successfully by achieving an average of 84% success in Text Line Segmentation and Layout Feature Identification, average of 74% success for OCR, average of 70% success for OCR Correction, average of 75% success for Keyword Extraction and average of 52% success for Summarization.

Read the paper · More papers on PaperTik