Identifying Sub-documents in a Composite Scanned Document Using Naive Bayes, Levenshtein Distance and Domain Driven Knowledge Base
Oluwatosin Ogundare, Nathanial Wiggins · 2018
This paper presents an approach to classification of scanned documents by identifying the title of the document using Naïve Bayes classifier, the predicted title is then validated using a domain driven knowledge base. Levenshtein distance is used to mitigate errors arising from the OCR (optical character recognition) algorithm. This approach produced significantly better results than using the Naïve Bayes classifier by itself. This study contributes resources to the intelligent processing of real estate documents in the form of rich domain specific knowledge base.