DeepDoc: Automated Document Segregation Using Deep Learning
Aayushi Shilledar, Abhiram Degwekar, Atharva Pande, Kalpit Bhawalkar · 2023
Document segregation has always been a significant issue as it is very tedious and time-consuming. It is something that is required almost everywhere as data is a prerequisite for manipulation, research, communication, and interpreting purposes. It is hence necessary to maintain and store it in a proper way for which document segregation plays a vital role. So, to help achieve this, the main aim is to develop a model which not only segregates the files belonging to the same document together but also helps us to distinguish whether or not given a pair of pages belong to the same document or not and this also saves manual human efforts is being developed. It is done using a standard Resnet-50 classifier and Image Processing algorithms. Phase-1 training and testing are completed for the model in which the model predicts whether or not given 2 pages belong to the same PDF. Whereas, Phase-2 segregates a given dump into their respective different folders.