Intelligent Document Processing with Small and Relevant Training Dataset

Lina Nicolaieff, Mohamed Mehdi Kandi, Younes Zegaoui, Christophe Bortolaso · 2022

Nowadays, companies deploy complex mechanisms to automate data collection, storage, and processing. Some of this data is unstructured and contained in pdf or scanned documents. Supervised object detection models exist to address the problem such as Faster-RCNN, but require retraining for each new use case. Annotation is a tedious and repetitive task done regularly when new document templates arrive. We present in this paper a method based on a Triplet-loss architecture to select a small and relevant subset of unstructured documents to annotate. To evaluate the method, we trained a model with many datasets. We compared the performance with different choices of document templates and dataset sizes. We show that a relevant and automated choice of document examples can avoid a huge annotation effort.

Read the paper · More papers on PaperTik