Estimation of errors in text and data processing

A. Slavova, B. Valkov, Krasimir Tonchev, Nediyana Daskalova, Mila Nikolova, M. Bivas, Plamen Mateev, Roumyana Yordanova, Stela Zhelezova · 2013

The company Adiss Lab Lts. obtained 1 000 000 medical reports that are either in free form text, or in XML format. One of the main goals of their development is to integrate an algorithm for information extraction (IE) in their platform. The verification of the algorithm’s output for a report is done by a medical doctor (MD) for a certain fee. Validating the correctness of all data would be overwhelming and very expensive. Hence, the problem, as presented by the company, is to provide a method (algorithm) which determines the minimum amount of reports that will validate the correctness of the IE algorithm and a procedure for selecting these reports. In order to solve the problem we have considered an algorithm-centric approach uses active learning and semi-supervised learning.

Read the paper · More papers on PaperTik