Structured Information Extraction from Medical Texts in Bulgarian

Svetla Boytcheva · Cybernetics and Information Technologies · 2012

Abstract This paper presents an approach for Information Extraction (IE) from Patient Records (PRs) in Bulgarian. The specific terminology and lack of resources in electronic format are some of the obstacles that make the task of current patient status data extraction in a structured format quite challenging. The usage of N-grams, collocations and words’ distances allows us to cope with this problem and to extract automatically the attribute-value pairs with relatively high precision.

Read the paper · More papers on PaperTik