Domain Adaptation of Parts of Speech Annotators in Hindi Biomedical Corpus: An NLP Approach

Pitambar Behera, Om Prakash Jena · 2021

The envisaged research demonstrates the development of bio-medically annotated Parts of Speech (POS) corpus in Hindi. The study presents the adaptation of POS tagger trained in general domain corpus to automatically annotate the corpus of health domain. The tagger is trained with 200,000 word tokens applied from the ILCI (Indian Languages Corpora Initiative) data of mixed domains (in addition to 50k newswire tokens of biomedical data) which provides a satisfactory accuracy of 92%. When adapted and tested with the fresh data of the biomedical domain, the tagger registers an accuracy of 86.5%. In addition, the paper also focuses light on the resource-poor scenario of Hindi and other Indian regional languages in general domain and biomedical corpus in particular. Furthermore, the study provides a detailed account of the issues and challenges encountered pertaining to interrater reliability, domain adaptation of corpus, linguistics, and NLP (Natural Language Processing).

Read the paper · More papers on PaperTik