Disentangling homonyms- using artificial neural networks to separate the cream from the crop in large text corpora
Uri Roll, Ricardo Correia, Oded Berger‐Tal · Proceedings of the 5th European Congress of Conservation Biology · 2018
ft through large corpora of academic texts and classify them to distinct topics. As an example, we explore the use of the word 'reintroduction' in academic texts. Reintroduction is used within the conservation context to indicate the release of organisms to their former native habitat, however an 'ISI' search using this word returns thousands of publications that use this term with other meanings and contexts. Using our method, we were able to quickly and correctly classify thousands of academic texts with more than 99% accuracy between conservation related and unrelated publications. Our approach can be easily used with any other homonym terms and can greatly facilitate sorting data in cases where homonyms hinder the harnessing of large text corpora. Beyond homonyms we see great promise in the combination of automated content analyses and machine learning methods in handling and screening big data for relevant information.