Supporting Information Retrieval of Emerging Knowledge and Argumentation
Nawroth, Christian · 2021
In research-oriented domains, e.g., the medical domain, new or emerging knowledge is permanently created through research and scientific discourse. This fact is, e.g., reflected by a permanent increase in scientific publications over the years. This overall increase and permanent creation of new knowledge make it hard for domain experts to find the right and relevant recent knowledge for a given task. In the medical domain, this could be the use of emerging knowledge in medical argumentation use cases, e.g., for or against a particular therapy. Supporting medical argumentation through textual evidence, in general, is the aim of the DFG-funded project RecomRatio, to which this thesis relates to. Hence, this work intends to make emerging knowledge in large medical document corpora available for evidence-based medical argumentation use cases. Therefore, it utilizes methods from the computer science subdomains of Information Retrieval, Natural Language Processing and Named Entity Recognition, Machine Learning, and Argumentation Mining to support evidence-based medical argumentation. The thesis introduces the motivation and challenges as addressed above, the research method, the research goal, research questions and objectives, and an outline of the thesis. The second chapter covers state-of-the-art research of the relevant fields in science and technology, i.e., Informational Behaviour and Information Retrieval, Vocabularies and Corpora, Machine Learning, Evaluation Methodologies, Natural Language Processing, Emerging Entities, and Argumentation Mining. Comparing state-of-the-art in these fields and the research objectives, the remaining research challenges are identified. These will be addressed in the following chapters. The third chapter conceptual design starts with different quantitative and qualitative studies that reveal the relevance of emerging knowledge in medical Information Retrieval and medical argumentation. Based on these insights, an innovative system is designed that integrates and adapts state-of-the-art approaches from Information Retrieval, Natural Language Processing, Machine Learning, and Argumentation Mining. The system’s core contribution is the design of a hybrid approach combining Natural Language Processing with Machine Learning on corpus-related features to extract emerging knowledge. In the system design, three real-world applications are conceptually designed that will be used for evaluation. In the following chapter, the conceptual system design is implemented prototypically using different technologies, i.e., Apache Solr (Java-based) and Python with the frameworks spaCy and Scispacy, sckikit-learn, and Keras/TensorFlow. The following evaluation covers the technical evaluation of the emerging knowledge extraction using a specifically designed evaluation strategy. Furthermore, a user-based evaluation of the system’s usefulness and usability is conducted. Also, an expert interview on the argumentation support’s outcome utilizing emerging knowledge is conducted. Overall, the evaluation concludes that the prototypical system is technically capable of extracting and utilizing emerging knowledge from medical document corpora using the hybrid approach of Natural Language Processing and Machine Learning on corpus related features. The user evaluation and the expert interview reveal that the system also fulfills users’ requirements regarding the support of emerging knowledge for medical Information Retrieval and argumentation. Hence, the conceptual design and the prototype could be used as an initial step for a real-world system. The thesis finishes with a summary of the contributions and an outline of future work.