Breaking Down Barriers: Next-Generation Techniques for Segmenting Medical Abstract Text using DeBERTa-V3
Ankit Anand, Syed Wali Ahmad Rizvi, Sanjiith Ravindhran, Rishi Ajith, Gokula Kumaran K G, Divij Goyal · 2024
The readability of medical research paper abstracts is often compromised by their dense, complex language and format, which consolidates key information into a single, challenging paragraph. This study introduces a novel application of Natural Language Processing (NLP) techniques to enhance the segmentation and readability of these abstracts, making them more accessible and skimmable for rapid review. Utilizing the advanced capabilities of PyTorch, we developed an NLP model leveraging the pretrained DeBERTa-v3 Base model from HuggingFace to segment medical abstracts into discrete, intelligible units. This approach was inspired by the architecture discussed in the State of the art Techniques(SOTA’s) which provided foundational insights into sentence classification within medical texts. Our model not only segments text but also maintains the logical and semantic continuity essential for understanding medical findings. Preliminary results demonstrate a significant improvement in the readability of segmented abstracts compared to their unsegmented counterparts when compared to Universal Sentence Encode(USE). This paper details the development process, the comparative analysis with USE, and the potential implications of this enhanced readability in medical literature.