Extracting Software Design from Text: A Machine Learning Approach

Islam Elmasry, Khaled Tawfik Wassif, Hanaa Bayomi · 2021

Extracting software design components from text is a goal of many research works. Utilizing Machine learning techniques for that goal can improve the accuracy of this process instead of using traditional Natural Language Processing (NLP) methods. In this paper, the proposed approach uses machine learning techniques to extract classes and attributes from plain text (e.g. software requirements documents) through two consequent classifiers, the first one classifies each word into a class or not using predefined features, then, the second classifier starts to classify words into an attribute or not. Finally, dependency parsing is used to define a set of rules applied to the document given the extracted classes and attributes to relate the attributes to the classes. One of the contributions of this paper is the created dataset in its final pre-processed form that could make it easier to use in the software design field in the future.

Read the paper · More papers on PaperTik