Calculating Requirements Similarity Using Word Embeddings

Sandeep Reddivari, Jeffery Wolbert · 2022 IEEE 46th Annual Computers, Software, and Applications Conference (COMPSAC) · 2022

Finding similar requirements from a requirements dataset is an important problem in requirements engineering (RE). Automatic processing and representation of requirements documents has become a necessity for requirements engineers (REs) due to the inability of one person (or a even small group of people) to completely oversee the requirements of a system manually because of the system's scale. This paper outlines a novel framework which can help REs to realize document similarity using word embeddings. This framework takes a collection of documents as input and produces a list of similarity ratings for each document. The similarity ratings provide REs an intuition of what could be related or how the requirements search space can be reduced for further RE activities. The framework uses TF-IDF to produce a list of important terms for each document, filters out totally unique terms, then uses spaCy's deep learning word vectorization to calculate similarity between requirements documents.

Read the paper · More papers on PaperTik