Identification of Data-Intensive Systems Requirements using Semantic Similarity Search

Renita Raymond, S. Margret Anouncia · Journal of Engineering Science and Technology Review · 2022

The phenomenal growth of big data in social applications and IT software platforms over the last few decades has emphasized the significance of a systematic requirement engineering strategy for analyzing the requirements of dataintensive systems, that deliver valuable insights to business entities.Classification of data-intensive requirements can aid in the development of a more systematic and transparent requirements engineering process, resulting in increased requirement compliance and software project completion.As a result, this paper provides a unique approach Word2Vector based Fast Similarity Search (WV-FaSS) for improving the process of software requirement categorization for dataintensive systems in two phases.Word2Vec begins by taking a corpus of software requirements as input and producing a well-trained high-dimensional vector space.Following that, the vectors that were extracted semantically are indexed.The query vector is then used to look for the most similar vectors within the index, and lastly, similar documents are obtained.Experiments on two benchmark datasets, PURE and WARC, as well as a dataset from the private IT industry, demonstrated that our model outperformed state-of-the-art techniques with precision, recall, and F1 values of 0.91, 0.9, and 0.9, respectively.Thus, the proposed model WV-FAISS enables developers to rapidly search for embeddings of similar requirements that are similar to one another, while also increasing the scalability of similarity search methods.

Read the paper · More papers on PaperTik