Analysis of Various Measures of Text Similarity for Comparing Topics of Computer Science Syllabuses

Ritu Sodhi, Jitendra Choudhary, Ritu Jain, Ruby Bhatt, Ritesh Joshi, Anil Patidar · International Journal of Engineering Trends and Technology · 2024

Text similarity measures are used to find out how much different texts are similar. There is a need to compare text for document comparison, text classification, text summarizing, information retrieval, question-answer sessions, clustering documents, etc. There is also a need to compare computer science terms; while plagiarism checks, website contents, comparing syllabuses of the same subject, notes, books, etc. This research focused on the text similarity measures to compare text related to computer science terms. This research executed some of the lexical and semantic similarity measures for comparing topics of the syllabus of programming using Python. And found after executing various approaches that spacy using a large English model and cos_similarity together gives a better result. In the future, this research can be improved by including more similarity measures and by increasing the size of the dataset for comparison of computer science terms.

Read the paper · More papers on PaperTik