Calculating the Similarity of Indonesian sentences using Latent Semantic Indexing based on KBBI

Muhammad Panji M, Arief Fatchul Huda · 2022 International Conference on Informatics, Multimedia, Cyber and Information System (ICIMCIS) · 2022

Calculating Semantic Similarity between sentences is a problem that is long discussed in the field of language processing. This paper presents two methodologies with semantic similarities using the traditional TF-IDF algorithm and LSI method of using TF-IDF with the term distribution of definitions. In this research, the data used comes from KBBI, in which each document is converted into tokens and gets definitions according to KBBI, then calculates the similarity of each definition. The accuracy value gained 75.9% for semantic similarity using the traditional TF-IDF and 80% for the LSI and TF-IDF methods using the term distribution weights.

Read the paper · More papers on PaperTik