IndoSBERT: Enhancing Indonesian Sentence Embeddings with Siamese Networks Fine-tuning
Kadek Denaya Rahadika Diana, Masayu Leylia Khodra · 2023
Sentence embeddings hold a crucial role in Indonesian NLP research, being applied in various domains such as essay scoring, summarization, text-to-image generation, text classification, and others. By this research, we present IndoSBERT, a modification of BERT that has been fine-tuned using the siamese network scheme inspired by SBERT. This model was fine-tuned with the STS Benchmark Dataset which was translated into Indonesian languange. Our model can provide meaningful semantic sentence embeddings for Indonesian sentences. IndoSBERT achieved higher scores in Semantic Textual Similarity task compared to existing and widely used models, such as IndoBERT by IndoNLU and IndoBERT by IndoLEM.