Indonesian Educational Chatbot for Sports, Health and Physical Education using Retrieval-Augmented Generation

Viny Christanti Mawardi, Bambang Riyanto Trilaksono, Ayu Purwarianti · 2025

This study evaluates the effect of retrieval embeddings and similarity metrics at the generative component (LLaMA3). No fine-tuning or decoding variation was applied to the generator, ensuring that any performance differences were due to retrieval quality alone. This design enables a focused evaluation of how retrieval configurations influence the final output in a controlled RAG setup, particularly for primary school topics in sports and health (PJOK). We point out that RAG is a technique that helps large language models make fewer mistakes by using outside information, which improves the accuracy of facts and the relevance of the context. The research compares various embedding models and retrieval similarity metrics to evaluate their effectiveness in retrieving relevant information. The baai/bge-m3 and gte-multilingual-base models demonstrate strong performance with 0.98 precision, while the indo-bge-m3 model shows high recall but lower precision. From 5 embeddings, Llama Embedding performs significantly worse than the other models. The dot product has higher recall than Euclidean and cosine similarity metrics but embedding model selection significantly influences overall performance. The study underscores the importance of embedding choice and retrieval quality in improving chatbot responses. Further refinements are suggested, including better filtering, reranking retrieved chunks, and fine-tuning the retriever, or LLaMA, to align with the target language domain, such as Bahasa Indonesia. The results show that even though RAG improves how well things are generated, the quality of what is retrieved is still a major problem, which means we need to keep making improvements for the best results in educational uses.

Read the paper · More papers on PaperTik