Transformer-based Semantic Classification of Sanskrit Compound Words
Shriganesh Devaru Bhat, B. Premjith · 2025
Compound word, a semantic combination of two nouns frequently encountered in everyday discourse, literature, articles, and scriptures. The linguistic analysis of compound words (CWs) in text is emphasized as determining the semantic class of a CW enhances clarity, since CWs convey distinct meanings based on their semantic class and contain essential information that makes a sentence semantically complete. Compound analysis involves multiple computational operations, including segmentation, parsing, type identification, and paraphrasing. The computational semantic classification of CWs in a low-resource language that includes Sanskrit is very challenging because of restricted resources, morphological complexity, and the language’s productivity. This study aims to enhance compound analysis by using the benchmark UoHyd dataset, previously used in compound analysis research, to implement a fine-tuned LLM approach for the semantic classification of Sanskrit CWs into six primary categories: adverbial, determinative, oppositional determinative, exocentric, coordinative, and others. This study involved the fine-tuning of five BERT variants: San-BERT, San-ALBERT, San-RoBERTa, DistilBERT, and IndicBERTv2. Indic-BERTv2 outperformed four other models, with an accuracy of 94.49% and an F1 score of 94.48%. The research findings underscore that, while adhering solely to the minimum requirements and without any other linguistic components, LLMs surpass methodologies that include additional linguistic aspects, effectively capturing the nuances of semantics in Sanskrit CWs.