Extracting Triplets from Domain-specific Texts using Fine-Tuned LLaMA Models: A Comparative Study

G. S. Veena, Aparna Rajendran, Deepa Gupta · Procedia Computer Science · 2025

This research centers on extracting triplets (Subject, Verb, Object) from domain-specific texts, particularly focusing on the film industry. We compiled a corpus of 4,300 sentences from 500 well-known movie articles and Wikipedia entries. The study targeted five key relations: Composed by, Written by, Released on, Made f rom, and Released in , along with their respective entities. Dependency parsing was employed to extract triplets from this dataset. These extracted triplets were then used for supervised fine- tuning of the LLaMa model. We assessed the evaluation loss across different versions of the LLaMA models and found that LLaMA 3 outperformed the others, showing the best performance for this task. For testing, the fine-tuned LLaMA model is used to extract triplets from unseen sentences. The extracted triplets can be used to create knowledge graphs, offering structured representations of information by mapping relationships between entities.

Read the paper · More papers on PaperTik