FLAME: Self-Supervised Low-Resource Taxonomy Expansion Using Large Language Models
Sahil Mishra, Ujjwal Sudev, Tanmoy Chakraborty · ACM Transactions on Intelligent Systems and Technology · 2024
Taxonomies represent an arborescence hierarchical structure that establishes relationships among entities to convey knowledge within a specific domain. They find utility in various real-world applications, such as e-commerce search engines and recommendation systems. Consequently, there arises a necessity to enhance these taxonomies over time. However, manually curating taxonomies with neoteric data presents challenges due to limitations in available human resources and the exponential growth of data. Therefore, it becomes imperative to develop automatic taxonomy expansion methods. Traditional approaches encounter difficulties stemming from limited resources, primarily due to the small size of existing taxonomies . This scarcity of training data often leads to over-fitting. In this article, we propose Fine-tuning large language Models for taxonomy Expansion ( FLAME ) , a novel approach for taxonomy expansion in low-resource environments (i.e., limited size of existing taxonomies, lack of robust representation capabilities of pre-trained language models) by harnessing the capabilities of large language models (LLMs) that are trained on extensive real-world knowledge. LLMs help compensate for the scarcity of domain-specific knowledge . Specifically, FLAME leverages prompting in few-shot settings to extract the inherent knowledge within the LLMs, ascertaining the hypernym entities within the taxonomy. Furthermore, it employs reinforcement learning to fine-tune LLMs, resulting in more accurate predictions. Experiments on four real-world benchmark datasets demonstrate the effectiveness of FLAME in real-world scenarios, achieving a remarkable improvement of 12.8% in accuracy and 5.6% in Wu & Palmer metric over 11 baselines. Furthermore, we discuss the strengths and weaknesses of FLAME through an extensive case study, error analysis, and ablation studies on the benchmarks.