TIFD: Tibetan Instruction-Following Dataset for Large Language Models Supervised Fine-Tuning
Wenhao Zhuang, Dawa Cairen, Yuan Sun · Data Intelligence · 2024
In addressing challenges within the field of Natural Language Processing (NLP), supervised fine-tuning is an efficient technique that allows pre-trained Large Language Models to adapt to specific tasks. This is especially crucial for low-resource languages, such as Tibetan, where the demand for high-quality fine-tuning datasets is particularly pronounced. This paper introduces the Tibetan Instruction-Following Dataset (TIFD), comprising 11,535 JSON objects, each with four attributes: a unique identifier, instructions, input, and output. These attributes correspond to the task’s unique ID, the instruction describing the task, supplementary input for the task instruction, and the response generated by GPT-4. In the data preprocessing phase, unsuitable data for Tibetan fine-tuning were first filtered out. To ensure the diversity of instructions, the Tibetan instructions were effectively vectorized using LaBSE, followed by the selection of diverse Tibetan instruction data using a method based on the K-Center-Greedy algorithm. After preliminary English-Tibetan machine translation of responses generated by GPT-4, a web program was developed to facilitate Tibetan proofreaders in referencing the original English and corresponding Chinese translations for rectifying errors in words, grammar, and structure in the Tibetan data. Subsequently, five Tibetan language professionals were invited to manually proofread the data, checking and modifying the data format and Tibetan content, and removing erroneous and meaningless tasks, thereby ensuring the quality and reliability of the TIFD. Using TIFD to supervise the fine-tuning of the Tibetan large language model TiLamb has enhanced TiLamb’s Tibetan dialogue and instruction-following capabilities. The TIFD serves as a vital resource for Tibetan data, playing a significant role in the development of Tibetan NLP technology and in the effective customization of Large Language Models (LLMs).