Linguistically-Informed Dataset Curation for Efficient LLM Fine-Tuning: Balancing Performance and Efficiency

Sultan Alshamrani · IEEE Transactions on Sustainable Computing · 2025

The rapid growth of AI systems, particularly large language models, has raised significant concerns about their environmental impact due to excessive energy consumption and carbon emissions. Despite these concerns, the trend in AI development continues to prioritize performance gains through increasingly resource-intensive approaches, such as utilizing more powerful hardware and larger datasets, often at the expense of efficiency. This study contributes to the efforts towards more efficient AI by proposing and empirically evaluating two dataset curation strategies, DS-1 and DS-2, which are linguistically-informed by syntactic features like Part-of-Speech (POS) tags to prune lexical content, for fine-tuning LLMs. The focus of this work is the empirical demonstration of how these linguistically-motivated curation approaches can create a balance between computational efficiency gains and performance maintenance during the LLM fine-tuning process. This approach represents a promising step in addressing the critical issue of AI's environmental impact. The evaluation is conducted on sentiment analysis tasks, serving as a focused case study. We evaluate two curation approaches, DS-1 and DS-2, applied to sentiment analysis tasks across three datasets, achieving substantial dataset size reductions while preserving essential linguistic information. Our assessment of six state-of-the-art LLMs—RoBERTa, ALBERT, ERNIE, DeBERTa, BERT, and GPT-2—on both curated and original datasets reveals that curated datasets can yield comparable performance to uncurated ones, with efficiency gains of up to 70% in tokenization time and up to 80% in training energy. Notably, the study uncovers varying resilience of different model architectures to curation, with the GPT-based model demonstrating perfect performance adaptation, while other advanced models like ERNIE and DeBERTa also showed strong resilience. This research is crucial in promoting the development of more sustainable and resource-efficient NLP systems, challenging the prevailing notion that larger datasets invariably lead to better performance in fine-tuning LLMs, and paving the way for more environmentally conscious AI development practices. The generalizability of these specific curation strategies to other NLP tasks is an avenue for future investigation.

Read the paper · More papers on PaperTik