Drop Noise For Cleaning LLMs Data

Tong Guo · 2025

In industry NLP application, our dataset by prompting large language models (LLMs) has a certain number of noise data. We present a simple method to find the noise data and remove them. We retain the data that contains certain common tokens between the LLMs data and the prediction results of a generative model trained on the LLMs data. We remove the data that does not contain certain common tokens between the LLMs data and the prediction results of a generative model trained on the LLMs data. We adopt T5-Base as our generative model. The experiment result shows our method is highly effective and does not require any manual annotation. For industry deep learning application, our method improves the NLP tasks accuracy from 88% to 98% under human evaluation, meanwhile the LLMs data source is sufficiently abundant.

Read the paper · More papers on PaperTik