Data Curation for NLP Corpora

Rachel Wagner-Kaiser, Tim Cerino · 2025

This chapter covers guidance for identifying and selecting the appropriate datasets to train effective models. The datasets created for training, validation, and testing will depend on the data distribution and corpus characteristics, with a focus on identifying and capturing as much variation in the corpus as possible as part of model training. This chapter also discusses how data selection impacts training, validation, and testing, and why it is important to align the dataset with the goal of the solution.

Read the paper · More papers on PaperTik