Distillation vs. Sampling for Efficient Training of Learning to Rank Models
Pooya Khandel, Andrew Yates, Ana-Lucia Varbanescu, Maarten de Rijke, Andy D. Pimentel · 2024
In real-world search settings, learning to rank (LtR) models are trained and tuned repeatedly using large amounts of data, thus consuming significant time and computing resources, and raising efficiency and sustainability concerns. One way to address these concerns is to reduce the size of training datasets. Dataset sampling and distillation are two classes of method introduced to enable a significant reduction in dataset size, while achieving comparable performance to training with complete data.