Pairwise Learning to Rank for Search Query Correction
Antonín Novák, Ján Šedivý · 2013
This article introduces a new algorithm for a Search Query Spelling Correction System. It is based on learning to rank approach and allows to use large number of various signals leading to an improved accuracy. The performance will be tested against the conventional solution - the Noisy Channel Model. The new system was developed on a Czech Internet search query set, but the feature vector structure and the algorithm can be easily adapted for any other language when sufficient data is available. We will describe the algorithm details, the training and validation data sets. Further, we will discuss the selection and impact of the new feature vector signals.