Developing Efficient Custom Solr Plugin for LTR Feature Extraction

Perrin G. Bignoli · SSRN Electronic Journal · 2020

LTR (Learning to Rank) has emerged as a major source of search relevancy improvement for LexisNexis Advance (LNA). Initially, feature extraction for LTR was performed in part using ElasticSearch via AWS-hosted ESS domains. Although a great deal of work was spent customizing ESS to provide the required feature information, limitations imposed by AWS effectively limited LTR to reranking the top 20 search results. Experimentation showed that greater hDCG gains could be obtained by expanding the scope of LTR to rerank the top 50 search results, but the P90 execution time for those requests was 940ms in ESS. In order to produce a system capable of extracting the LTR features for the top 50 results of every LNA search, members of Team Europa, Team Saturn, and Team Titan partnered with the SRP Team to create a Solr-based system that was more performant than the ESS-based system. This talk describes in detail the evolution of Solr custom plugins that reduced the P90 execution time from 470ms to 150ms to 50ms in Solr. This involved the development of 3 progressively more efficient techniques were explored in Solr: 1.). a custom Term Vector Component, 2.) Function Sub-queries and custom Explorer Queries, and 3.) custom Value Sources.

Read the paper · More papers on PaperTik