TaoSR1: The Thinking Model for E-commerce Relevance Search
Chenhe Dong, Shaowei Yao, Pengkun Jiao, Jianhui Yang, Yiming Jin, Zerui Huang, Xuejun Zhou, Dan Ou, Haihong Tang, Bo Zheng · 2026
Query-product relevance prediction serves as a foundational technology in e-commerce search engines, enabling users to discover desired products and ensuring optimal user experience. Previous approaches have primarily relied on BERT-based models which excel at textual and basic semantic matching but demonstrate poor performance in understanding and reasoning capabilities for more complex queries. Consequently, numerous recent studies have explored the application of Large Language Models (LLMs) in search systems. However, most still adopt discriminative paradigms or ultimately distill knowledge to BERT models for deployment. In this paper, we propose an optimization framework based on large language models and directly deploy these models in online systems. Nevertheless, practical deployment presents several challenges, including online deployment, error accumulation in Chain-of-Thought (CoT) leading to performance degradation, and discriminative hallucination. To address these challenges, we propose an LLM-based optimization framework called Taobao Search Relevance Model v1 (TaoSR1) comprising three stages: (1) Supervised Fine-Tuning (SFT) with CoT to endow models with reasoning capabilities; (2) Offline multiple sampling based on a pass@N strategy, combined with Direct Preference Optimization (DPO), to enhance model generation quality; and (3) Difficulty-based dynamic sampling integrated with Group Relative Policy Optimization (GRPO) to further mitigate model's discriminative hallucination problems. Finally, by incorporating post-CoT processing and a relevance tier partitioning method based on cumulative probability, our model achieves more feasible and efficient online deployment. Experimental results demonstrate that our proposed model significantly outperforms baseline methods on challenging offline evaluation datasets, while achieving substantial improvements in online side-by-side human evaluations. Our proposed framework introduces a novel optimization paradigm for incorporating CoT reasoning into relevance classification tasks; we contend that this methodology provide valuable insights into the application of LLMs for other classification tasks.