How Reliable Is Semantic Search in Industrial Computing Domain? A Statistical Evaluation Pipeline
Elif Yozkan, Ilham Supriyanto · 2025
Statistical evaluation is crucial for improving semantic search, especially in highly technical domains such as industrial computing. Although qualitative human assessments are often used, they are time-consuming, expensive, and cover only a limited range of scenarios. A key challenge in conducting statistical evaluation is generating a realistic ground truth test set that accurately represents complex domain-specific terminology, including users’ persona and background knowledge. However, once this test set is established, it enables systematic experimentation to refine AI search configurations. Our study finds that increasing text volume improves search accuracy, leading to a precision increase from 77% to 82%. However, an excess of specialized terms, indicated by a relatively higher token-to-word conversion rate, can potentially weaken semantic understanding. To mitigate this, a hybrid approach that integrates semantic search with the traditional lexical method is utilized, with OpenAI embeddings to further enhance search performance.