Optimized Spoken Query Cross-Lingual Document Retrieval using BM25 and Neural Re-Ranking with AdamW
Anurag Kaushish, Aaditya Vijayvargiya, Kanika Rawat, Mithun Kumar, Vartika Jain · 2025
Cross-lingual document search from Hindi to English is challenging, mainly due to sparsity in training data, linguistic subtleties, and script differences. To improve this, we provide a hybrid system combining traditional retrieval methods with cutting-edge neural ones to counter Hindi-English language imbalance. Our method starts by converting spoken Hindi queries to text using speech recognition software, followed by English translation. Translated queries are employed to search a preindexed English document collection. For initial retrieval, we employ the Whoosh search engine with an adapted BM25F algorithm, which yields more accurate document content over titles to compensate for missing metadata. For re-ranking, we introduce a neural re-ranking model trained with AdamW optimization, which estimates semantic relevance between translated queries and retrieved documents. Large-scale innovations involve BM25F parameter tuning (b1 and k1) with Bayesian optimization over synthetic data - a workaround for the inaccessibility of labeled Hindi-English data. Our two-phase model surpasses the BM25 baseline and keyword-based methods, especially when vague queries or uncertain translations are encountered. Through enriching the context with lost senses and supplemental content, our method provides uniform performance, all entirely automated. Our light-weight method incorporates efficient indexing, retrieval optimization, and neural re-ranking harmoniously, intending to have efficient cross-lingual access for Hindi speakers. For future work, we aim to generalize this approach to other Indian languages and study how we can decrease our reliance on labeled data, making it deployable in resource-scarce settings.