CSRef: Contrastive Semantic Alignment for Speech Referring Expression Comprehension

Lihong Huang, Sheng-hua Zhong, Yan Liu · ACM Transactions on Multimedia Computing Communications and Applications · 2025

Referring Expression Comprehension (REC) aims to locate a target object in an image based on a natural language description. While existing REC methods primarily rely on textual input, spoken language remains an underexplored modality, despite its inherent naturalness and accessibility. To bridge this gap, we introduce a novel task, Speech Referring Expression Comprehension (SREC), which enables object localization using spoken language as input. To advance this task, we propose a new method, CSRef, alongside datasets and evaluation criteria tailored to SREC. CSRef integrates a global Contrastive Semantic Alignment (CSA) mechanism with the SREC framework, enabling direct extraction of semantic information from speech for visual grounding. This approach streamlines the semantic processing pipeline and reduces complexity compared to conventional methods that depend on automatic speech recognition followed by textual REC. We conduct extensive experiments on three widely used REC datasets extended with synthetic speech, as well as three constructed face-centric SREC datasets. The results demonstrate that CSRef outperforms transcription-based baselines in both efficiency and accuracy. Furthermore, we evaluate CSRef in a downstream application, language-guided face blurring, and compare it with the MLLM-Guided Image Editing (MGIE) approach. CSRef achieves superior performance, both in the precision of region modification and in the overall quality of the face-blurring effect. These findings establish CSRef as an effective and scalable solution for SREC, representing a promising step toward speech-based visual grounding in real-world human-computer interaction scenarios. The code is publicly available at https://github.com/macrorise-lh/CSRef .

Read the paper · More papers on PaperTik