Do modern speech LLMs and re-scoring techniques improve bilingual ASR performance for Basque and Spanish in domain-specific contexts?
Ander González-Docasal, Juan Camilo Vásquez-Correa, Haritz Arzelus, Aitor Álvarez, S. A. Moreno-Acevedo · Computer Speech & Language · 2025
This paper presents an extended evaluation of Vicomtech’s automatic speech recognition (ASR) systems developed for the Albayzín 2024 Bilingual Basque-Spanish Speech-to-Text (BBS-S2T) Challenge, a task focused on transcribing bilingual parliamentary recordings featuring frequent intra- and inter-sentential code-switching between Basque and Spanish. These recordings, drawn from Basque Parliament plenary sessions, pose significant challenges due to the abrupt language alternations, the limited availability of digital resources for Basque, and the absence of contextual and speaker information. The study incorporates additional analysis of state-of-the-art ASR architectures, namely Phi4-multimodal and CrisperWhisper, fine-tuned on the challenge dataset. Furthermore, the systems were evaluated on a complementary benchmark to assess model robustness. A detailed comparison of automatic hypothesis selection techniques, including both traditional n -gram and large language model (LLM)-based approaches, is also provided. Results demonstrate that optimal word error rate (WER) does not always correlate with the most accurate transcriptions, highlighting the complexity of evaluating ASR performance in code-switching scenarios. • Leading ASR systems have been assessed on a challenging bilingual domain. • Fusion and re-scoring frameworks were applied using statistical and large language models. • External language models prioritised meaning over transcription accuracy. • An extensive and dedicated analysis had been performed to further evaluate recognition errors. • Optimal WER may not reflect the top quality transcriptions in specific applications.