Arabic Speech-to-LLM Alignment via Adapter-Based Bridging: An Empirical Study of Grounding and Robustness

Haram Altaf · 2026

This study investigates Arabic speech-to-LLM alignment by adapting and evaluating the SparQLe paradigm in an Arabic-specific setting. Following the SparQLe recipe, the system connects a frozen speech encoder to a frozen instruction-tuned large language model through a trainable query-based bridge, enabling speech-conditioned text generation without relying on a conventional automatic speech recognition pipeline. In the proposed configuration, a frozen ArTST-v2 encoder extracts Arabic speech representations, which are processed by an ArBERT-based Q-Former and a projection layer before being passed to a frozen LLaMA-3-8B-Instruct decoder. Training is conducted in two stages: first, the bridge is optimized for speech--text alignment using contrastive, matching, and language-modeling objectives; second, the aligned representations are mapped into the language model through cross-entropy supervision while the encoder and decoder remain fixed. The evaluation emphasizes alignment quality and meaning preservation alongside transcription-oriented diagnostics, reflecting the goal of semantic speech understanding rather than only word-level reproduction. Results show stable convergence across both training stages, with a high semantic similarity and a positive speech-text alignment margin, showing that the learned adapter captures meaningful correspondences between Arabic speech and text. Across clean speech, cross-domain data, and code-switched inputs, the system demonstrates grounded Arabic generation from speech in a parameter-efficient setting. Overall, this thesis shows that Arabic speech can be effectively aligned with a frozen instruction-tuned LLM through an adapter-centric design, offering a practical foundation for Arabic speech-enabled applications and exploratory prompt-based tasks such as translation, summarization, and question answering.

Read the paper · More papers on PaperTik