LibriSQA: A Novel Dataset and Framework for Spoken Question Answering with Large Language Models
Zihan Zhao, Yiyang Jiang, Heyang Liu, Yu Wang, Yanfeng Wang · IEEE Transactions on Artificial Intelligence · 2024
While large language models (LLMs) have demonstrated impressive performance across various domains and tasks, they still struggle with multimodal tasks, particularly the spoken question answering (SQA) task, which requires precise alignment and deep interaction between speech and text. In this article, we address the SQA challenge by curating a novel free-form and open-ended SQA dataset, LibriSQA, which is composed of 214k SQA pairs covering a wide range of topics. It consists of two parts. Part I is designed for natural conversational formats and Part II focuses on multiple-choice questions with answers and analytical segments. Considering the limited availability of speech-text LLMs, we propose a lightweight, end-to-end framework to perform the SQA task on the LibriSQA dataset, achieving significant results. By transforming automatic speech recognition (ASR) into the SQA format, we further demonstrate the framework’s capability in handling ASR tasks. Our empirical findings support the idea that LLMs can effectively align and comprehend speech information, paving the way for the development of universal multimodal LLMs. Our LibriSQA dataset can be found athttps://github.com/ZihanZhaoSJTU/LibriSQA.