Assessing Audio Hallucination in Large Multimodal Models
Sakuto Hanamaki, Namesa Kirishima, Sora Narumi · 2024
Speech recognition systems have become increasingly integral in various applications, from virtual assistants to automated transcription services, necessitating the development of models capable of accurately processing and transcribing spoken language. The introduction of multimodal models like ChatGPT-4 and Gemini 1.5 Flash represents a significant advancement in this field, yet challenges such as audio hallucination, pronunciation handling, and punctuation placement remain critical hurdles. This study provides a comprehensive evaluation of ChatGPT-4 and Gemini 1.5 Flash, focusing on their performance in accurately transcribing English audio inputs under varying conditions. By employing rigorous statistical and qualitative analysis, including metrics like Word Error Rate (WER) and Character Error Rate (CER), the study reveals that ChatGPT-4 exhibits superior accuracy and reliability in handling complex speech patterns. Detailed examination of pronunciation handling and punctuation accuracy further elucidates the specific areas where each model excels or faces challenges. The findings demonstrate the importance of continuous refinement and enhancement of multimodal models to improve their practical applicability and reliability in real-world scenarios. This research contributes valuable insights into the strengths and limitations of leading speech recognition technologies, providing a benchmark for future developments in the field.