Floras 50: A Massively Multilingual Multitask Benchmark for Long-Form Conversational Speech
William Chen, Brian Yan, Chihchen Chen, Shinji Watanabe · 2024
A common criticism for current speech recognition benchmarks is the reliance on settings which do not generalize well to real-world conversational environments, such as read speech and pre-segmented utterances. These issues are more apparent in multilingual benchmarks due to the expenses of annotation, which raises the difficulty of building practical speech technologies for more languages. This paper presents the FLORAS 50 evaluation set, the first massively multilingual longform speech processing benchmark for automatic speech recognition (ASR), speech translation (ST), and speech summarization (SSUM). FLORAS contains 32000 hours of YouTube audio mined from the YODAS ASR dataset, and thus contains a diverse array of recording conditions and speaking styles across 50 languages. Each recording in FLORAS has a minimum duration of 5 minutes, which encourages the development of new methods that are both accurate and memory efficient. To expand the task coverage of FLORAS to long-form multilingual ST and SSUM, we also introduce a scalable human-in-the-loop pseudo-labeling method with Large-Language Models. As we find that current model architectures are either cannot handle the long sequences in FLORAS or capture long-form dependencies well, we propose the LongBranchformer architecture that can efficiently model both local and global relationships. We establish baselines on FLORAS using the LongBranchformer and state-of-the-art pretrained models like Whisper, showing the limitations of current techniques in long-form settings. The dataset is available at https://huggingface.co/datasets/espnet/floras.