A Chinese Expressive Long-dialogue Speech Dataset with Scripts

Jin Li, Tianrui Wang, Meng Ge, Chenrui Cui, Jiahui Zhao, Jianrong Wang, Longbiao Wang, Jianwu Dang · 2025

With the advancement of large-scale models, the demand for emotionally rich, long-context, and highly natural communication in human-computer interaction increases. However, the exploration of long-context or script-level speech conversation tasks remains limited due to the lack of specific supervised data. To address this, we introduce a three-stage data processing pipeline for creating a Chinese expressive long-dialogue speech dataset with scripts (CELSDS). We collect videos from TV series, manually annotate speaker information for each character, apply Optical Character Recognition (OCR) to extract speech content, annotate episode summaries, and use a large language model (LLM) to generate sentence-level scenario descriptions. To our knowledge, this is the first Chinese long-context dialogue dataset that incorporates speaker and content annotations, script-level episode summaries, and sentence-level scenario details. Using this dataset, we develop a baseline model for both speech-to-script and script-to-speech generation tasks. The annotations and data production code are open-sourced at: https://github.com/lijin0120/CELSDS.

Read the paper · More papers on PaperTik