From CHAT towards ASR: A Hybrid Pipeline for Constructing the HUKILC-CO Hungarian Child Speech Dataset

Yue Luo, Xihang Xia, Kinga Jelencsik-Mátyus, Shaimaa Safaa Ahmed Alwaisi, Péter Mihajlik · 2025

Developing robust automatic speech recognition (ASR) systems for children remains challenging due to the scarcity of annotated datasets, particularly for underrepresented languages like Hungarian. This paper introduces HUKILC-CO, derived from the original HUKILC dataset, as the first Hungarian child speech corpus tailored for ASR. It isolates child speech from adult dialogue and non-verbal annotations while retaining environmental noise to reflect real-world conditions. We address key challenges, including noisy environments, overlapping speech, and CHAT formatting—via a hybrid alignment pipeline integrating Rev AI’s ASR and Batchalign2’s forced alignment. Our methodology combines fuzzy text matching and temporal buffering to synchronize transcripts with audio, achieving precise utterance- and word-level alignment. The dataset comprises 8.3 hours of high-quality child speech from 62 speakers (4.5–5.5 years), partitioned into speaker-independent training, validation, and test sets to ensure generalization. We further standardize labels by removing non-ASR annotations (e.g., corrections, symbols). HUKILC-CO bridges critical gaps in low-resource ASR research, enabling reproducible benchmarks for child speech technologies. The dataset and alignment pipeline are available for research to support educational and assistive applications.

Read the paper · More papers on PaperTik