Diartk: an open source toolkit for research in multistream speaker diarization and its application to meetings recordings
Deepu Vijayasenan, Fabio Valente · 2012
The speaker diarization task consists of inferring “who spoke when ” in an audio stream without any prior knowledge and has been object of several NIST international evaluation campaigns is last years. A common trend for improving performances has been the use of several different feature streams as diverse as speaker location features, visual features or noise robust acous-tic features. This paper describes an open source toolkit re-leased under GPL license aiming at facilitating research in mul-tistream speaker diarization and reproducing state-of-the-art re-sults. In contrary to other related diarization toolkits, it is ex-plicitly designed to handle an arbitrary number of features with very different statistics while limiting the computational com-plexity. The release includes a set of scripts to replicate bench-mark results on previous NIST evaluations and is intended to provide an easy to use software to study and include novel fea-tures into diarization systems.