ODSS: An Open Dataset of Synthetic Speech
Artem Yaroshchuk, Christoforos Papastergiopoulos, Luca Cuccovillo · Zenodo (CERN European Organization for Nuclear Research) · 2023
ODSS is a multilingual, multispeaker dataset of synthetic and natural speech, designed to foster research and benchmarking of novel studies on synthetic speech detection. ODSS comprises audio utterances generated from text by state-of-the-art synthesis methods, paired with their corresponding natural counterparts. The synthetic audio data includes several languages, with an equal representation of genders. Natural and synthetic speech audio files within ODSS are released under the CC-BY-SA 4.0 license: Usage, extension and redistribution by the research community are strongly encouraged.