Mind the data gap(s): Investigating power in speech and language datasets

Nina Markl · 2022

Algorithmic oppression is an urgent and persistent problem in speech and language technologies.Considering power relations embedded in datasets before compiling or using them to train or test speech and language technologies is essential to designing less harmful, more just technologies.This paper presents a reflective exercise to recognise and challenge gaps and the power relations they reveal in speech and language datasets by applying principles of Data Feminism and Design Justice, and building on work on dataset documentation and sociolinguistics.

Read the paper · More papers on PaperTik