ACM TechBrief: Automated Speech Recognition

Allison Koenecke, Niranjan Sivakumar, Jingjin Li, Shaomei Wu · ACM eBooks · 2025

ACM's Technology Policy Committees provide cutting-edge, apolitical, non-lobbying scientifc information about all aspects of computing to policy makers in the United States and Europe.To tap the deep expertise of ACM's 100,000 members worldwide, contact ACM's Global Policy Ofce at [email protected] +1 202.580.6555.Automated Speech Recognition (ASR) technology uses sophisticated machine learning algorithms -such as Generative Artifcial Intelligence (AI) -to understand and transcribe spoken language to text.ASR's reach is far beyond that of voicebased personal assistants like Siri or Alexa; it is now embedded to conduct virtual job interviews for hiring purposes, to surveil incarcerated people's phone calls to make further carceral decisions, and to record doctors' notes about a patient's health.Inaccurate transcriptions in these tasks can lead to a range of real-world consequences -from mild annoyances like having to repeat oneself over the phone to an automated voice assistant, to serious systemic issues like doctors making medical treatment decisions based on incorrect transcriptions.Furthermore, some speaker demographics are disproportionately affected by inaccurate ASR transcriptions, raising concerns about technological bias. ASR in High-Stakes Applications has Variable PerformanceTo convert spoken audio fles to written text transcriptions, ASR is used in high-stakes applications:• Medical: Over half a million doctors have used a Microsoft-owned AI scribe to transcribe patient visits [22].• Hiring: Over 60% of Fortune 100 companies use an ASR-based AI hiring tool, HireVue [32].• Policing & Criminal Justice: 80% of police offcers wear body cameras [36], and the leading US-based seller (Axon) offers an OpenAI-based ASR product that automatically flls out police reports using body camera audio [33].AI products also exist for audio-based prison surveillance [34] and court reporting [4]. Key ConsiderationsASR performance is generally measured by the Word Error Rate ("WER"), which is the share of words mis-transcribed by the ASR [13].A higher WER is worse; the lowest possible WER (0%) indicates perfect accuracy, while 50% suggests that roughly half of the words are transcribed incorrectly.Leading ASR systems report average WERs in the 5-20% range [23, 25,37].Below, we summarize fndings from audit studies conducted in the past 5 years, spanning ASRs (a superset of base models developed by Amazon, Apple, AssemblyAI, BlueJeans, Facebook, Google, IBM, Microsoft, OpenAI, Rev AI,

Read the paper · More papers on PaperTik