The Unintended Irregularities of Automatic Speech Recognition
Silja Vase · Proceedings of the World Congress on Electrical Engineering and Computer Systems and Science · 2020
The present study examines the emerging role of Automatic Speech Recognition (ASR) and the unintended irregularities that arise when the algorithm is configured in healthcare practices.Once you consider health information technology, the mind often drifts towards new digital devices or tools that are implemented to improve quality or create efficiency.Scholars within Human-Computer Interaction (HCI) reflect widely upon how digital technology has become a vital part of healthcare exploring the Internet of Things (IoT) [1], visually directed applications [2] and usability frameworks [3] to name a few perspectives.However, this study focusses on ASR, an algorithm that is used by physicians to conduct Electronic Health Records (EHRs), which allows the patients to access their records in real-time [4].ASR substitutes dictation and typing when physicians produce these medical records.The central objective is to understand how physicians navigate in relationship to ASR and how the algorithm affects practices.The study is based on ethnographically collected data at a Danish hospital in Jutland, where the American developed Nuance SpeechMagic algorithm was implemented approximately ten years ago and has become a mandatory medium for physicians when they conduct EHRs.The study is based on the data collected at a multi-sited study that starts at an orthopedic department and follows the algorithm around through several departments and ends at the supplier.In short, ASR enables natural language to be translated into text, which subsequently can be edited by physicians to train the algorithm for a better translation in the future [5].The specific algorithm studied is based on a linear algorithm, trained by corrections made by physicians who represent diverse practices in various fields.The algorithm is thus expected to cover the disciplines and semantic representations within these when documenting EHRs.Further, ASR is expected to comprehend interferences or malfunctions of other technologies, e.g., phone calls, alerts, 'frozen' accounts or screens, uneven connection, timeworn microphones, etc.The study shows how the performance measurements visualized by ASR do not reflect the physicians' use of it, which is demonstrated as practice-related challenges.The challenges emerge from 1) lack of technical literacy, 2) interference made by the surroundings, or 3) technical malfunction.This study shows how the present algorithm does not meet expectations made by physicians towards the technology and how it further fails to deliver a reduction of time spent on producing medical records.What is more, individual physicians need to spend more time training the algorithm than others, as their speech does not fit into predetermined gender-related categories, which in turn complicates their time spent on patients.The findings further meet concerns stated by previous voices in the field, such as how non-native speakers can experience a lower recognition of their voices [6] and how noise also distorts the recognition ability [7].Upcoming deep learning ASR algorithms are based on the current design and will continue to produce unintended irregularities.As this study is based on a linear ASR algorithm, future studies should thus focus on the cognitive aspects to research significant implications for future healthcare practices.