Improving Access to Engineering Education: Unlocking Text and Table Data in Images and Videos

Uchechukwu Uche-Ike, Lawrence Angrave · 2024

Accessibility of media, including visual media, is a significant concern in creating engineering education that is inclusive and accessible; without access, students who are blind or have low vision are unable to learn from today's engineering materials.This project presents a new accessibility tool, implemented as a user-friendly browser extension to extract and create accessible structured text -with a focus on tables of information -when embedded within web-based images and video media.The app can extract information to be used in a speech-to-text output of a screen reader; or text-to-braille device; pasted into a programming editing environment (e.g., Matlab, Jupyter notebook, Microsoft Code, or text editor); and further, used as the initial source input for manual or automated construction of audio descriptions to accompany the original media for future dissemination.In addition to improving access to engineering education, this project is a productivity tool for students seeking more access to textual data presented in image form.For example, it serves as a tool for all engineers and student engineers who seek to extract and re-use tabular information embedded inside an image or video that otherwise would require manual entry.The system uses the React.jsframework and Tesseract Optical Character Recognition (OCR) engine.The tool preserves privacy because it runs entirely inside the browser: no image data leaves the client.It can extract information from any web page on any website supported by the Chrome browser.Users can screenshot images and videos, extract text and numbers in scientific notation, determine the tabular structure, and meaningfully extract text from tables in tabular form (separated into rows and columns).We describe its design: the user interface features that allow its use by people with low-vision and access specialists.Upon initiation, the user selects a media HTML element to process; the user can choose to extract text from this frame or change the focus to another object on the page.The extraction area is adjustable by keyboard or pointer/touch interaction.Finally, the user chooses to either screen-capture the area and download the image, extract text from the area, or extract tabular information.To make the tool intuitive, interaction is structured as a wizard where users are guided along the stepwise process but can go back to previous steps.We provide examples of the best and worst output of our accessibility tool when applied to engineering education content and evaluate its accuracy and performance to extract tabular information from image samples from engineering disciplines.

Read the paper · More papers on PaperTik