Adapting Open-source Multimodal Large Language Models for Enhanced Vision Assistants
Md. Repon Islam, Muhammad Sheikh Sadi · 2024
This paper presents a novel wearable system designed to enhance the lives of visually impaired individuals by integrating advanced vision understanding through Multimodal Large Language Models (MLLMs) and Optical Character Recognition (OCR). An open-source 8 billion parameters MLLM has been deployed on a consumer-grade GPU-enabled server which excels in interpreting visual data and delivering intelligent responses. The wearable device, powered by an Orange Pi 5B single-board computer (SBC), incorporates advanced OCR, Automatic Speech Recognition (ASR), Text-to-Speech (TTS), and an on-device 0.5B Large Language Model (LLM) for real-time interaction. Text recognition is achieved using PaddleOCR lite, optimized through a series of algorithms for edge performance. The on-device LLM handles offline question-answer or summarization tasks, while server integration provides a more powerful, multimodal understanding of the environment, based on user queries. This hybrid approach balances low-latency, real-time processing with the computational demands of advanced multimodal models, offering a scalable, cost-effective solution. Experimental evaluations demonstrate character recognition accuracy of 95.55% for OCR and efficient response generation at 18 tokens per second for LLM on the edge. The deployed MLLM scores 9.1 out of 10 on a carefully curated subset of the VizWiz VQA dataset, which consists of real-world image-question pairs captured by blind individuals across the globe based on their actual needs. The system represents a significant advancement in assistive technologies for the visually impaired community, leveraging a hybrid computing framework and state-of-the-art LMMs for robust, real-time support.