Transformer-based Models for Visual and Textual Data Extraction

Meet R. Patel, Manisha J. Nene · 2025

In an era driven by digital transformation, the demand for intelligent and automated data processing has increased significantly. This research explores a unified approach to extracting both visual and textual information from images, enabling better analysis and interpretation in multiple domains. As a case study, we focus on CCTV surveillance, where traditional monitoring methods are often limited to object detection and manual inspection. Our proposed framework enhances this process by generating detailed image descriptions using transformer-based models, extracting multilingual textual data through Optical Character Recognition (OCR), and translating it into English for broader accessibility. By unifying these components, the frame-work offers a more comprehensive understanding of surveillance footage, improving both real-time monitoring and post event investigations. Fine-tuned on a domain-specific CCTV dataset, the system surpasses smaller state-of-the-art models in generating accurate and context-rich captions. Additionally, its lightweight design allows for deployment on edge devices, making it suitable for real-time monitoring and security applications without relying on high-end computational resources. This research demonstrates the potential of multimodal data extraction to improve situational awareness and decision making in security and other real-world scenarios.

Read the paper · More papers on PaperTik