Multimodal AI For Image And Video Inference: Generate Text Output From Visual Inputs
Sumayya M Kadampur, Meenaxi M Raikar, Vishwanath P. Baligar · 2025
Multimodal AI unites computer vision with large language models to process and produce information in multi-modality fashion: images, videos, and texts. The work takes the advancement in Large Multimodal Models (LMMs) towards CLIP-ViT-L-336px having MLP projection coupled with timestamp-aware encoders so that challenges towards scalability, compositional reasoning, and temporal dynamics can be better overcome. Vision-Language connectors and task-specific datasets reach state-of-the-art results in image tasks, whereas sliding Q-Formers and timestamp binding enhance temporal localization and dense video captioning. Efficient training on publicly available datasets ensures scalability, which can be applied to video summarization, temporal grounding, and visual question answering. This work establishes a scalable foundation for real-world multimodal AI applications.