Exploring the concept of multi-modality using grounding dino with evaluation of its performance in the context of different types and complexity of input text prompts
Abdul Rahim, Amanjot Singh, Ankur Sharma, Rachit Garg, Ch. Ravisankar, Nagendar Yamsani · 2024
We have known artificial intelligence, deep learning models can be trained with a certain type of input format to perform a task e.g., OCR models takes input in the form of image to read the text characters from the image, a weather prediction model can take weather data like wind speed, moisture, etc. as an input which is in text format to predict future weather data. But the multi-modality concept which is comparatively newer in the market paves the way for models to be built which can take more than one type of inputs e.g., text and audio, text and image, etc. In the same concept one of the popular architectures is specially designed for detection of any hindrance, model grounding dino about which we will be discussing in this setting. It is a zero-shot object detection model which can perform inference just by taking a text prompt input from the user describing what they want the model to detect in the provided image. In this study, we will be discussing about its capabilities, architecture and also its performance. We will be testing and quantifying its performance using standard self-created input text prompts with variations including description of the object, complexity of the object, etc. on a standard image.