Object Identification Using Generative AI: An Application of Computer Vision
Shekharesh Barik, Manisha Behera, Pooja Kumari, Soumya Sanjit Mohapatra, Tapas Behera · 2024
The exponential growth of image data in recent years, driven by the widespread adoption of digital photography and the proliferation of social media platforms, presents a pressing challenge: how to efficiently identify objects within this vast sea of visual information. Traditional methods of object identification often struggle to cope with the scale and complexity of modern image datasets. Generative AI, a subset of machine learning, focuses on creating models that can generate data resembling real examples from a given domain. The most prominent approaches within Generative AI which are Variational Autoencoders and Generative Adversarial Networks have shown remarkable capabilities in generating images which are realistic and capturing complex data distributions, making them promising candidates for improving object identification tasks. Object detection has undergone significant evolution in algorithms, aiming to enhance both accuracy and speed. Thanks to the relentless numerous researcher’s effort, deep learning algorithms have swiftly advanced, resulting in improved object detection capabilities. These advancements find widespread application in various domains such as medical imaging, pedestrian detection, self-driving cars, and face detection, robotics among others, thereby streamlining tasks previously reliant on human effort. Given the expansive nature of the field and the plethora of state-of-the-art algorithms, comprehensively covering them all at once proves to be a daunting endeavor. This paper seeks to provide a foundational overview on the methodologies of object detection, focusing on the detection stages. In the realm of two-stage detectors, algorithms like RCNN, Fast RCNN, and Faster RCNN are explored, prioritizing accuracy. Conversely, one-stage detectors such as YOLO versions emphasize speed.