Improving Object Detecting by Structuring and Training YOLO Model
Zihao Zhou, Degong Gu, Yuancheng Shi, Huaqin Zhou, Kang Chen, Hongyi Qu, Hongyan Ren · 2024
Since OpenAI released large language models, artificial intelligence products have been progressing at an astonishing speed in recent years. A text-to-video generative AI model named Sora, which can create videos depicting realistic or imaginative scenes based on textual instructions, demonstrating the potential for simulating the physical world, shocked the world. However, the generated video may have some flaws especially in the relationships between objects. YOLO is one of the most famous objects detecting model in recent years. This paper will evaluate original YOLO model and discuss potential and practical solution for dealing with the shortage and how to improve the accuracy by using our object detecting framework. We first analyze the flaws in object detection in the Sora released video and the potential reasons by reviewing the background technology. Then, we create and trained new model for handling the limitations. Lastly, we will discuss the pros and cons of our model.