Automatic Semantic Video Annotation
Kalaivani Anbarasan · 2018
The rapidly increasing quantity of publicly available videos has driven research into developing automatic tools for indexing, rating, searching and retrieval. Textual semantic representations, such as tagging, labeling and annotation, are used to represent appropriate semantics for search and retrieval. The semantics should be inspired by the human cognitive way of perceiving to describe videos. The difference between the low-level visual contents and the corresponding human perception is referred to as the ‘semantic gap’. Tackling this gap is harder in the case of unconstrained videos due to lack of semantics knowledge. Video based applications such as video surveillance, road traffic control, sports events detection require a strong human intervention when a semantic understanding of contents is needed to detect objects, actions or events within a video stream. Manual analysis of video sequences is a very time consuming task and it often leads to inaccurate results due to the -video blindness-. In the video surveillance domain, for example, it has been stimulated that an operator can miss up to 95% of scene activities after only 22 minutes of analysis. In the last years, great efforts by the computer vision research community leads to the development of robust and reliable algorithms for video analysis tasks at different levels: 1) Low-level video analysis methods address the ability to find the image regions corresponding to objects of interest (detection) and then track them across different frames while maintaining the correct identities (tracking). 2) Mid-level video analysis methods face the problem of recognizing simple or “atomic” events or activities. 3)High-level video analysis methods concentrate on the detection of “complex” events or activities