Combining Features at Search Time: PRISMA at Video Copy Detection Task.

Juan Manuel Barrios, Benjamín Bustos, Xavier Anguera · 2011

Most of current Video Copy Detection systems (VCD) perform a multimodal detection by dividing the system into subsystems. Each subsystem performs a copy detection using a different feature (either visual or audio), and the sets of candidates are combined (fused) to create the final result. We present a VCD system that fuses visual and audio descriptors at the similarity search level. The system produces the copy candidates by comparing video segments using visual and audio descriptors instead of fusing copy candidates from independent subsystems. We submitted four Runs to TRECVID 2011 CCD task: • PRISMA.m.balanced.EhdGry: a combination of two visual global descriptors. Two detection candidates per query. • PRISMA.m.balanced.EhdRgbAud: a combination of two visual global descriptors and one audio descriptor. Two detection candidates per query. • PRISMA.m.nofa.EhdGry: a combination of two visual global descriptors. One detection candidate per query. • PRISMA.m.nofa.EhdRgbAud: a combination of two visual global descriptors and one audio descriptor. One detection candidate per query. Our Runs achieve good detection effectiveness, especially for NoFA profile, and they are among the fastest Runs. To the best of our knowledge, this is the first VCD system that successfully fuses audio and visual descriptors at an earlier stage than decision level. Additionally, we have performed a joint submission with Telefonica Research team, under the name Telefonica-research.m.balanced.joint, which tests the combination at the decision level of Telefonica’s local descriptor, audio descriptor, and PRISMA’s EhdRgb global descriptors. 1.

Read the paper · More papers on PaperTik