Patch-based visual microphone for improving quality of sound

Juhyun Ahn, Yong-Joong Kim, Daijin Kim · 2016

Visual microphone is a technique introduced to recover the sound from a silence video. And traditional method of sound recovery involves extracting and combining subtle motion signals from the entire image. However, there are two possible drawbacks of recovering the sound using the entire image. First, motion signals extracted from plain and edge regions may contain noise due to an aperture problem. Although plain regions are penalised by their squared amplitude values, summation of all the pixels present in the image could introduce notable amount of noise. Second, it is unclear which part of the surface of an object is hit by the sound wave. Utilizing only the region hit by the sound wave is expected to lead to a better sound recovery. The proposed patch-based visual microphone framework addresses these two problems by recovering the sound from a sub-region (patch) in the image centered at a key point (corner). Since we are unable to know which sub-region in the image is good for sound recovery, speeches are recovered from patches centered at each key point (corner), and then the best speech with the least noise is selected as a recovered speech. Extensive experiment results show that utilizing motion signals from a small region in the image near a key point can improve quality of the recovered speech.

Read the paper · More papers on PaperTik