Early vision using distributions
Carlo Tomasi, Mark A. Ruzon · 2000
For over thirty years computer vision researchers have been proposing methods for “early vision” tasks such as detecting edges and corners. One key assumption shared by most previous methods is that image neighborhoods are constant in color or intensity, with deviations modeled as noise. Due to computational considerations that encourage the use of small neighborhoods where this assumption holds, these methods remain popular. This research models a neighborhood as a distribution of colors. Our goal is to show that the increase in accuracy of this representation translates into higher-quality results for early vision tasks on difficult, natural images, especially as neighborhood size increases. We emphasize large neighborhoods because small ones often do not contain enough information. We emphasize color because it subsumes greyscale as an image range and because it limits the number of valid models we should consider; using only greyscale images allows assumptions that do not hold for color. We start by developing the compass operator, a color edge detector that computes the orientation of the diameter of a circle that maximizes the distance between two color distributions. This distance is computed by using the earth mover's distance, which finds the minimal amount of work needed to transform one distribution into another. We continue with color corner detection, a generalization of color edge detection in which the two sides no longer have the same size. Extracting corners is more involved than extracting edges because multiple responses to the same corner are not allowed. We show that corners are dependent on edge evidence and, therefore, an edge model is required to confirm the existence of a corner. Finally, we extend blue screen matting, a technique that extracts an object from a constant color background, to backgrounds that are almost arbitrary. Given coarse knowledge of the boundary between two objects, we compute the color distributions on either side of a potential boundary pixel and estimate alpha, the proportion in which a color from each side mixed to form that pixel's color. As a result, a user can more easily move objects from one image to another while maintaining photorealism.