A fuzzy approach to perceptual organization for object recognition
Ellen Walker, Hang-Bong Kang · 1993
This thesis discusses a fuzzy approach to perceptual organization for extracting structure from a single grayscale image. Perceptual organization generally refers to the human visual ability to find structure and groupings from image data. In the context of computer vision, it involves partitioning of image data and describing the associations among the various parts in terms of primitive image elements. For perceptual organization to be effective, it must be stable, regardless of small perturbations, and must explain the observed structures in the image data. Even though those requirements are developed well in some computer vision systems, there still exist some limitations in performing perceptual organization on image elements (or tokens). The limitations are: (1) image tokens are only classified into completely true or completely false with respect to given grouping criteria, (2) grouping is usually performed by bottom-up processing in finding structure, (3) an uncertainty computation method for evaluating extracted structures is not developed, (4) higher level groupings based on high-level geometric features are not exploited, and (5) a rich structural description is not maintained during grouping. To eliminate the limitations, our approach to perceptual organization is based on fuzzy set theory, which is the generalization of the usual Boolean logic. In our fuzzy approach, a grade membership value (or goodness value) with respect to given grouping criteria is measured for every image token. Using a goodness value, we can control the degree of approximation in grouping based on high-level knowledge. The control of the degree of approximation enables top-down processing in performing grouping. Furthermore, we derive a formula to compute the uncertainty in extracted geometric structures. The formula deals with unknown and incompatible parts between grouped image tokens, and an accidental viewpoint for grouped image tokens. The uncertainty value in extracted geometric structures will be useful in reliable higher level reasoning. In addition to those advantages inherited from fuzzy set theory, we add two features to our grouping system: a higher level grouping performed on primitive grouping results, and the maintenance of structural description. The higher level grouping is used for extracting complex geometric structures because most images usually have complex objects. The extracted complex structures will be useful in efficiently reducing search space for object recognition. Another feature is the maintenance of a rich structural description. All information acquired from grouping, including emergent features, the characteristics of image tokens, and the information about grouped tokens, is stored within a frame- based knowledge representation scheme. This rich information provides a good basis for knowledge-based image understanding systems. Furthermore, this makes our grouping system robust because re-grouping or un-grouping is possible when the grouping result is not desirable.