Statistical and computational theories for image segmentation, texture modeling and object recognition

David B. Mumford, Song Chun Zhu · 1996

Presented in this thesis are the statistical theories and computational schemes for three fundamental problems in computational vision: image segmentation, texture modeling, and two dimensional object recognition. Though these problems can be studied separately, they, like many other vision problems, depend on each other, therefore if our objective is to build a consistent and sophisticated vision system, the solutions to these problems should obey a common theory. The first chapter of this thesis proposes a general unified theory and a computational framework for solving vision problems. Lying at the core of this unified theory is a pyramidal description scheme, where various visual concepts are described by random variables, continuous or discrete, in a hierarchic structure. Visual computation is then posed as a statistical inference problem according to the Bayesian theory. It suggests that given the observed images, we should be able to infer the random variables of various levels all together. This computational scheme automatically incorporates concepts like the multiple intermediate solutions and the top-down/bottom-up loop. Guided by this unified theory, an algorithm for image segmentation, called region competition, is proposed in chapter 2, This algorithm is derived by minimizing a generalized Bayes/MDL criterion using the variational principle. It combines the most attractive features of the existing algorithms such as snakes/balloons and region growing, and is also related to edge detection using filters. Hence addresses the existing 4 kinds of approaches to image segmentation as different aspect of the same problem. Then in chapter 3, a unified theory called FRAME (filters, random fields and maximum entropy) is proposed for texture modeling, and it combines filtering theory and random field modeling through the maximum entropy principle. It interprets many previous concepts and methods for texture analysis and synthesis from a unified point of view. A novel strategy for filter pursuit and probability approximation is also proposed in chapter 3. Chapter 4 reports a flexible object recognition and modeling system (FORMS) which represents and recognizes objects from their contours. This consists of a forward model for generating the shapes of animate objects which gives a formalism for solving the inverse problem-object recognition in a topdown/bottom-up loop. In FORMS, the forward model includes principal component analysis for shape deformation and stochastic grammar for shape generation. The inverse process includes a novel method for skeleton extraction and part segmentation.

Read the paper · More papers on PaperTik