A multi-module minimization neural network for motion-based scene segmentation

Yen-Kuang Chen, Sun‐Yuan Kung · 2002

A competitive learning network, called multi-module minimization (MMM) neural network, is proposed for unsupervised classification. Our objective is to provide a general framework to divide a set of input patterns into a number of clusters such that the patterns of the same cluster exhibit any pre-specified similarity measure (i.e. not limited only to RBF). As an example of non-RBF measure, let us look into a motion-based segmentation problem. The image frame can be divided into different regions (segments) each of which is characterized by a consistent affine motion. Algebraically, this leads to an LBF similarity criterion-because each region can be characterized by a 3-dimensional hyperplane. In order to apply the traditional RBF clustering techniques (e.g. VQ, k-mean), it requires a preprocessing step such as taking the Hough transform, which itself creates additional ambiguity. This problem is avoided in a direct approach such as the proposed MMM neural network. It allows us to directly cluster the tracked features into different moving objects by means of an LBF cost function. In general, the primary cost function should be carefully chosen to reflect the true physical model of the application. By minimizing the cost function, we can categorize a set of input patterns into a number of clusters. Because the primary similarity measure is no longer Euclidean type, it may become necessary to take spatial neighborhood into account as a secondary cost function. Still, a third cost function, reflecting the MDL type criterion, needs to be added so that noisy or spurious patterns will not be mistakenly modeled as a meaningful class. Accordingly, we have proposed an EM-type learning algorithm which uses all or part of the three cost functions mentioned above. A convergence proof for this algorithm is provided. Simulation results demonstrate that the MMM neural network does capture different motions and yield fairly accurate segmentation and motion-compensated frames.

Read the paper · More papers on PaperTik