Motion capture from multi image video sequences
Nicola D'Apuzzo · Repository for Publications and Research Data (ETH Zurich) · 1998
The process is composed of 5 steps: (1) acquisition of video sequences, (2) calibration of the system, (3) surface measurement of the human body for each frame, (4) 3-D surface tracking and filtering, (5) 3-D joints tracking. Photogrammetry and image processing techniques are used for these purposes. Our image acquisition system is currently composed of three synchronized CCD cameras and a frame grabber which acquires a sequence of triplet images. Before stepping into the processing phase, the camera system has to be calibrated. Self calibration methods (Maas, 1998) are applied to gain exterior orientation of the cameras, the parameters of internal orientation and the parameters modeling the lens distortion (Brown, 1971). From the video sequences, we extract two kinds of 3-D information: a three dimensional surface measurement of the visible parts of the body for each triplet and 3-D trajectories of points on the body. Our approach for surface measurement is based on multi-image matching, using the adaptive least squares method (Gruen, 1985). A semi automated matching process (D’Apuzzo, 1998) determines a dense set of corresponding points in the triplets, starting from few manually selected seed points. In case of poor natural texture, local contrast enhancement of the images is required for the least squares matching. The 3-D coordinates of the matched points are then computed by forward ray intersection using the orientation and calibration data of the cameras. A filter is applied to reduce the remaining noise in the 3-D data and get a more uniform density of the point cloud. Figure 1 shows an example of the surface measurement process using a pair of images.