Shot-Level Camera Motion Estimation Based on a Parametric Model.

Ling‐Yu Duan, Jinqiao Wang, Yan-Tao Zheng, Changsheng Xu, Qi Chuan Tian, Jesse Sheng Jin, Hanqing Lu · 2005

This paper presents a system for shot-level camera motion analysis via the estimation of global motion from image sequences. An affine model is used to parameterize the global motion resulting from motion of a camera (e.g. pan, zoom, tilt, rotation). A singular value decomposition (SVD) based gradient descent is executed to minimize the prediction error. To assure the convergence of the gradient descent algorithm, we employ a modified n-step search matching to estimate initial translation and apply the gradient de-scent over a pyramid of input images. M-estimator is used to remove the influence of outliers derived from local object motion. The recovered parameters can be tied to the qualitative analysis of physically meaningful camera motion by simple transformation. Towards the shot-level characterization of camera motion, several simple rules are carried out to condense the series of transformed parameters extracted from a whole shot. Those rules mainly concern the magnitudes and temporal persistency of each meaningful parameter. We submit seven runs to TRECVID 2005. Their run-ids are D0, D1, …, D6. They use the same set of algorithms with different parameters tuned for the rules described above. Based on the results, we have the findings: 1) the design of rules and related parameters ’ tuning is significant but difficult to cover many cases from diverse scene contents, 2) various sizes of camera shots (e.g. close-up, medium shot, long shot) makes it infeasible to secure a proper parametric assumption of camera motion as the global motion in large amounts of video data, and 3) the context dependent training could improve application performance wherein the context information may include video genre, shot size, and shot categories according to a priori knowledge of a scene. 1.

Read the paper · More papers on PaperTik