Linear Global Translation Estimation with Feature Tracks

Zhaopeng Cui, Nianjuan Jiang, Chengzhou Tang, Ping Tan · 2015

Global structure-from-motion (SfM) algorithms register all cameras simultaneously, which are potentially more efficient and less prone to drifting than incremental SfM methods. Global SfM methods often solve the camera orientations and positions separately. This paper focuses on the problem of global position (i.e. translation) estimation. Essential matrix based global translation estimation methods (e.g. [1]) usually degenerate at collinear camera motion because the translation scale is not determined by an essential matrix. Trifocal tensor based methods (e.g. [3]) usually rely on a strongly connected camera-triplet graph, where two triplets are connected by their common edge. The 3D reconstruction will distort or break into disconnected components when such strong association among images does not exist. The recent 1DSfM method [4] designs a smart filter to discard outlier essential matrices and solves scene points and cameras together by enforcing orientation consistency. However, this method requires abundant association between input images, e.g.∼O(n2) essential matrices for n cameras, which is more suitable for Internet images and often fails on sequentially captured data. The data association problem of [4] and [3] is exemplified in Figure 1. The Street example on the top is a sequential data where each image is only matched upto 4 neighbors. 1DSfM fails on this example due to insufficient image association. In the Seville example on the bottom, those Internet images are mostly captured from two viewpoints (see the two representative sample images) with weak affinity between images at different viewpoints. This weak data association causes seriously distorted reconstruction for the triplet-based method in [3]. This paper introduces a direct linear algorithm to address the presented challenges. It avoids degeneracy at collinear motion and deals with weakly associated data. Our method capitalizes on constraints from essential matrices and feature tracks. As shown in Figure 2 (a), the location of a scene point p can be computed as the middle point of the mutual perpendicular line segment AB of the two rays passing through p’s image projections:

Read the paper · More papers on PaperTik