Vision-assisted modeling for model-based video representations
Shawn Coniff Becker, V. Michael Bove · DSpace@MIT (Massachusetts Institute of Technology) · 1997
Abstract The objective of this thesis is to develop algorithms and software tools for acquiring a structured scene representation from real imagery which can be used for 2-D video coding as well as for interactive 3-D presentation. A new approach will be introduced that recovers visible 3-D structure and texture from one or more partially overlapping uncalibrated views of a scene composed of straight edges. The approach will exploit CAD-like geometric properties which are invariant under perspective projection and are readily identifiable among detectable features in images of man-made scenes (i.e. scenes that contain right angles, parallel edges and planar surfaces). In certain cases, such knowledge of 3-D geometric relationships among detected 2-D features allows recovery of scene structure, intrinsic camera parameters, and camera rotation/position from even a single view. Model-based representations of video offer the promise of ultra-low bandwidths by exploiting the redundancy of information contained in sequences which visualize what is largely a static 3-D world. An algorithm will be proposed that uses a previously recovered estimate of a scene's 3-D structure and texture for the purpose of model-based video coding. Camera focal length and pose will be determined for each frame in a video sequence taken in that known scene. Unmodeled elements (e.g. actors, lighting effects, and noise) will be detected and described in separate compensating signals.