Scene understanding from image and video: Segmentation, depth configuration and inpainting
Maria Oliver Parera · Dialnet (Universidad de la Rioja) · 2018
The visual information we extract from our perception of the world is formed by a continuum of objects interacting among them, instead of the small particles that form it. Objects play a crucial role in our understanding of the environment: we perceive them as the single entities that allow us to interact with our surroundings. When we take a picture or film a video we want to capture the reality we are observing. In this way, despite we store the information in pixels, we look for the objects conforming it. That is, we look for the identifiable portions of the image that can be interpreted as single units. Obviously, these basic units will change depending on the application. For example, if we are working on face recognition we will be interested in the eyes, mouth and nose of the face, but, if we work in people tracking we would be more interested in considering the whole person as a single unity. Changing from the pixel unity to objects level has a lot of advantages in many areas, not only in computer vision but also in robotics or industrial engineering. For instance, despite robots see the world through sensors that receive pixel-data, we expect them to perceive the surrounding world in the same way we do. We need the robots to be able to recognize and interact with the objects conforming the scenes. Besides, objects have a number of attributes such as volume and shape, texture or interrelation properties such as adjacency or T-junctions. These attributes make objects to be richer instances than individual pixels, which can help in a classification process. This manuscript is focused on working from the object level perspective. We propose a segmentation model to decompose the image scene into regions or shapes. Then, we propose to solve two other problems which have as input an image or video classified in objects or shapes. Image Segmentation consists in partitioning the image into regions that share common features, such as color or texture. To this goal we propose a variational method that considers adaptive patches to characterize, in an affine invariant way, the local structure of each region of the image. The patches are computed using an affine covariant structure tensor defined at every pixel of the image domain, so that they can automatically adapt its shape and size. The proposed segmentation model uses an $L^1$-norm fidelity term and the total variation of relaxed fuzzy membership functions as an approximation to the length of the boundaries of the segmented regions. The output of the method is a partition of the image in regions together with a patch representative of the texture of each region. The Scene Structure problem involves the recovery of the relative order structure of the objects from a planar image, where some objects may occlude others. We propose to estimate the interpretation of the scene by integrating some global and local cues while also providing both the complete disoccluded objects that form the scene and their ordering according to depth. Our method first computes several distal scenes which are compatible with the proximal planar image. To compute these different hypothesized scenes, we propose a perceptually inspired object disocclusion method, which works by minimizing the Euler's elastica as well as by incorporating the relatability of partially occluded contours and the convexity of the disoccluded objects. Then, to estimate the perceptually preferred scene we rely on a Bayesian model and define probabilities taking into account the global complexity of the objects in the hypothesized scenes as well as the effort of bringing these objects in their relative position in the planar image. Inpainting is the problem of recovering an image or video that is partially damaged. It is also used to remove undesired objects from the image or video and recovering the occluded objects. In video inpainting the missing information in the frames produces an incomplete optical flow. Therefore, video inpainting involves an extra challenge: recovering the optical flow. First, we propose a variational model for the completion of moving shapes through binary video inpainting that works by smoothly recovering the objects into the inpainting hole, taking into account the optical flow and motion occlusions. We solve it by a dynamic shape analysis algorithm based on threshold dynamics. The resulting inpainting algorithm diffuses the available information along the space and the visible trajectories of the pixels in time. Finally, we present an automatic method for optical flow inpainting. Given a video, each frame domain is endowed with a Riemannian metric based on the video pixel values. The missing optical flow is recovered by solving the Absolutely Minimizing Lipschitz Extension (\textsc{amle}) partial differential equation on the Riemannian manifold. An efficient numerical algorithm is proposed using eikonal operators on finite graphs for nonlinear elliptic partial differential equations. MANUSCRIPT OUTLINE This document is organized in three parts. Each of them is devoted to analyze one of the problems. In Part I we approach the segmentation problem. In Chapter 1 related works to the segmentation problem are reviewed, together with a motivation to use patch-based methods comparison. Chapter 2 is devoted to introduce the affine invariant tensors that are used to compute the patches and perform the segmentation. In Chapter 3 we explain our model, which uses the distance patch measure presented in Chapter 2, to decide which region each pixel belongs to. The optimization algorithm is also explained in Chapter 3. Chapters 4 and 5 are respectively dedicated to the results and conclusions of the proposed method. Finally, as our algorithm involves the computation of the vector median, Appendix A is dedicated to the explanation of a fast algorithm to compute it. Part II is devoted to the method that computes the most probable scene structure of an image. In Chapter 6 we do a review of the psycophysical studies that are involved in the image recovery of the 3D information given a single image. As our model also proposes to disocclude the occluded objects we also do a review of inpainting problems. In Chapter 7 we present the disocclusion method used to obtain the completed objects of the scene, together with a probabilistic method, which decides which is the most probable scene interpretation. The detailed algorithm is presented in Chapter 8. In Chapter 9 we show results on synthetic and real images and, finally, the conclusions are presented in Chapter 10. The last Part of this Thesis, Part III, is devoted to present the methods for video inpainting. Chapter 11 introduces the binary video inpainting and optical flow problems. Chapter 12 is fully dedicated to the binary video inpainting problem. We present our method, which completes the shapes in a smooth way using both, the temporal and the spatial information. As we are considering binary videos it includes an extra difficulty for the minimization part. For this reason, we propose to use the Allen-Cahn equation to be able to find a solution of our problem. We explain the proposed minimization strategy, which is based on Threshold Dynamics. And, finally we provide some details on the code and examples on different applications. In Chapter 13 we present our model of optical flow inpainting. We propose to use the information from the frames in order to complete the optical flow on the missing areas using a geodesic distance. Finally, in Chapter 14 we provide the conclusions of this part together with some proposals of future work. CONTRIBUTIONS Publications: [1] Oliver M., Raad L., Ballester C., Haro G., Motion Inpainting by an Image-Based Geodesic AMLE Method. IEEE International Conference on Image Processing, 2018. (Accepted) [2] Oliver M., Haro G., Fedorov V., Ballester C., L1 Patch-Based Image Partitioning Into Homogeneous Textured Regions. IEEE International Conference on Acoustics, Speech and Signal Processing, pp. 1558--1562, 2018. [3] Oliver M., Palomares R. P., Ballester C., Haro G., Spatio-Temporal Binary Video Inpainting Via Threshold Dynamics. IEEE International Conference on Acoustics, Speech and Signal Processing, pp. 1822--1826, 2017. [4] Oliver M., Haro G., Dimiccoli M., Mazin B., Ballester C., A Computational Model for Amodal Completion. Journal of Mathematical Imaging and Vision. 56(3):511--534, 2016. Conference Presentations [1] Oliver M., Haro G., Fedorov V., Ballester C., L1 Patch-Based Image Partitioning Into Homogeneous Textured Regions. SIAM Conference on Imaging Science, June 2018, Bologna (Italy) (Poster presentation) BEST POSTER AWARD (2nd position) [2] Oliver M., Palomares R. P., Ballester C., Haro G., Spatio-temporal binary video inpainting via threshold dynamics. Annual Catalan Meeting on Computer Vision, September 2017, Barcelona (Spain. (Poster presentation) [3] Oliver M., Haro G., Dimiccoli M., Baptiste M., Ballester C., A Computational Model of Amodal Completion. Annual Catalan Meeting on Computer Vision, September 2016. Barcelona (Spain). (Poster presentation) [4] Oliver M., Haro G., Dimiccoli M., Baptiste M., Ballester C., A Computational Model of Amodal Completion. SIAM Conference on Imaging Science, minisymposium: Geometry-based Models in Image Processing. May 2016, Albuquerque (New Mexico) (Oral presentation)