A distributed information model for robust and flexible perceptual interfaces
Petros Faloutsos, Gabriele Nataneli · 2011
Modern artists turn to technology to express their ideas effectively and find in it a medium that goes beyond the limits of traditional pencil and paper. They seek tools that are natural to use and intuitive, but also need solutions that are accurate and controllable. Their demands on technology are often broad and conflicting. As researchers in computer graphics, we are expected to supply solutions supporting creative ideas from the early and crude stages of storyboarding to the demanding and precise requirements of final production. While the industry has already produced several mature applications that are routinely employed in complex projects, these tools are tedious to use and mostly inadequate to yield the true potential of individual artists. In fact, the current limits of technology are only compensated by large budgets and extended work forces. In recent years, the research community has shifted its focus to promising technologies based on machine perception, which are commonly known as perceptual interfaces. In this thesis, I propose a new paradigm for perceptual interfaces as well as practical algorithms and theoretical concepts that are designed to tackle these key challenges. I begin by presenting several practical algorithms and perceptual technologies that aid the creative process at various stages of development in a fixed workflow. Subsequently, I extend these ideas by presenting a framework that enables a perceptual workflow to be configured dynamically around the needs of artists. To this end, I introduce a new distributed representation of complex information that organizes distinct levels of abstraction in a unified model and a reasoning system that can query, modify, and adapt this representation dynamically during an interactive session. This representation coupled with the reasoning system can configure new interaction modes on-the-fly, and lets users control the artistic process at the most appropriate level of abstraction, without sacrificing the degree of control that is often lost in traditional perceptual interfaces. As a distributed representation, this formalism is robust and fault tolerant, but most importantly can be used to pinpoint when perceptual algorithms fail to produce consistent results. When errors do occur, the reasoning system will not proceed by blind guesses, but will make an informed choice and pick the simplest interaction mode that can enable the artist to correct the error predictably and with precision. This approach is seeded by a small knowledge base that describes the assets, algorithms, and external tools that are available to the system. The knowledge base informs the reasoning system which in turn tracks the flow of information during an interactive session and specializes the knowledge base as needed. The system thus learns about the user's intent and the problem domain dynamically, while it is being used. Lastly, I present VisualDive, a full-fledged application that implements these ideas demonstrating how this new paradigm can be used in practice and in harmony with the rich ecosystem of other more traditional and well-respected software tools in computer graphics. Specifically, I show how VisualDive can enhance a workflow based on free-hand sketching to pose facial expressions. I conclude by outlining how these concepts lay the algorithmic foundations for tackling more traditional problems in computer vision that depend on robustness and require a higher degree of automation than in the context of computer graphics. VisualDive is a large cross-platform application built on a flexible plug-in architecture and modern principles of software design. The main view of VisualDive revolves around a node-based interface similar to what found in other major graphics packages, but with profoundly different capabilities and purpose. Unlike traditional applications, the workflow in VisualDive begins with the intent of the user, and progresses through an interactive dialog that is tailored around the user's needs. The reasoning system in VisualDive orchestrates and configures both algorithms that are directly exposed to VisualDive through the plug-in architecture as well as external applications that can cooperate with VisualDive without modification. As a result, the user of VisualDive can easily access a large body of content creation tools blurring the lines between perceptual interfaces and traditional tools. Nonetheless, VisualDive successfully hides such considerable complexity by presenting stark innovation within the comfort and familiar conventions of graphical user interfaces. (Abstract shortened by UMI.)