Statistical Structure and Task Dependence in Visual Cue Integration 1

Paul Schrater, Daniel Kersten · 1999

A full Bayesian approach to vision requires consideration of potential interactions between all the variables in both the scene and image. A complete model of the interactions, however, would seem computationally intractable because of the large dimensionality of image measurements and scene properties. As a consequence, both experimental studies and theoretical models of human vision have relied on an assumption of modularity in which a particular scene property, such as object depth, is estimated from a restricted set of image measurements, such as image size. The computational problem is not hopeless, however, and can be surmounted by restricting the task and taking advantage of the statistical structure of the problem. In a Bayesian context, modularity falls out of the conditional independencies in the joint distribution of scenes and images p(S, I). By conditioning the joint distribution with respect to particular inference tasks, further modularity is possible while preserving optimal cue combination. We illustrate the problem of modularity and cue combination for the perception of depth from two highly disparate cues, cast shadow position and image size. While strong modularity would suggest ad hoc or no cue combination, we find that the performance of human subjects is better predicted by near-optimal cue combination.

Read the paper · More papers on PaperTik