Crowding: Recent advances and perspectives
Michael H. Herzog, Bilge Sayim · Journal of Vision · 2022
In crowding, the perception of a target strongly deteriorates in the presence of neighboring elements.Crowding is the standard situation in everyday life since elements are rarely seen in isolation (except in psychophysics laboratories).Crowding is not only crucial for normal object recognition, but also, for example, in reading, visual search, and perhaps numerosity estimation.Hence, any theory of vision needs to account for crowding.What causes crowding has been debated heavily for more than a century.Not surprisingly, crowding is mainly discussed within the dominant framework of vision: In this classic framework, dating back to the work of Nobel Laureates Hubel and Wiesel, crowding is explained by lateral interactions, such as lateral inhibition, between neurons coding for similar features, for example, orientation and spatial frequency.Another explanation is the pooling of neural responses from lower to higher level neurons, with the latter having larger receptive fields and thus lower resolution, which causes crowding.However, such simplistic models cannot account for a large variety of data, which have shown that crowding depends on the complex spatial and temporal layout of elements across large parts of the visual field.For example, crowding is strong when a target, such as a Gabor, vernier, or letter, is flanked by other Gabors, verniers, or letters, respectively.A release of crowding occurs when more flankers are added that make up a group from which the target ungroups.Simple local approaches fail because the single flankers are contained in the multiflanker configurations (for a review, see Herzog, Sayim, Chicherov, & Manassi, 2015).These results are more or less well-accepted.However, the mechanisms of crowding are as controversially debated as before, ranging from very basic to highly complex mechanisms (Levi, 2008;Whitney & Levi, 2011;Herzog et al., 2015), as this special issue shows.Because local pooling models have clear problems to explain why complex configurations determine crowding strength, complex pooling models, such as the texture tiling model (TTM), were proposed that pool information in different feature channels across large regions of the visual field (Balas, Nakano, & Rosenholtz, 2009; Rosenholtz, Yu, & Keshvari, 2019).Whereas the TTM is a one-stage model, other models propose that crowding occurs only within perceptual groups, requiring at least two processing stages to account for crowding: first perceptual groups need to be computed and then interference occurs through a different mechanism (Herzog et al., 2015).Attentional accounts of crowding propose that imprecise attention to the target location (Strasburger, 2005) or insufficient attentional resolution, in contrast with sufficient visual resolution, underlie crowding (He, Cavanagh, & Intriligator, 1996).In this special issue, some of the major questions are where and how crowding occurs in the processing hierarchy (low level vs. high level), whether it occurs in different channels (magnocellular vs. parvocellular; fovea vs. periphery; different color systems), what is the role of other mechanisms, such as attention and redundancy masking, whether crowding needs only one or at least two processing stages (e.g., a grouping stage), and when and under what conditions crowding rules hold and are specific to crowding.Even though complex interactions seem to prevail in crowding, there are approaches that attempt to explain crowding by low-level mechanisms.For example, Rodriguez and Granger (2021) propose that grouping effects can be explained by a generalized contrast-still a rather simple-mechanism, which does not involve an explicit grouping stage.Instead, basic center-surround mechanisms in the early visual stream underlie their proposed mechanism.Previously, it was argued that local pooling cannot explain crowding but complex pooling models, such as the TTM (Balas et al., 2009), can.As mentioned, the TTM does not require any explicit grouping stage-one stage is sufficient (Rosenholtz et al., 2019).However, Bornet et al. (2021) tested the TTM in six key experimental studies that highlighted high-level effects in crowding.The TTM performed similar to simple pooling mechanisms and was thus not able to account for the results.