Contemplation or Action: Rethinking the Evaluation of 3D Scene Through Dynamic Game Interaction
Dawei Liu · Uppsala University Publications (Uppsala University) · 2026
As AI-generated 3D scenes increasingly populate interactive applications from video games to virtual training, current scene-level evaluation predominantly relies on objective and quantifiable criteria—such as spatial collision, physical support, and semantic consistency—which implicitly assumes that algorithmic standards align with human perception under all viewing conditions. Yet the validity of these static benchmarks under dynamic human engagement remains an open question. The research objective of this study is to validate whether established plausibility evaluation criteria remain applicable when humans actively engage with dynamic three-dimensional environments, thereby offering new insights and perspectives for the development of AI algorithms that generate 3D scenes. To examine this, the research proposes a novel human evaluation methodology grounded in natural game behavior, employing the Prop Hunt mod as an ecologically valid, time-pressured environment. This study establishes the POSP framework (Position, Orientation, Support, Possibility), operationalizes classical scene perception theories for interactive contexts, and analyzes 400 hiding cases from gameplay videos across four evaluation dimensions. This study addresses two research questions: (RQ1) To what extent do static evaluation criteria, align with human perceptual judgments of scene in dynamic interactive environments? (RQ2) To what extent does simple additive aggregation of dimension scores align with human perceptual judgments of scene in dynamic interactive environments? Results show that no single POSP dimension significantly predicted hiding outcomes (all p > 0.05), indicating that static criteria lose predictive validity under dynamic interaction. Combinatorial analysis further reveals that intermediate compliance levels yielded comparable success rates, demonstrating that simple additive models fail to explain dynamic outcomes. Consequently, scene plausibility during active engagement depends on configurational harmony rather than isolated categorical verification, calling for a fundamental reconsideration of static evaluation standards in interactive contexts.