Exploring the Limits of Large Language Models’ Ability to Distinguish Between Objects

Hyeongjin Ju, Incheol Park, Yağız Nalçakan, Youngwan Jin, Sanghyeop Yeo, Shiho Kim · Applied Sciences · 2025

This paper explores the capability of large language models (LLMs) to accurately classify objects in challenging visual scenarios, focusing on two main tasks: differentiating real objects from artificial replicas and distinguishing human figures from human-like entities (e.g., mannequins, banners). We evaluate a diverse set of vision–language models (VLMs) ranging from large-scale architectures to parameter-efficient systems across multiple question prompts designed to probe object identification, authenticity verification, and multi-object reasoning. Our experiments reveal that while many models perform reasonably well in identifying single objects, their accuracy declines substantially under more complex conditions, such as multi-object scenes or tasks requiring fine-grained judgments of authenticity. Even top-tier models exhibit noticeable performance drops from around 100% to below 15% accuracy when forced to discern real from fake items among multiple candidates, and from 100% to 83.33% accuracy on providing positional details for human-like figures. We further discuss how these performance limitations indicate gaps in current LLM-based vision systems to highlight the need for more robust spatial reasoning and attribute analysis. Our findings underscore the significance of broadening these models’ multimodal understanding and refining prompts, with an eye toward improving real-world applications—from automated quality control to surveillance—where nuanced visual classification is crucial. By comparing a variety of architectures under consistent evaluation settings, this study offers insights into the barriers LLMs face when confronted with increasingly complex visual information.

Read the paper · More papers on PaperTik