Shopsense: Retail Analytics Using a Novel Joint Aware Temporal Encoder (JATE)
Kartick S. Rajen, John Charles Thomas, Alice K. · 2025
This paper presents ShopSense, an AI-driven video analytics system for customer behavior analysis. A key contribution of this work is the novel Joint-Aware Temporal Encoder (JATE), a model which, to our knowledge, is state-of-the-art in terms of model size for efficient, pose-based action recognition. ShopSense integrates this model into an end-to-end pipeline with YOLOv8 for detection, a Centroid Tracker for tracking, and Mediapipe for pose estimation. JATE's architecture is specifically engineered to classify key retail actions (walking, standing, reaching) from pose sequences with exceptional computational efficiency, making it a powerful alternative to resource-intensive spatiotemporal models. The system aggregates these outputs to produce tangible insights, including heatmaps and customer pathing, all while being designed for deployment on resourceconstrained edge devices.