Shoe-CLIP: Attribute-Based Semantic Similarity Evaluation With Explainable Evidence for Sneaker Design
Hui-Jun Kim, Sang-Heon Oh, Sung-Hee Kim · IEEE Access · 2026
Design similarity evaluation largely depends on designers’ intuition, while existing image-based metrics insufficiently capture attribute-level semantic differences and lack interpretability. This study proposes Shoe-CLIP, a domain-specialized semantic embedding model for footwear design similarity evaluation. The model computes quantitative similarity scores using a CLIP encoder fine-tuned with footwear attribute captions and generates attribute-level explanations through a CLIP–BART-based architecture. A GPT-based summarization module further structures commonalities and differences between designs. Experimental results show that Shoe-CLIP achieves the lowest error relative to attribute-based Ground Truth scores compared to conventional image processing metrics and general-purpose CLIP. The explanation model demonstrates high semantic alignment, and in a blind preliminary comparison with designers, Shoe-CLIP achieved the lowest absolute error among the evaluated metrics, while the limited sample size requires cautious interpretation of inferential claims. Generated explanations closely reflect designers’ core evaluation criteria. These findings indicate that domain-specialized semantic embeddings can simultaneously ensure quantitative accuracy and explainability in design similarity evaluation, with potential applications in design validation, intellectual property risk assessment, and attribute-based retrieval systems.