PEEB: Part-based Image Classifiers with an Explainable and Editable Language Bottleneck

Thang M. Pham, Peijie Chen, Tin Nguyen, Seunghyun Yoon, Trung Bui, Anh‐Tu Nguyen · 2024

CLIP-based classifiers rely on the prompt containing a {class name} that is known to the text encoder.Therefore, they perform poorly with new classes or the classes whose names rarely appear on the Internet (e.g., scientific names of birds).For fine-grained classification, we propose PEEB-an explainable and editable classifier to ( 1) express the class name into a set of text descriptors that describe the visual parts of the class; and (2) match the embeddings of the detected parts with their textual descriptors in each class to compute a logit score for classification.In a zero-shot setting where the class names are unknown, PEEB significantly outperforms CLIP, achieving a 10-fold increase in top-1 accuracy.Compared to part-based classifiers, PEEB not only achieves state-of-the-art (SOTA) accuracy in the supervised-learning setting-88.80% and 92.20% accuracy on and Dogs-120 , respectively-but also the first to enable users to edit the text descriptors to form a new classifier without any re-training.Compared to concept bottleneck models, PEEB is also the SOTA in both zero-shot and supervised learning settings.

Read the paper · More papers on PaperTik