CateSift: An interactive steering approach for classifying large scale text

Chundong Wang, Yuhan Tian, Xumeng Wang, Yixuan Song, Haotian Zhang, Yongxin Zhao · Visual Informatics · 2025

Concept management for large-scale text data is critical in domains such as healthcare informatics, digital libraries, and news classification. However, the variability in concept structures and the diversity of application requirements pose challenges for existing automated methods, which often lack the flexibility to accommodate customized needs. Meanwhile, manual classification remains resource-intensive and inefficient. To address this issue, we propose CateSift, an interactive approach that integrates public knowledge to streamline the classification process and incorporates expert knowledge to formulate classification models. The main contributions of this work are as follows: (1) a visualization interface, called CateSift , that facilitates users in constructing and refining classification models for large-scale data, and (2) A prompt-based model that can integrate expert knowledge to iteratively refine hierarchical classification structures. Specifically, CateSift provides users with a hierarchical concept tree that highlights concepts with uncertain classifications and invites users to optimize the classification models by injecting knowledge. To address the issue of large-scale data, CateSift allows users to steer the classification model by adjusting the classification tree or annotating classifications. Case studies indicate that the proposed approach effectively and efficiently supports classification for large-scale data. • This study presents an interactive classification framework, CateSift , which employs multi-level prompt templates and integrates prompt-based large language models with expert knowledge to perform hierarchical classification on large-scale datasets. This method effectively overcomes the limitations of existing automated approaches regarding accuracy and hierarchical structure flexibility, as well as the inefficiencies associated with manual classification. • An interactive prototype system is introduced to support users in interpreting and guiding model classification outcomes, enabling efficient detection of potential errors and unstable hierarchical components. The system facilitates iterative model refinement through user-provided corrections and annotations, accommodating diverse domain-specific needs and further decreasing user effort.

Read the paper · More papers on PaperTik