FDTs: A Feature Disentangled Transformer for Interpretable Squamous Cell Carcinoma Grading

Pan Huang, Xin Luo · IEEE/CAA Journal of Automatica Sinica · 2025

Dear Editor, This letter proposes an end-to-end feature disentangled Transformer (FDTs) for entanglement-free and semantic feature representation to enable accurate and trustworthy pathology grading of squamous cell carcinoma (SCC). Existing vision transformers (ViTs) can implement representation learning for SCC grading, however, they all adopt the class-patch token fuzzy mapping for pattern prediction probability or window down-sampling to enhance the representation to contextual information. Such a mechanism results in humanly incomprehensible decision and poorly semantics, and eventually leads to severely entangled feature representation. Motivated by this critical issue, this paper innovatively propose the FDTs relying on two-fold ideas: 1) Building a Semantic Instance-based feature disentangled learning (SIFDL) framework that accurately divides the image into multiscale instances for parallel multi-objective optimization, followed by instance aggregations to enable human-interpretable semantics of the feature representations; 2) Integrating an instance attention block (IAB) to discover the relationship between semantic instances and grading patterns at the instance level to reduce the entanglements of low-effect instances (like lymphatic or muscle) on the grasped feature representations. Experiments on two real-world SCC datasets indicate that the proposed FDTs has superior performance in addressing the task of trustworthy SCC grading with state-of-the-arts.

Read the paper · More papers on PaperTik