DenseViT: A Hybrid CNN-Vision Transformer Model for an Improved Multisensor Lithological Classification
Michael Appiah-Twum, Wenbo Xu, Edward Mensah Acheampong · 2024
In remote sensing, complex spatial layouts and feature interpretation raise concerns in scene classification. Convolutional neural networks (CNNs) do a great job at capturing global features yet, they have a propensity to overlook long-range contextual detail. In contrast, Vision Transformers (ViTs) excel in contextual information extraction with computational complexity and local feature capturing as their downside. This study therefore seeks to strike a balance between CNNs and ViTs with a proposed DenseViT model for an efficient lithological mapping integrating Landsat-9 and ASTER geodata. The proposed model is evaluated against ViT, ResNet50, Random Forest (RF) and KNN, using metrics including area under the curve (AUC), accuracy, sensitivity, specificity and confusion matrices. The result metrics show that DenseViT averages a 1.92% increase in performance relative to the comparative models in this study with an 83.66% accuracy. The proposed DenseViT model demonstrates computational efficiency and encapsulates local and global features remarkably.