| International Journal of Computer Applications |
| Foundation of Computer Science (FCS), NY, USA |
| Volume 187 - Number 135 |
| Year of Publication: 2026 |
| Authors: Tu Thanh Tri, Duong Thi Thuy Nga, Chau Phuong Toan, Dang Kim Lien |
10.5120/ijcac919f6ae2058
|
Tu Thanh Tri, Duong Thi Thuy Nga, Chau Phuong Toan, Dang Kim Lien . An Explainable Vision Transformer Framework for Skin Lesion Classification using Dual-Map Fusion Strategy. International Journal of Computer Applications. 187, 135 ( Aug 2026), 35-42. DOI=10.5120/ijcac919f6ae2058
Accurate and interpretable skin lesion classification is essential for early melanoma diagnosis and clinical decision support. This study proposes an explainable computer-aided diagnosis framework based on a Vision Transformer (ViT) for binary classification of benign and malignant skin lesions. To improve model transparency, Grad-CAM and transformer attention maps are integrated through a Dual-Map Fusion strategy, providing complementary local and global visual explanations of the model's predictions. The proposed framework was evaluated using a combined HAM10000 and ISIC dermoscopic image dataset. Experimental results demonstrated an overall classification accuracy of 92.55% and a malignant lesion recall of 94.68%, indicating reliable diagnostic performance with a reduced risk of missed malignant cases. In addition, the fused explanation maps provided more informative and interpretable visual evidence than Grad-CAM or attention maps alone, facilitating a better understanding of the model's decision-making process. The proposed framework combines the strong classification capability of Vision Transformer with enhanced explainability, providing an effective and transparent decision-support tool for automated skin lesion diagnosis. These findings demonstrate the potential of explainable transformer-based models for reliable clinical application in dermatological image analysis.