Vision Transformer for 2D Biomedical Image Classification

 




 

Teh, Yee You (2026) Vision Transformer for 2D Biomedical Image Classification. Final Year Project (Bachelor), Tunku Abdul Rahman University of Management and Technology.

[img] Text
Teh Yee You_Thesis.pdf
Restricted to Registered users only

Download (6MB)

Abstract

This research investigates the efficacy of Vision Transformers (ViTs) and Hybrid Vision Transformers (HViTs) for the classification of 2D brain magnetic resonance imaging (MRI) scans. While traditional Convolutional Neural Networks (CNNs) excel at identifying local spatial patterns, they often lack the global context and long-range dependency modeling required for precise clinical diagnosis. To address this, a modular deep learning pipeline was developed to conduct a comprehensive comparative analysis across seven distinct architectures: AlexNet, ResNet50, ConvNeXt, an Attention-based CNN, a Train-from-scratch ViT, a Pretrained ViT, and a Hybrid-ViT. The study utilized an open-access brain tumor dataset, processed through a robust pipeline involving data augmentation to ensure model generalizability. The proposed Hybrid-ViT architecture integrates a ResNet50 backbone for initial feature extraction with a Transformer encoder to capture global anatomical relationships. Comparative findings demonstrate that while the Attention-CNN offered superior computational efficiency, the Hybrid-ViT achieved the highest peak classification accuracy of 96%. Analysis of alternative models revealed that the Train-from-scratch ViT underperformed significantly (87% accuracy), underscoring the critical role of transfer learning and hybrid structures in medical imaging tasks where data may be limited. Furthermore, the integration of Explainable AI (XAI) techniques, specifically Grad-CAM and LIME, provided transparency into the decision-making process of each model. Visualizations confirmed that the Hybrid-ViT and Pretrained architectures consistently focused on pathologically relevant regions, whereas simpler models often relied on anatomically inconsistent features. This multi-model evaluation provides a clear benchmark for transformer-based diagnostic systems, demonstrating that hybrid designs successfully bridge the gap between local detail and global context to improve both diagnostic accuracy and interpretability in clinical environments

Item Type: Final Year Project
Subjects: Technology > Electrical engineering. Electronics engineering
Faculties: Faculty of Engineering and Technology > Bachelor of Electrical and Electronics Engineering with Honours
Depositing User: Library Staff
Date Deposited: 24 Jul 2026 08:43
Last Modified: 24 Jul 2026 08:43
URI: https://eprints.tarc.edu.my/id/eprint/38003