Cross-attention fusion vision transformer for explainable and efficient multi-class eye disease detection

dc.contributor.advisorRasel, Annajiat Alim
dc.contributor.advisorAgomoni, Ahmed Mayeesha Reza
dc.contributor.authorRahman, Yasin
dc.contributor.authorMozahedul Hoque, Md.
dc.contributor.authorChowdhury, Moriyum
dc.contributor.authorMustakim Al Mahmud
dc.contributor.authorMahi, Mahidul Islam
dc.date.accessioned2026-04-22T06:56:13Z
dc.date.available2026-04-22T06:56:13Z
dc.date.issued2026-01
dc.descriptionCataloged from PDF version of thesis.
dc.descriptionIncludes bibliographical references (pages 91-95).
dc.descriptionThis thesis is submitted in partial fulfillment of the requirements for the degree of Bachelor of Science in Computer Science and Engineering, 2026.
dc.description.abstractEarly and accurate detection of retinal fundus diseases is critical for preventing irreversible vision loss and supporting effective clinical decision-making. Retinal fundus imaging is widely used for large-scale screening due to its non-invasive nature; however, the diversity and structural complexity of retinal pathologies pose significant challenges for automated analysis. Convolutional Neural Networks (CNNs) have been extensively employed for fundus image classification owing to their strong local feature extraction capabilities, however, their limited ability to model long-range contextual dependencies constraints performance in complex multi-class disease scenarios. Vision Transformers (ViTs), on the other hand, leverage self-attention mechanisms to capture global contextual information but often suffer from high computational costs and reduced effectiveness in limited-data medical imaging settings. The study proposed a lightweight hybrid Cross-Attention Fusion Vision Transformer architecture can be used to classify multi-class retinal diseases through fundus images. The suggested model combines CNN-based local feature extraction and transformerbased global contextual modelling with the cross-attention fusion mechanism, which allows both fine-grained pathological features and holistic retinal structure to interact and at the same time keep computational efficiency. The hybrid model is specifically designed with small-scale medical data in mind and uses attention-based interpretability, which helps to include explainable AI in the future to improve clinical transparency. Experimental evaluation on publicly available fundus datasets demonstrates that the proposed hybrid approach achieves a favorable balance between accuracy, robustness, and efficiency compared to standalone CNN and Vision Transformer models, highlighting its suitability for automated retinal disease screening applications.
dc.identifier.otherID 23341115
dc.identifier.otherID 22101379
dc.identifier.otherID 21301210
dc.identifier.otherID 24241323
dc.identifier.otherID 21301542
dc.identifier.otherhttps://dspace.bracu.ac.bd/server/api/core/items/3687c47c-4b8c-48f8-a099-f7c580730b24
dc.identifier.urihttp://hdl.handle.net/10361/28029
dc.language.isoen
dc.publisherBRAC University
dc.sourceBRAC University Institutional Repository
dc.subjectRetinal fundus imaging
dc.subjectEye disease classification
dc.subjectConvolutional neural networks
dc.subjectVision transformer
dc.subjectLightweight deep learning
dc.titleCross-attention fusion vision transformer for explainable and efficient multi-class eye disease detection
dc.typeThesis

Files

Original bundle

Now showing 1 - 1 of 1
Thumbnail Image
Name:
23341115, 22101379, 21301210, 24241323, 21301542_CSE.pdf
Size:
1.38 MB
Format:
Adobe Portable Document Format