Comparative analysis of attention-based, convolutional, and SSM-based models for multi-domain image classification

dc.contributor.advisorChakrabarty, Dr. Amitabha
dc.contributor.authorBiswas, Mondrita
dc.contributor.authorRahman, Sayeedur
dc.contributor.authorTarannum, Syeda Farhat
dc.contributor.authorNishanto, Dipro
dc.contributor.authorSafwaan, Md Aqeed
dc.date.accessioned2025-06-16T10:12:09Z
dc.date.available2025-06-16T10:12:09Z
dc.date.issued2025-01
dc.descriptionCataloged from PDF version of thesis.
dc.descriptionIncludes bibliographical references (pages 110-116).
dc.descriptionThis thesis is submitted in partial fulfillment of the requirements for the degree of Bachelor of Science in Computer Science, 2025.
dc.description.abstractThe increasing frequency and severity of environmental and societal challenges, such as natural disasters, medical diagnostics, and agricultural threats require the development of efficient and scalable detection and classification systems. Lightweight and fast models deployed on edge devices, such as surveillance drones, portable diagnostic tools, or agricultural sensors, can address constraints of network delays, adverse conditions, and bandwidth limitations often faced by autonomous technologies. Transformer-based models using attention mechanisms, trade off computational costs to achieve high accuracies in these classification tasks. Recently, State-space models (SSMs) have emerged as a promising alternative in areas where long-range dependence on data is crucial but computational efficiency is particularly important. This research explores the application of Attention-Based, Convolutional, and SSMs, particularly Vision Mamba (ViM), in diverse domains: wildfire detection, plant disease identification, and skin cancer diagnostics. Finally, the feasibility of knowledge distillation in ViM is examined using the information gathered from a thorough evaluation and model comparisons. Evaluations highlight that while CNN models consistently achieved the highest accuracy, ViM Tiny is the most memory-efficient, requiring only 0.03GB of GPU memory. ViM Tiny (7.60M params) achieved 70.60% accuracy in wildfire detection, matching DeiT Base’s (85.80M Params) 70.62% accuracy. The SSM-based models also had the fastest convergence rate. These models achieved promising accuracies in plant disease classification (98.71%–99.65%) and skin cancer detection (87.03%–90.16%), highlighting their potential for efficient and scalable vision tasks. In the context of wildfire detection, knowledge distillation with EfficientNet B7 as a teacher model further improved ViM Tiny’s accuracy from 70.6% to 85.32%, highlighting its potential for lightweight, high-performance applications in critical scenarios.
dc.identifier.otherID: 21101056
dc.identifier.otherID: 21101281
dc.identifier.otherID: 21101016
dc.identifier.otherID: 21101032
dc.identifier.otherID: 21101066
dc.identifier.otherhttps://dspace.bracu.ac.bd/server/api/core/items/83953f81-3e68-4473-8fca-f9e163e2c1c6
dc.identifier.urihttp://hdl.handle.net/10361/26059
dc.language.isoen
dc.publisherBRAC University
dc.sourceBRAC University Institutional Repository
dc.subjectState space models
dc.subjectVision Mamba
dc.subjectKnowledge distillation
dc.subjectMulti-domain applications
dc.subjectAttention-based models
dc.subjectConvolutional models
dc.titleComparative analysis of attention-based, convolutional, and SSM-based models for multi-domain image classification
dc.typeThesis

Files

Original bundle

Now showing 1 - 1 of 1
Thumbnail Image
Name:
21101056, 21101281, 21101016, 21101032, 21101066_CSE.pdf
Size:
9.85 MB
Format:
Adobe Portable Document Format