Exploring deep learning models for handwritten Bengali compound character classification: a comparative study

dc.contributor.advisorRasel, Annajiat Alim
dc.contributor.authorHossain, Md Iftekhar
dc.date.accessioned2026-01-14T09:26:16Z
dc.date.available2026-01-14T09:26:16Z
dc.date.issued2025-10
dc.descriptionCataloged from PDF version of thesis.
dc.descriptionIncludes bibliographical references (pages 51-54).
dc.descriptionThis thesis is submitted in partial fulfillment of the requirements for the degree of Bachelor of Science in Computer Science, 2025.
dc.description.abstractRecognizing isolated juktoborno in Bangla script, which consists of 334 unique compound characters, remains a challenging problem in Optical Character Recognition (OCR). This research explores the efficacy of deep learning-based approaches by comprehensively evaluating nine state-of-the-art architectures: EfficientNet-B0, MobileNet V3, DenseNet-121, ConvNeXt V2 Tiny, ViT-Small-DINOv3, Custom CNN, ResNet-50, VGG-16, and SqueezeNet. We employed transfer learning using ImageNet pre-trained weights, fine-tuning all models on the MatriVasha Dataset containing 120 compound character classes with 306,461 grayscale images at 128×128 resolution. The dataset was partitioned into 70 All models were trained using AdamW optimizer with cosine annealing learning rate schedule, data augmentation (RandomAffine, RandomPerspective, RandomErasing), and early stopping on high-performance hardware (dual RTX 5090 GPUs). EfficientNet-B0 achieved the highest accuracy of 97.68% with only 4.16M parameters, demonstrating superior efficiency. MobileNet V3 secured second place at 97.01% while being the fastest to train (5.7 minutes), and DenseNet-121 achieved 96.40% with effective feature reuse. This study contributes to advancing Bangla OCR technology through comprehensive architecture comparison and establishes new performance benchmarks for compound character recognition. The findings will help in practical applications like document digitization, automated translation, text-to-speech conversion, and search-based text retrieval, with EfficientNetB0 recommended for maximum accuracy and MobileNet V3 for optimal speed-accuracy balance in production deployment.
dc.identifier.otherID 23241157
dc.identifier.otherhttps://dspace.bracu.ac.bd/server/api/core/items/7ab9485a-bbdc-49ff-816e-d8a1168d689d
dc.identifier.urihttp://hdl.handle.net/10361/27439
dc.language.isoen
dc.publisherBRAC University
dc.sourceBRAC University Institutional Repository
dc.subjectDeep learning model
dc.subjectCompound characters
dc.subjectCharacter recognition
dc.subjectBengali language
dc.subjectHandwritten characters
dc.subjectBengali alphabet
dc.subjectCompound letters
dc.titleExploring deep learning models for handwritten Bengali compound character classification: a comparative study
dc.typeThesis

Files

Original bundle

Now showing 1 - 1 of 1
Thumbnail Image
Name:
23241157_CSE.pdf
Size:
729.74 KB
Format:
Adobe Portable Document Format