WHISNER-BN: parameter-efficient end-to-end spoken named entity recognition for low-resource languages with morphology-aware alignment

dc.contributor.advisorAlam, Md. Golam Rabiul
dc.contributor.authorImran, Shah
dc.date.accessioned2026-04-01T07:14:00Z
dc.date.available2026-04-01T07:14:00Z
dc.date.issued2025-12
dc.descriptionCataloged from PDF version of thesis.
dc.descriptionIncludes bibliographical references (pages 73-75).
dc.descriptionThis thesis is submitted in partial fulfillment of the requirements for the degree of Bachelor of Science in Computer Science and Engineering, 2025.
dc.description.abstractNamed entity recognition from speech remains underdeveloped for low-resource languages such as Bengali, despite its importance for voice search, conversational AI, and accessibility. This thesis investigates Bengali spoken NER through three paradigmsdiscriminative structured prediction, generative multi-task learning, and multimodal instruction-tuned approaches-with the primary contribution whisNer-bn, a parameterefficient architecture employing Low-Rank Adaptation of Whisper encoders (7.1% trainable parameters), BiLSTM contextual encoding, and Conditional Random Field decoding with explicit BIO constraints. To enable end-to-end training, we introduce bnSpAligner, a morphology-aware forced alignment algorithm achieving 78% tokenlevel accuracy for Common Voice and 83% for SUBAKKO through adaptive thresholding and phonetic equivalence classes, and BSSC-Annotator, a multi-agent framework leveraging cross-lingual transfer to produce 50.57 hours of annotated Bengali speech comprising 38,504 entities across 36,237 utterances at less than 10% of manual annotation cost. Evaluated on 5,000 samples from combined test partitions using models trained on 20% of available data (4,989 samples), whisNer-bn achieves 0.681 F1, outperforming generative multi-task approaches by 15.6 percentage points and cascaded ASR-NER pipelines by 5.8 percentage points. The results demonstrate that discriminative structured prediction with joint acoustic-linguistic modeling provides superior inductive biases for entity recognition in morphologically complex languages, establishing the first systematic benchmark and transferable methodology for Bengali spoken NER with implications for thousands of underserved languages.
dc.identifier.otherID 23366027
dc.identifier.otherhttps://dspace.bracu.ac.bd/server/api/core/items/d81ec9b6-f9b6-4233-81bb-9a675193d511
dc.identifier.urihttp://hdl.handle.net/10361/27696
dc.language.isoen
dc.publisherBRAC University
dc.sourceBRAC University Institutional Repository
dc.subjectMulti-task learning
dc.subjectNamed entity recognition
dc.subjectNER
dc.subjectBengali speech processing
dc.titleWHISNER-BN: parameter-efficient end-to-end spoken named entity recognition for low-resource languages with morphology-aware alignment
dc.typeThesis

Files

Original bundle

Now showing 1 - 1 of 1
Thumbnail Image
Name:
23366027_CSE.pdf
Size:
15.7 MB
Format:
Adobe Portable Document Format