Redefining Bangla Text Processing with a New Stemming Methodology

dc.contributor.authorTanvir, Md Istiak
dc.date.accessioned2026-06-25T04:34:23Z
dc.date.available2026-06-25T04:34:23Z
dc.date.issued2025-01-13
dc.descriptionProject Report
dc.description.abstractStemming is a basic NLP technique that normalizes the linguistic structure and improves the performance of text analysis by stripping words to their roots or base forms by removing suffixes which are quite vital in applications such as text normalization, information retrieval and keeping linguistic consistency. The Bangla language has a special stemming problem due to its extensive morphological structure with many inflectional changes and intricate grammatical rules. Hence, finding the root form of a word in Bangla involves much more complex affixation patterns and compound word creation than is evident in other languages with simpler grammatical systems. To handle the linguistic complexities of Bangla better than the previous approaches. In this work a new stemming method is developed for the language. It is more adaptable and practical for root word extraction in the real world because it has been equipped with the most updated methods to recognize and handle both the Bangla suffixes and morphological changes. Very encouraging results are obtained from the complex experiment carried out on a dataset containing 1,000 unique Bangla words. It’s extremely high F1-score of 85.66% and high accuracy rate of 87.2% do, in fact, justify its superior performance over traditional stemming algorithms. It is apparent from the present study that this newly proposed stemming strategy significantly enhances the efficacy and efficiency of all Bangla text processing systems, including search engines, information retrieval platforms, and all NLP applications. This study paves the way for further development in the Bangla Language Processing problem and emphasizes the cruciality of continued research in developing language-specific Natural Language Processing tools.
dc.identifier.otherhttp://dspace.daffodilvarsity.edu.bd:8080/handle/123456789/17455
dc.identifier.urihttp://dspace.daffodilvarsity.edu.bd:8080/handle/123456789/17455
dc.language.isoen_US
dc.publisherDaffodil International University
dc.sourceDIU Institutional Repository
dc.subjectBangla Stemming
dc.subjectNatural Language Processing (NLP)
dc.subjectMorphological Analysis
dc.subjectRoot Word Extraction
dc.subjectBangla Language Processing
dc.subjectInformation Retrieval
dc.titleRedefining Bangla Text Processing with a New Stemming Methodology
dc.typeOther

Files

Original bundle

Now showing 1 - 1 of 1
No Thumbnail Available
Name:
221-15-4720.pdf.txt
Size:
63.03 KB
Format:
Adobe Portable Document Format

Collections