Comparison of Unigram, Bigram, HMM and Brill's POS tagging approaches for some South Asian languages

dc.contributor.authorHasan, Muhammad Fahim
dc.contributor.authorNaushad UzZaman
dc.contributor.authorKhan, Mumit
dc.date.accessioned2010-10-05T05:03:09Z
dc.date.available2010-10-05T05:03:09Z
dc.date.issued2007
dc.descriptionIncludes bibliographical references (page 6-8).
dc.description.abstractPart-of-Speech (POS) Tagging is a process that attaches each word in a sentence with a suitable tag from a given set of tags. POS Tagging is important in various areas of Natural Language Processing. Different methods of automating the process have been developed and employed for English and other Western languages. Some similar work, most of which utilize the stochastic approaches for POS Tagging has also been done in the same area for South Asian languages. We experimented with some of the widely-used approaches for POS Tagging on three South Asian languages, Bangla, Hindi and Telegu, using corpora of different sizes. We observed the performance of the approaches and found the Brill’s transformation based tagger’s performance to be superior to the other approaches in all of our experiments, though the use of this approach has been very limited until recently.
dc.identifier.otherhttps://dspace.bracu.ac.bd/server/api/core/items/cadfcdb7-51d3-4745-bd20-6a4f09534a96
dc.identifier.urihttp://hdl.handle.net/10361/330
dc.language.isoen
dc.publisherBRAC University
dc.sourceBRAC University Institutional Repository
dc.subjectPart-of-speech tagging
dc.subjectLanguage processing
dc.titleComparison of Unigram, Bigram, HMM and Brill's POS tagging approaches for some South Asian languages
dc.typeArticle

Files

Original bundle

Now showing 1 - 1 of 1
Thumbnail Image
Name:
Comparison of Unigram, Bigram, HMM and Brill’s POS.pdf
Size:
140.26 KB
Format:
Adobe Portable Document Format