Text classification with an efficient preprocessing technique for cross-language and multilingual data

dc.contributor.advisorAshraf, Faisal Bin
dc.contributor.authorKhan, Towhid
dc.contributor.authorMallick, David Dew
dc.contributor.authorKhan, Md.Shakiful Islam
dc.contributor.authorHasan, Md Mahadi
dc.date.accessioned2023-10-17T08:43:07Z
dc.date.available2023-10-17T08:43:07Z
dc.date.issued9/28/2022
dc.descriptionCataloged from PDF version of thesis.
dc.descriptionIncludes bibliographical references (pages 43-44).
dc.descriptionThis thesis is submitted in partial fulfillment of the requirements for the degree of Bachelor of Science in Computer Science, 2022.
dc.description.abstractThe procedure of eradicating extraneous textual elements and preparing or process- ing the values to be fed into the classifier model is often indicates the concept of text-preprocessing. There are several preprocessing methods, however not all of them are effective when used with cross-language and multilingual datasets. Run- ning a cross-lingual or multilingual dataset through a single pre-processing method and text classification model is rather challenging. What if a technique could be used to better classify data from multilingual and cross lingual datasets? In order to accelerate the process of improving accuracy, we tested various combinations of data pre-processing with text classification models on datasets in Bangla, English, and cross-lingual (Native language written in English letters). We may infer from our experiment that mLSTM functioned effectively for datasets in Bangla and English. Thus, mLSTM can be a helpful preprocessing method for datasets containing a variety of languages.
dc.identifier.otherID 18201035
dc.identifier.otherID 18201045
dc.identifier.otherID 18201198
dc.identifier.otherID 18201062
dc.identifier.otherhttps://dspace.bracu.ac.bd/server/api/core/items/57dfcf8b-ea6a-4b17-980b-2578ba04c5f9
dc.identifier.urihttp://hdl.handle.net/10361/21865
dc.language.isoen
dc.publisherBRAC University
dc.sourceBRAC University Institutional Repository
dc.subjectRandom forest
dc.subjectLogistic regression
dc.subjectTF-IDF
dc.subjectSVM
dc.subjectXGB
dc.subjectmLSTM
dc.subjectLSTM
dc.subjectInformation retrieval
dc.subjectSentiment analysis
dc.subjectNLP
dc.titleText classification with an efficient preprocessing technique for cross-language and multilingual data
dc.typeThesis

Files

Original bundle

Now showing 1 - 1 of 1
Thumbnail Image
Name:
18201035, 18201045, 18201198, 18201062_CSE.pdf
Size:
8.12 MB
Format:
Adobe Portable Document Format