Text classification with an efficient preprocessing technique for cross-language and multilingual data
| dc.contributor.advisor | Ashraf, Faisal Bin | |
| dc.contributor.author | Khan, Towhid | |
| dc.contributor.author | Mallick, David Dew | |
| dc.contributor.author | Khan, Md.Shakiful Islam | |
| dc.contributor.author | Hasan, Md Mahadi | |
| dc.date.accessioned | 2023-10-17T08:43:07Z | |
| dc.date.available | 2023-10-17T08:43:07Z | |
| dc.date.issued | 9/28/2022 | |
| dc.description | Cataloged from PDF version of thesis. | |
| dc.description | Includes bibliographical references (pages 43-44). | |
| dc.description | This thesis is submitted in partial fulfillment of the requirements for the degree of Bachelor of Science in Computer Science, 2022. | |
| dc.description.abstract | The procedure of eradicating extraneous textual elements and preparing or process- ing the values to be fed into the classifier model is often indicates the concept of text-preprocessing. There are several preprocessing methods, however not all of them are effective when used with cross-language and multilingual datasets. Run- ning a cross-lingual or multilingual dataset through a single pre-processing method and text classification model is rather challenging. What if a technique could be used to better classify data from multilingual and cross lingual datasets? In order to accelerate the process of improving accuracy, we tested various combinations of data pre-processing with text classification models on datasets in Bangla, English, and cross-lingual (Native language written in English letters). We may infer from our experiment that mLSTM functioned effectively for datasets in Bangla and English. Thus, mLSTM can be a helpful preprocessing method for datasets containing a variety of languages. | |
| dc.identifier.other | ID 18201035 | |
| dc.identifier.other | ID 18201045 | |
| dc.identifier.other | ID 18201198 | |
| dc.identifier.other | ID 18201062 | |
| dc.identifier.other | https://dspace.bracu.ac.bd/server/api/core/items/57dfcf8b-ea6a-4b17-980b-2578ba04c5f9 | |
| dc.identifier.uri | http://hdl.handle.net/10361/21865 | |
| dc.language.iso | en | |
| dc.publisher | BRAC University | |
| dc.source | BRAC University Institutional Repository | |
| dc.subject | Random forest | |
| dc.subject | Logistic regression | |
| dc.subject | TF-IDF | |
| dc.subject | SVM | |
| dc.subject | XGB | |
| dc.subject | mLSTM | |
| dc.subject | LSTM | |
| dc.subject | Information retrieval | |
| dc.subject | Sentiment analysis | |
| dc.subject | NLP | |
| dc.title | Text classification with an efficient preprocessing technique for cross-language and multilingual data | |
| dc.type | Thesis |
Files
Original bundle
1 - 1 of 1
- Name:
- 18201035, 18201045, 18201198, 18201062_CSE.pdf
- Size:
- 8.12 MB
- Format:
- Adobe Portable Document Format
