Comparative study of toxic comments classification using machine learning algorithms

dc.contributor.advisorChakrabarty, Amitabha
dc.contributor.authorRazzak, Razia
dc.contributor.authorSadril, Md.
dc.contributor.authorShakil, Mahmudul Hasan
dc.contributor.authorRahman, Mahfuzur
dc.contributor.authorTaki, Sabiha Tul Omman
dc.date.accessioned2021-07-15T06:18:46Z
dc.date.available2021-07-15T06:18:46Z
dc.date.issued2021-01
dc.descriptionCataloged from PDF version of thesis.
dc.descriptionIncludes bibliographical references (pages 54-56).
dc.descriptionThis thesis is submitted in partial fulfillment of the requirements for the degree of Bachelor of Science in Computer Science and Engineering, 2021.
dc.description.abstractThe rapid growth of information technology and the disruptive transformation of social media have happened in recent years. Websites like Facebook, Twitter, Instagram, where people can express their thoughts or feelings by posting text, photos or videos, have become incredibly popular. But unfortunately, it has also become a place for hateful activity, abusive words, cyberbullying and anonymous threats. There are many existing works in this field but those are not fully successful yet to provide accuracy in satisfactory level. In this work, we employ natural language processing (NLP) with convolution neural networking (CNN), extreme gradient boosting (XGBoost) and support vector machine (SVM) for segmenting toxic comments at first and then classifying them in six types from a large pool of documents provided by Kaggle’s regarding Wikipedia’s talk page edits. Using this dataset, the hamming score of CNN model is 89% ,XGBoost model is 87% and SVM model is 84%.
dc.identifier.otherID: 16101291
dc.identifier.otherID: 16301032
dc.identifier.otherID: 16301026
dc.identifier.otherID: 16101206
dc.identifier.otherID: 17101519
dc.identifier.otherhttps://dspace.bracu.ac.bd/server/api/core/items/7fa8d312-698f-4add-9675-9fc7c44c3df1
dc.identifier.urihttp://hdl.handle.net/10361/14810
dc.language.isoen_US
dc.publisherBRAC University
dc.sourceBRAC University Institutional Repository
dc.subjectCyberbullying
dc.subjectNatural Language Processing
dc.subjectWord Embedding
dc.subjectConvolutional Neural Networks
dc.subjectXGBoost
dc.subjectSupport Vector Machine
dc.titleComparative study of toxic comments classification using machine learning algorithms
dc.typeThesis

Files

Original bundle

Now showing 1 - 1 of 1
Thumbnail Image
Name:
16101291, 16301032, 16301026, 16101206, 17101519_CSE.pdf
Size:
1.95 MB
Format:
Adobe Portable Document Format