A Modified Naïve Bayesian-based Spam Filter using Support Vector Machine

dc.contributor.authorHossain, Md. Sabir
dc.contributor.authorZubair, Md.
dc.contributor.authorRahman, Mohammad Obaidur
dc.contributor.authorPatwary, Muhammad Kamrul Hossain
dc.contributor.authorRajib, Md. Golam Sarwar
dc.date.accessioned2026-07-06T21:13:20Z
dc.date.available2026-07-06T21:13:20Z
dc.date.issued3-May-2019
dc.description.abstractThe ever-growing problem which is threatening the
dc.description.abstractcurrent mailing system is spam. Spam is nothing but an
dc.description.abstractunsolicited bulk e-mail frequently sent in a financial nature
dc.description.abstractwhich generates the need for creating an anti-spam filter.
dc.description.abstractAmongst many spam filtering techniques, the most advanced
dc.description.abstractmethod "Naïve Bayesian filtering" using the Support Vector
dc.description.abstractMachine (SVM) have been implemented. Spammers are very
dc.description.abstractcareful about the filtering techniques. For that very reason,
dc.description.abstractdynamic filtering is needed and the proposed method meets the
dc.description.abstractdemand. The algorithm splits the received email into tokens and
dc.description.abstractuses Bayes' theorem of probability to calculate the probability of
dc.description.abstractspam for each token to determine the total spam probability of
dc.description.abstractthe mail. Implementation of SVM instead of corpora is one of the
dc.description.abstractadded features of the algorithm. The most challenging feature
dc.description.abstractwas to take the words as well as whole sentences as input in the
dc.description.abstractSVM as tokens and feature vectors. The inclusion of sentences in
dc.description.abstractthe dataset training has increased the accuracy of detecting spam
dc.description.abstractand ham. Natural Language Tool Kit (NLTK) has been used as a
dc.description.abstractuseful language processing tool to tokenize the sentences and
dc.description.abstractalso to understand the meaning of the same types of sentences to
dc.description.abstractsome extent. As a test mail is being compared by word to word
dc.description.abstractand also sentence to sentence from the training datasets to
dc.description.abstractdetermine if the mail is spam or not, it will improve the
dc.description.abstractperformance of the filter. With some simple modifications, the
dc.description.abstractfilter can be used in both server and client end. The efficiency
dc.description.abstractincreases gradually with the increased number of email it
dc.description.abstractprocesses.
dc.identifier.otherhttp://103.99.128.19:8080/jspui/handle/123456789/337
dc.identifier.urihttp://103.99.128.19:8080/xmlui/handle/123456789/337
dc.publisherEWU
dc.sourceCUET Digital Repository
dc.subjectSpam
dc.subjectBayesian Approach
dc.subjectSVM
dc.subjectTokenization
dc.subjectSpamicity
dc.subjectDataset
dc.titleA Modified Naïve Bayesian-based Spam Filter using Support Vector Machine
dc.title.alternative1st International Conference on Advances in Science, Engineering and Robotics Technology (ICASERT 2019)
dc.title.alternativeICASERT 2019

Files

Original bundle

Now showing 1 - 1 of 1
Thumbnail Image
Name:
A Modified Naïve Bayesian-based Spam Filter.pdf
Size:
1.1 MB
Format:
Adobe Portable Document Format