Thesis 2019

Browse

Search Results

Now showing 1 - 10 of 20
  • Thumbnail Image
    Item
    Analyzing Effect of Feature Selection in Software Fault Detection
    (East West University, 2019-12-24) Cynthia, Shamse Tasnim; Rasul, Md. Golam
    The quality of software is enormously affected by the faults associated with it. Detection of faults at a proper stage in software development is a challenging task and plays a vital role in the quality of the software. Machine learning is now a days a commonly used technique for fault detection and prediction. However, the effectiveness of the fault detection mechanism is impacted by the number of attributes presented in the dataset. This paper thoroughly gives the importance to compare between different machine learning approaches and by observing their performances we can conclude which models perform better to detect fault in the selected software modules and investigates the effect of various feature selection techniques on software fault classification by using NASA’s some benchmark publicly available datasets. Various metrics are used to analyze the performance of the feature selection and classification techniques. The experiment discovers that some particular classifiers can detect the presence of the faults more effectively and by selecting the best features and solving the class imbalance problem can ensure better quality of the software.
  • Thumbnail Image
    Item
    A Comparative Study of Hybridized Neural Networks in Estimating Traffic Accident Severity
    (East West University, 2019-12-24) Anik, Md. Mydul Islam; Akram, Wasim; Md. Ashikuzzaman
    The increasing number of populations causing increase of vehicles which leads to traffic accident. As transportation system expands, it needs to be monitored to assure safety to citizen. Cities are trying to adopt technological advancement in order to minimize traffic accident. Traffic accidents have become one of the largest national health issues and many factors like weather condition, road condition, light condition, etcetera is related to it. In the current paper, several hybridize machine learning models are used on dataset of city Leeds, UK to estimate traffic accident severity. Hybridize Machine learning models are Artificial Neural Network (ANN) with Gradient decent, Principle Component Analysis (PCA) with ANN, Genetic Algorithm with ANN, Particle Swarm Optimization with ANN. These models are also compared with other machine learning models such as Support Vector Machine (SVM), Naïve Bayes, Nearest Centroid, Logistic Regression, K Nearest Neighbor Classification and Random Forest. Comparison was done considering performance evaluation of each model’s accuracy result. Genetic Algorithm with ANN showed promising result of 86.63% accuracy which is the highest score of all model results. Whereas, Nearest Centroid Method gave 55% of accuracy resulting lowest of all. The Results and findings obtained in this study are significant which can provide invaluable information on reducing traffic accident.
  • Thumbnail Image
    Item
    Bangladeshi Vehicle License Plate Detection and Recognition
    (East West University, 2019-12-24) Hasan, Md. Mehedi; Das, Anik Kumar; Alam, Md. Zahid
    recognition (OCR) to identify vehicles using their license plate. In this project, an algorithm has been proposed to detect and read Bengali license plate. Nowadays automatic vehicle monitoring system has drawn the importance due to maintain numerous traffic that are increasing rapidly in Bangladesh. Every vehicle has unique identification (license plate).So it is very easy and simple task to track and monitor any particular vehicle using their license plate. Because it is an automatic system that save a lot of time and workers instead of tracking and monitoring manually. For these purpose Automatic Number Plate Recognition (ANPR) is designed and developed. Automatic Number plate Recognition plays a vital part for controlling and managing vehicle and vehicle related works. This automatic detection system can be efficient for toll collection system, vehicle parking system, traffic control system, vehicle-tracking system, finding lost vehicle as well as detecting guilty vehicle involved in crime for police etc. There are mainly two types license plate exist in Bangladeshi Vehicle. The green background license plate is for commercial use and the white background license plate is for personal used vehicle. These license plates have four major part ordering in two line. First line describes the area name and type of vehicle and the second line describes vehicle class and registration number. Many approaches are used for ANPR but it needs a simple method with low complexity for detecting vehicle as fast as possible. This algorithm has three major steps- detection, segmentation and recognition. For detection firstly, it uses the feature of Haar Cascade, which is a machine learning based approach where a cascade function is trained from a lot of positive and negative images. It is used to detect objects in other images. It is well known for being able to detect faces and body parts in an image, but can be trained to identify almost any object. Then boundary box approach is use to segment these characters separately for area name, type, class and registration number and convert to text file from image. Finally, with the help of Template matching algorithm it recognizes the word and number used in Bengali license plate. This system has less complexity and high efficiency for number plate detection and recognition.
  • Thumbnail Image
    Item
    Paraphrased Question Answering Using Recurrent Neural Network in Bangla Language
    (East West University, 2019-12-24) Patwary, Nazmus Sakib; Hasan, Md. Mohaiminul; Rahman, Tanvir
    Recent studies on QA (Question Answering) system in English language have been emerged extensively with the composition of NLP (Natural Language Processing) and IR (Information Retrieval) by amplifying miniature sub tasks to accomplish a whole AI-system having capability of answering and reasoning complicated and long questions through understating paragraph. In our proposed study, we present a general heuristic framework, an end-to-end model used for paraphrased question answering using single supporting line which is the initial appearance ever in Bangla language. Corpus dataset was scrapped from Bangla wiki and then questions were generated corresponding context have been used to learn the model. Translated bAbI dataset (1 supporting fact) [5][6] in Bangla language has been also incorporated with to experiment the proposed model manually. To predict appropriate answer, model is trained with question-answer pair and a supporting line. For comparing our task applying variation of basic RNN (Recurrent Neural Network): LSTM (Long Short-Term Memory) and GRU (Gated Recurrent Unit) different accuracy has been found. For further accomplishment, synthetic and semantic word relevance in high dimension vector space: Bangla Word2vec (word embedding system) is added to the system as sentence representation along with PE (Positioning Encoding) and which outperforms both memory network GRU and LSTM precisely.
  • Thumbnail Image
    Item
    A Trust and User Personality Trait based Collaborative Filtering Recommender System
    (East West University, 2019-12-24) Afrin, Afsana; Rahman, Salma; Kamal, Farha
    Recommendation system is a system where a user gets suggestions for the product based on his/her previous preferences to the items. With the monumental growth of web services, developing adequate methods for the recommendation has become dominant in the research area. In terms of Collaborative Filtering(CF) user and item-based methods are the most presiding approaches used in RS. In order to get an improved recommendation, trust value and personality traits of users can play a key role in nding similarity between users.It solves cold start problem where neighbors of the new user are di cult to nd as they have not rated any item yet. In our work, the implicit(based on personality),latent(based on trust) and explicit(based on rating) features of user behaviour have been utilized to tackle the problems of Collaborative Filtering.The major bene t of the proposed method is its consideration of direct and indirect trust values and personality similarity compared to traditional collaborative ltering approaches.A com- parative review of traditional rating based recommender system,personality based,trust based and their combined approach is presented. Empirical analysis shows that using trust propagation and personality traits substantially increases the e ciency of the CF recommender system.
  • Thumbnail Image
    Item
    Credibility Identification of Online News Portal Using Website Traffic Metrics
    (East West University, 2019-12-24) Rahman, Farhad; Ragib, Md. Ashfak
    Online news portal sites are increasing progressively as like meshwork all over the world. The propagation of misrepresenting information such as social media feeds, news blogs, and online newspapers have made it challenging to distinguish reliable news sources, thus enhancing the need for computational tools able to provide insights into the trustworthiness of online content. It’s a much difficult task to determine the authenticity of the newspaper article content directly rather than knowing the actual source reliability of the article. To distinguish reliable and unreliable media sources, an approach is proposed in this paper to scale online news portal reliability using website metrics data that has been collected through the Alexa website traffic statistics tool. Initially, all Bengali online news portal websites were listed and a dataset has been created containing all relevant website metrics information of all those websites. Using domain knowledge and context analysis vital features have been extracted for scaling reliability of those websites using an unsupervised learning algorithm. In this context, the k-means algorithm is been used for making several clusters of the unlabeled dataset. Each cluster got a label depending on the metrics information correlation in terms of the real-world scenario. Comparing the experimental result with the theoretical knowledge, the proposed approach satisfied the research intention.
  • Thumbnail Image
    Item
    Study of Influence of Dimension Reductions of High Dimensional Datasets in Classification Problem
    (East West University, 2019-11-24) Bhuiyan, Mohd. Salman Hossain; Al Raian, Nabil; Leon, Shahad Iqbal
    In our day to day life we develop many applications based on datasets. In case of high dimensional dataset, we often face some problems while building any data mining model. When there are too many attributes in the dataset, then there may be dependency between attributes. There may be some irrelevant attributes too. So, we get less accuracy in the data mining because of the influence of dependent and irrelevant attributes. So, in order to solve this problem, we need to reduce the dimensions of the dataset. In this work, we experimentally tested influence of dimension reduction on classification problems. For this purpose, we used 4 different datasets. We used backward elimination method to reduce the dimension of the dataset down to seven dimensions. We have experimented with Multi-Layer Perceptron, Naïve Bayes, Decision Tree, K-Nearest Neighbor, and Support Vector Machine classification methods. We used 10-fold validation to train and test our dataset. Experimental results show that when the dimension is reduced, then the accuracy is improved for some classification algorithm like Multi-Layer Perceptron, Naïve Bayes and Random Forest. We come up with a conclusion that if we exclude the less significant attributes, then the classification model gives better accuracy than it does without dimension reduction.
  • Thumbnail Image
    Item
    Data Clustering Using Hybrid Genetic Algorithm with k-Means and k-Medoids Algorithms
    (East West University, 2019-09-26) Islam, Md. Touhidul; Basak, Pappu Kumar; Bhowmik, Priom
    Clustering methods separate a set of data points into groups or clusters, where data points of each cluster have the similar properties and are dissimilar from those of other clusters. In general k-means and k-medoids methods are used for data clustering. These clustering methods are heuristic and may stuck in a local optimum. To avoid this problem, we propose a hybrid Genetic Algorithm (HGA) for data clustering. For this purpose, we propose a genetic encoding of the clustering problem, where data points are separated into k clusters. The cluster centers of the generated clusters are determined using the techniques of both k-means and k-medoids methods. The fitness of the clustering is calculated using the sum of Euclidian distances of each data point from its cluster center. We experiment with Iris, Seeds, and Ionosphere datasets. Experimental results show that the proposed HGA generates 2.67% to 28.68% higher clustering accuracies than the clustering accuracies previously reported in the literature.
  • Thumbnail Image
    Item
    Investigation of Different Machine Learning Approaches for Analyzing Human Attitude
    (East West University, 2019-12-24) Ahmed, Ifthakhar; Mostafa, Golam
    Sentiment analysis is widely used in data science, where data is used and analyzed, which is available in various social media and internet. Sentiment analysis is a qualitative processing of text data that extracts and defines subjective details in the source material and allows a company or something like that to recognize the feeling of its brand or products or services when tracking online conversations and feedback. Social media analysis is usually limited to simple analysis of sentiment or count based metrics. Everyone is articulate one way or another in in this time. Most social media and android apps like Facebook or WhatsApp or Twitter have a lot of information which is accessible in this highly developed and re imagined world. Twitter is one of the most common and international networks. Twitter is that kind of social media where many users can express their opinion and feelings through small tweets. These tweets can be analyzed using different machine learning algorithms. Twitter sentiment analysis is a very popular research work now. Most of the work is on two types of sentiment detection, it can be either positive or negative. This paper includes neutral sentiment. So, the proposed idea is to find out a tweet is positive or negative or neutral with a better accuracy. In this paper tweeter data has been encoded using label encoder OneHot encoder. After that applying different types of pre-processors, we have feed that numeric form to the machine learning classifier algorithms. Although our accuracy is quite low, in future we shall develop the algorithm for better accuracy.
  • Thumbnail Image
    Item
    Sentiment Analysis Based on Feature Selection and Machine Learning Techniques
    (East West University, 2019-11-24) Tabassum, Thajiba; Shohag, Md; kafi, Abdullahhil
    This research focuses on sentiment analysis of amazon food review dataset. For this work, it used some machine learning algorithms but specifically used sentiment analysis and LSTM. It was implemented by some machine learning algorithm with sentiment analysis. People are now a day more depends on restaurant food because of their busy life. For this they taste many kinds of restaurants and without knowledge they don’t know where they can get good food with good service. They don’t know which is good or which is bad and sometimes the reviews are so confusing that they do not understand. To overcome the above problems researchers have used machine learning algorithms to classify positive or negative food review with binary classification. In this study it discussed about long short term memory (LSTM) and four machine learning algorithm to solve this problem it was also use aspect based sentiment analysis. There are a lot of approaches developed for binary classification but it used Naïve Bayes (Bernoulli), Perceptron, Decision tree, Logistic regression, Long short term memory(LSTM). Recurrent Neural Networks(RNN) have exposed as widely used architectures and are united with sequence-based models. The main research aims to develop this machine learning algorithm to gives the best result for binary classification in restaurant food review. The purpose of this research is binary classification and sentiment analysis. To classify text, it has been used Amazon food review data set. It has been used some machine learning algorithms. It has been divided the work into three stages, these were data preprocessing, sentiment analysis and binary classification.