Early and late fusion ensemble methods for predicting bug severity from bug report

dc.contributor.advisorAzmain, Md. Aquib
dc.contributor.authorFarooq, Md. Farhan
dc.contributor.authorNabil, MD. Shahariar Nawshad
dc.date.accessioned2025-12-29T06:13:41Z
dc.date.available2025-12-29T06:13:41Z
dc.date.issued2025-06
dc.descriptionCataloged from PDF version of thesis.
dc.descriptionIncludes bibliographical references (pages 47-49).
dc.descriptionThis thesis is submitted in partial fulfillment of the requirements for the degree of Bachelor of Science in Computer Science, 2025.
dc.description.abstractAccurate prediction of software bug severity is essential for optimizing resource allocation, enhancing bug triaging, and improving project management within the software development lifecycle. This study introduces a robust methodology for predicting bug severity by leveraging textual data from bug reports, employing advanced natural language processing (NLP) techniques and machine learning models. We evaluate several approaches, including Word2Vec with XGBoost (68% accuracy, 64% precision, 68% recall), TF-IDF with Logistic Regression/SVM (77% F1 score), DistilBERT (73% accuracy, 70% F1 score), and DistilRoBERTa (76% accuracy, 73% F1 score), each demonstrating strengths in capturing semantic and contextual nuances of bug descriptions. To further improve performance, we propose a fusion-based ensemble learning framework, combining early fusion (integrating TFIDF, Word2Vec, and transformer embeddings into a unified feature vector) and late fusion (aggregating predictions from independently trained models). The hybrid Ensemble Fusion model achieves the highest performance, with an accuracy of 79% and an F1 score of 76%, excelling in generalizing across diverse bug severity, including challenging short and long durations. Our methodology encompasses rigorous data preprocessing, feature engineering, and techniques to mitigate class imbalance, utilizing a comprehensive dataset of bug reports with rich textual and metadata attributes. The results underscore the efficacy of integrating diverse feature representations and model predictions, providing a scalable, robust, and actionable solution for predicting bug severity, ultimately enhancing software development efficiency and reliability.
dc.identifier.otherID 20101083
dc.identifier.otherID 20201191
dc.identifier.otherhttps://dspace.bracu.ac.bd/server/api/core/items/98718a6d-d027-4c36-b962-4bacde0ab005
dc.identifier.urihttp://hdl.handle.net/10361/27380
dc.language.isoen
dc.publisherBRAC University
dc.sourceBRAC University Institutional Repository
dc.subjectNLP
dc.subjectNatural language processing
dc.subjectMachine learning
dc.subjectBug reports
dc.subjectEnsemble fusion
dc.subjectLate fusion
dc.subjectEarly fusion
dc.subjectSoftware development
dc.subjectSoftware testing
dc.subjectBug severity prediction
dc.titleEarly and late fusion ensemble methods for predicting bug severity from bug report
dc.typeThesis

Files

Original bundle

Now showing 1 - 1 of 1
Thumbnail Image
Name:
20101083, 20201191_CSE.pdf
Size:
656.58 KB
Format:
Adobe Portable Document Format