Machine learning dataset for criminal law suggestions using case studies within the context of Bangladesh

dc.contributor.advisorMostakim, Moin
dc.contributor.authorTasfin, Ahnaf Arif
dc.contributor.authorIslam, Wasif
dc.date.accessioned2025-05-22T03:11:53Z
dc.date.available2025-05-22T03:11:53Z
dc.date.issued2025-02
dc.descriptionCataloged from PDF version of internship report.
dc.descriptionIncludes bibliographical references (pages 37-38).
dc.descriptionThis internship report is submitted in partial fulfillment of the requirements for the degree of Bachelor of Science in Computer Science and Engineering, 2025.
dc.description.abstractWith the growth of Language Modeling and the upcoming natural language processing assisted tools which aim for text generation, can someday render bureaucracy mean ingless while also decreasing human workload. To bring about that day, we need datasets which are able to train those models. Especially in the case of Bangladesh where there are very few datasets based on Bangladesh’s legal case studies. In this paper, we have created a Multi-Label and Binary classification dataset through data augmentation using verified case studies from the Manupatra database with the criminal law subject within the context of Bangladesh and have tested them with a few models such as DistilBERT, BERT, GPT-2, GPT-3, XLNet and Mamba to classify the acts involved and court in a case study for a 2000 data split and 3000 data split. So far, we have collected 3000 criminal case studies for our augmented dataset. Experimental results showed that out of all the models, Mamba performed the best while GPT-2 came up with the worst results. DistilBERT showed almost similar results to BERT and XLNet despite their computational differences during the benchmarking process of our augmented dataset.
dc.identifier.otherID 21101156
dc.identifier.otherID 21101199
dc.identifier.otherhttps://dspace.bracu.ac.bd/server/api/core/items/2a175dc3-a391-4638-9e65-8b3714909cf5
dc.identifier.urihttp://hdl.handle.net/10361/25974
dc.language.isoen
dc.publisherBRAC University
dc.sourceBRAC University Institutional Repository
dc.subjectMachine learning
dc.subjectNatural language processing
dc.subjectBidirectional encoder representations
dc.subjectNeural networks
dc.subjectManupatra
dc.titleMachine learning dataset for criminal law suggestions using case studies within the context of Bangladesh
dc.typeInternship Report

Files

Original bundle

Now showing 1 - 1 of 1
Thumbnail Image
Name:
21101156,21101199_CSE.pdf
Size:
261.7 KB
Format:
Adobe Portable Document Format