Bangla Dataset Generation for Natural Language Inference

dc.contributor.authorIslam, Md. Shohidul
dc.contributor.authorKhan, Abdun Nayeem
dc.contributor.authorNizami, Md Shaidur Rahman
dc.date.accessioned2024-09-02T05:46:02Z
dc.date.available2024-09-02T05:46:02Z
dc.date.issued2023-05-30
dc.descriptionSupervised by Dr. Hasan mahmud, Associate Professor, Prof. Dr. Kamrul Hasan, Department of Computer Science and Engineering(CSE), Islamic University of Technology(IUT), Board Bazar, Gazipur-1704, Bangladesh Board Bazar, Gazipur, Bangladesh
dc.description.abstractUnderstanding entailment and contradiction is fundamental to understanding nat ural language, and inference about entailment and contradiction is a valuable test ing ground for the development of semantic representations. However, machine learning research in this area has been dramatically limited by the lack of resources in Bangla. To address this, we propose to introduce our own corpus curated for natural language inference which is labeled pairs of sentences with a label that depicts their inner entailment. Our goal is to create a dataset that has over 30K instances and to do so we have now created a Bangla dataset by machine trans lating the SNLI corpus into Bangla. After that, we show that benchmark models can be used to evaluate and do the task of inference in Bangla . We hope that our dataset will catalyze research in Bangla sentence understanding by providing an informative standard evaluation task.For this we provided two baseline models which are both considered integral in the task of inference in any langauge.
dc.identifier.otherhttps://repository.iutoic-dhaka.edu/server/api/core/items/a92dbfd6-cfff-419c-8756-544e93966a65
dc.identifier.urihttp://hdl.handle.net/123456789/2147
dc.language.isoen
dc.publisherDepartment of Computer Science and Engineering(CSE), Islamic University of Technology(IUT), Board Bazar, Gazipur-1704, Bangladesh
dc.sourceIUT Institutional Repository
dc.subjectentailment, contradiction, neutral, natural language, inference, seman tic representations, machine learning, Bangla, corpus, labeled pairs of sentences, inner entailment, dataset, instances, SNLI corpus, machine translation, benchmark models, evaluation task, baseline models, sentence understanding
dc.titleBangla Dataset Generation for Natural Language Inference
dc.typeThesis

Files

Original bundle

Now showing 1 - 1 of 1
Thumbnail Image
Name:
Thesis 2023_CSE_180041113_180041139_170041055 - MD. SHAIDUR RAHMAN NIZAMI, 180041139.pdf
Size:
861.25 KB
Format:
Adobe Portable Document Format

Collections