Detecting misleading information from Large Language Models responses

dc.contributor.advisorAzmain, Md. Aquib
dc.contributor.advisorAnwar, Md. Tawhid
dc.contributor.authorAdor, Muntasir Ahmed
dc.contributor.authorHasan, Fahim
dc.contributor.authorMahamud, Syed Ashik
dc.contributor.authorNazmin, Iffat Ara
dc.contributor.authorMuntasir Arin, Md. Mahim
dc.date.accessioned2025-06-18T05:08:49Z
dc.date.available2025-06-18T05:08:49Z
dc.date.issued2025-02
dc.descriptionCataloged from PDF version of thesis.
dc.descriptionIncludes bibliographical references (pages 31-34).
dc.descriptionThis thesis is submitted in partial fulfillment of the requirements for the degree of Bachelor of Science in Computer Science and Engineering, 2025.
dc.description.abstractThe arrival of large language models (LLMs) have been a game-changer in natural language processing (NLP). It revolutionized the way we comprehend and generate content. However, LLMs can hallucinate-that is, contradict the reality or input provided by the user. This is a big problem because these models are being used these days in many diverse industries, including for medical and legal purposes where accuracy is paramount. Hallucinations can damage user trust and lead to the spread of incorrect facts. Although it is not a major issue in ordinary situations, it raises serious concerns in sensitive areas like healthcare and legal advice. Also it can inadvertently become part of the training corpus for future models if not carefully filtered. This creates a feedback loop where errors in one generation of models can propagate and potentially amplify in subsequent iterations. To solve this problem, we have created an all-rounded dataset with questions from SQuAD (Stanford Question Answering Dataset), HotpotQA, and TriviaQA, among other datasets. We will use the state-of-the-art LLM GPT-4o mini to generate answers. Finally, to this end, the paper describes several rules for the annotation of a corresponding dataset, its resulting characteristic properties and the classification quality that can be achieved when using the dataset for fine-tuning different models.
dc.identifier.otherID: 20101259
dc.identifier.otherID: 20201144
dc.identifier.otherID: 20301124
dc.identifier.otherID: 21101106
dc.identifier.otherID:24141109
dc.identifier.otherhttps://dspace.bracu.ac.bd/server/api/core/items/6c199a74-5fca-47e1-a078-febd024ecea9
dc.identifier.urihttp://hdl.handle.net/10361/26078
dc.language.isoen
dc.publisherBRAC University
dc.sourceBRAC University Institutional Repository
dc.subjectMachine learning
dc.subjectNatural language processing
dc.subjectLarge language models
dc.subjectAI hallucination
dc.subjectGPT
dc.subjectTransformer
dc.subjectALBERT
dc.titleDetecting misleading information from Large Language Models responses
dc.typeThesis

Files

Original bundle

Now showing 1 - 1 of 1
Thumbnail Image
Name:
20101259, 20201144, 20301124, 21101106, 24141109_CSE.pdf
Size:
805.25 KB
Format:
Adobe Portable Document Format