Evaluating Retrieval-Augmented Generation Variants for Biomedical Multiple-Choice Question Answering
Date
2025-10-25
Journal Title
Journal ISSN
Volume Title
Publisher
Department of Computer Science and Engineering(CSE), Islamic University of Technology(IUT), Board Bazar, Gazipur-1704, Bangladesh
Abstract
The scarcity of reliable biomedical Question-Answering (QA) resources in low-resource
languages like Bangla, spoken by over 230 million, limits access to accurate medical
knowledge. Large Language Models (LLMs) like GPT and LLaMA often produce
factually incorrect answers due to static training data, a critical issue in medical
QA where accuracy is vital. To address this, we evaluated Retrieval-Augmented
Generation (RAG) to combine external knowledge retrieval with generative rea
soning, enhancing answer verifiability. We introduce BanglaMedQA, a curated
dataset of 1,000 medical Multiple Choice Questions (MCQs) with rationales from
Bangladeshi entrance exams (MBBS, BDS, AFMC) spanning from 1990–2024, and
Bangla MMedBench, atranslated MMedBench dataset for complex reasoning. Using
a Bangla biology textbook corpus and web search, we tested RAG variants- Tra
ditional, Zero-Shot Fallback, Agentic, Iterative Feedback, and Aggregate retrieval
with LLMs. Agentic RAG achieved the highest accuracy (89.54% with openai/gpt
oss-120b), followed by Iterative RAG (88.73%) across the BanglaMedQA dataset,
with every method outperforming Zero-Shot. Despite challenges in translation fi
delity and retrieval robustness, this work enhances Bangla medical QA accuracy and
accessibility.
Description
Supervised by
Mr. Tareque Mohmud Chowdhury,
Assistant Professor,
Department of Computer Science and Engineering (CSE)
Islamic University of Technology (IUT)
Board Bazar, Gazipur, Bangladesh
This thesis is submitted in partial fulfillment of the requirement for the degree of Bachelor of Science in Computer Science and Engineering, 2025
