Chained semantic retrieval for rare disease and gene identification using clinical phenotype

dc.contributor.advisorAlam, Md. Golam Rabiul
dc.contributor.authorSaiful, Md.
dc.date.accessioned2026-03-04T04:28:09Z
dc.date.available2026-03-04T04:28:09Z
dc.date.issued2025-10
dc.descriptionCataloged from PDF version of thesis.
dc.descriptionIncludes bibliographical references (pages 57-58).
dc.descriptionThis thesis is submitted in partial fulfillment of the requirements for the degree of Master of Science in Computer Science, 2025.
dc.description.abstractAccurate identification of rare diseases and associated genes from patient phenotypic data represents a critical challenge in precision medicine and genomic research. Current diagnostic approaches face substantial limitations in scalability, interpretability, and real-time processing when integrating heterogeneous phenotypic and genotypic databases. This study presents a novel artificial intelligence-driven system employing chained semantic retrieval and knowledge graph integration to identify rare diseases, genes, and Human Phenotype Ontology (HPO) terms from patient phenotype descriptions. The methodology leverages sentence transformers (based on Bidirectional and Auto-Regressive Transformers) for generating vector embeddings, Facebook AI Similarity Search (FAISS) for efficient similarity computation, and a fine-tuned Llama 3.2 model integrated with a Retrieval-Augmented Generation (RAG) pipeline. The system implements a tiered scoring mechanism that chains retrieval across Human Phenotype Ontology, disease, and gene databases to progressively refine predictions through contextual enhancement. Evaluation on 50 patient phenotypes with confirmed Duchenne Muscular Dystrophy diagnosis, consisting of 30 true positive and 20 true negative cases, demonstrated strong performance with 28 true positives, 17 true negatives, 2 false negatives, and 3 false positives in top-ten retrieval results. The system achieved 90 % overall accuracy, 93.3 % recall, 85 % specificity, 90.3 % precision, and an F1-score of 91.8 %, with an average computational efficiency of 3.2 seconds per response. The proposed framework effectively addresses critical gaps in cross-database integration while maintaining interpretability through tiered confidence scoring for clinical decision support applications.
dc.identifier.otherID 24266049
dc.identifier.otherhttps://dspace.bracu.ac.bd/server/api/core/items/a99ea401-4efe-4313-a14e-1a472ac9206e
dc.identifier.urihttp://hdl.handle.net/10361/27585
dc.language.isoen
dc.publisherBRAC University
dc.sourceBRAC University Institutional Repository
dc.subjectRare diseases
dc.subjectPhenotype embeddings
dc.subjectGene identification
dc.subjectHuman phenotype ontology
dc.subjectSemantic retrieval
dc.subjectGenotype-phenotype mapping
dc.subjectData integration
dc.subjectClinical data
dc.subjectLarge language models
dc.titleChained semantic retrieval for rare disease and gene identification using clinical phenotype
dc.typeThesis

Files

Original bundle

Now showing 1 - 1 of 1
Thumbnail Image
Name:
24266049_CSE.pdf
Size:
744.15 KB
Format:
Adobe Portable Document Format