A cross-modal attention-based multimodal deep learning framework for early prediction of adolescent mental health disorder

dc.contributor.advisorRabiul Alam, Md. Golam
dc.contributor.authorTaosif, Md.
dc.contributor.authorChaman, Ummay Maimona
dc.contributor.authorProva, Nazifa Anjum
dc.contributor.authorTaher, Sidrat Moon
dc.date.accessioned2026-04-20T08:35:31Z
dc.date.available2026-04-20T08:35:31Z
dc.date.issued2025-06
dc.descriptionCataloged from PDF version of thesis.
dc.descriptionIncludes bibliographical references (pages 73-77).
dc.descriptionThis thesis is submitted in partial fulfillment of the requirements for the degree of Bachelor of Science in Computer Science and Engineering, 2025.
dc.description.abstractAdolescent mental health conditions are frequently underdiagnosed due to the limitations of single diagnostic techniques and the reduced efficiency of procedures that fail to combine biological, psychological, and social aspects. We introduce a cross-modal attention-based multimodal deep learning framework that uses structural brain imaging and phenotypic data to identify anxiety disorders in the early stages. The proposed method incorporates the QTAB dataset to concurrently model T1-weighted magnetic resonance imaging (T1w MRI), behavioral questionnaire responses, and demographic variables, revealing related patterns of risk across modalities. The framework consists of three independently trained modality-specific encoders: a 3D convolutional neural network for structural MRI that generates 768- dimensional neuroanatomical representations, and two self-attention enhanced prototypical learning modules for behavioral and demographic data that generate discriminative metric embeddings of 32 and 64 dimensions, respectively. Cross-modal integration happens at the decision level, allowing for more reliable fusion in sample sizes that are limited. The experimental evaluation shows that the proposed framework obtains an AUC of 0.8935 for predicting anxiety disorders, representing a 15.1% improvement over the most robust unimodal baseline (questionnaire: AUC = 0.7766), with 85.7% sensitivity and 87.3% specificity. The best fusion weights were 23% MRI, 63% questionnaire, and 14% demographics, demonstrating the questionnaire data’s higher predictive signal. Results demonstrate that optimized weighted late fusion of well-calibrated modality predictions surpasses intricate learnt fusion techniques, underscoring the significance of proper weight optimization and calibration in small-sample multimodal psychiatric modelling. The results illustrate the efficacy of multimodal late fusion in the early detection of anxiety risk in adolescents and offer a clinically significant approach for scalable mental health screening systems.
dc.identifier.otherID 22301178
dc.identifier.otherID 22301719
dc.identifier.otherID 22301675
dc.identifier.otherID 22301747
dc.identifier.otherhttps://dspace.bracu.ac.bd/server/api/core/items/e181874e-4c74-4135-98a3-288e15cb6ca8
dc.identifier.urihttp://hdl.handle.net/10361/27970
dc.language.isoen
dc.publisherBRAC University
dc.sourceBRAC University Institutional Repository
dc.subjectPrototypical network
dc.subjectNeuroimaging
dc.subjectTwin-split
dc.subjectMultimodal fusion
dc.subjectAdolescent mental health
dc.subjectMultimodal deep learning
dc.titleA cross-modal attention-based multimodal deep learning framework for early prediction of adolescent mental health disorder
dc.typeThesis

Files

Original bundle

Now showing 1 - 1 of 1
Thumbnail Image
Name:
22301178, 22301719, 22301675, 22301747_CSE.pdf
Size:
1.04 MB
Format:
Adobe Portable Document Format