Automated species identification in camera trap images for wildlife conservation

dc.contributor.advisorAlam, Md. Golam Rabiul
dc.contributor.advisorAhmed, Md Sabbir
dc.contributor.authorAmin, Nowshin
dc.contributor.authorOyshi, Nafisa Tabassum
dc.contributor.authorZidan, Tahmid Abrar
dc.contributor.authorNoor, Miftaun
dc.contributor.authorShafin, Md. Abrar Rahman
dc.date.accessioned2025-07-29T05:29:00Z
dc.date.available2025-07-29T05:29:00Z
dc.date.issued2025-06
dc.descriptionCataloged from PDF version of thesis.
dc.descriptionIncludes bibliographical references (pages 51-53).
dc.descriptionThis thesis is submitted in partial fulfillment of the requirements for the degree of Bachelor of Science in Computer Science and Engineering, 2025.
dc.description.abstractWildlife conservation involves protecting, preserving, and managing wildlife species and their habitats. With today’s rapid pace of human development, climate change, and other unsustainable practices, the need for wildlife conservation has heightened. Despite significant progress in species identification using deep-learning models, significant challenges still remain in effectively detecting small animals in low-contrast trap images due to limited feature extraction capabilities. This thesis presents a novel end-to-end framework integrating a shifted window based local self-attention mechanism along with enhanced feature fusion in a object detection head and incorporating multimodal large language model to address these limitations. The proposed architecture involves a Swin-BiFPN backbone integrated in a Faster RCNN detection network, coupled with a visual semantic extraction module driven by the LLaVA v1.5 (13B) multimodal large language model. The detection framework, capable of extracting crucial features in challenging trap images, demonstrates consistently high results and robust generalization capabilities. Furthermore, the visual semantic extraction module provides zero-shot detection capability, as well as providing valuable insights and emergent cues of the animal’s behavior, further supporting the conservation effort. The MLLM evaluation was conducted using both traditional NLP metrics (precision, recall, F1, and SBERT similarity) and subjective scoring by LLM-based judges (GPT-4.1 and GROK 3.0), across five MLLMs, demonstrating the model’s strong performance in visual description generation. The proposed framework improves detection accuracy across low-contrast trap images and small animals while also demonstrating zero-shot detection capability leveraging the MLLM.
dc.identifier.otherID 21201234
dc.identifier.otherID 21201179
dc.identifier.otherID 21201056
dc.identifier.otherID 21241021
dc.identifier.otherID 21201080
dc.identifier.otherhttps://dspace.bracu.ac.bd/server/api/core/items/e64fe11f-a5cb-4d33-8c19-9632edf6689c
dc.identifier.urihttp://hdl.handle.net/10361/26508
dc.language.isoen
dc.publisherBRAC University
dc.sourceBRAC University Institutional Repository
dc.subjectWildlife conservation
dc.subjectDeep learning
dc.subjectCNN
dc.subjectSwin Transformer
dc.subjectBiFPN
dc.subjectFaster-RCNN
dc.subjectMLLM
dc.subjectAnimal identification
dc.subjectCamera Trap Images
dc.titleAutomated species identification in camera trap images for wildlife conservation
dc.typeThesis

Files

Original bundle

Now showing 1 - 1 of 1
Thumbnail Image
Name:
21201234_21201179_21201056_21241021_21201080_CSE.pdf
Size:
8.29 MB
Format:
Adobe Portable Document Format