Automated species detection improves wildlife monitoring with attention models

Automated Species Identification in Camera Trap Images for Wildlife Conservation

Computer Vision and Pattern Recognition

Summary

Identifying animals in photos taken by camera traps is important for protecting wildlife but hard when animals are small or blend into the background. The authors created a new method that uses a special attention-based computer vision model combined with a large language model to better spot these animals and understand their behavior from images. This approach works well even on hard-to-see animals and can identify animals it has never seen before. Their system also provides useful insights into animal actions that can help conservation efforts.

What this means in practice

  • For wildlife monitoring teams: Improve detection accuracy of small and camouflaged animals in camera trap photos used for habitat and species monitoring.
  • For environmental consulting companies: Provide enhanced automated reports on wildlife presence and behavior from trap images, supporting biodiversity assessments at development sites.$Commercial implications: This enables sale of advanced wildlife analysis services to clients requiring ecological impact assessments.

Authors

Nowshin Amin, Nafisa Tabassum Oyshi, Tahmid Abrar Zidan, Miftaun Noor, Md. Abrar Rahman Shafin

Abstract

Wildlife conservation involves protecting, preserving, and managing wildlife species and their habitats. With today's rapid pace of human development, climate change, and other unsustainable practices, the need for wildlife conservation has heightened. Despite significant progress in species identification using deep-learning models, significant challenges still remain in effectively detecting small animals in low-contrast trap images due to limited feature extraction capabilities. This thesis presents a novel end-to-end framework integrating a self-attention mechanism to address these limitations. The proposed architecture involves a Swin-BiFPN backbone integrated in a Faster RCNN detection network, coupled with a visual semantic extraction module driven by the LLaVA v1.5 (13B) multimodal large language model. The detection framework, capable of extracting crucial features in challenging trap images, demonstrates consistently high results and robust generalization capabilities. Furthermore, the visual semantic extraction module provides zero-shot detection capability, as well as providing valuable insights and emergent cues of the animal's behavior, further supporting the conservation effort. The MLLM evaluation was conducted using both traditional NLP metrics (precision, recall, F1, and SBERT similarity) and subjective scoring by LLM-based judges (GPT-4.1 and GROK 3.0), across five MLLMs, demonstrating the model's strong performance in visual description generation. The proposed framework improves detection accuracy across low-contrast trap images and small animals while also demonstrating zero-shot detection capability leveraging the MLLM.