Preliminary Use of Vision Language Model Driven Extraction of Mouse Behavior Towards Understanding Fear Expression

Integration of diverse data is pivotal for advancing scientific discovery. This work establishes a vision-language model (VLM) for classifying mouse behavior directly from video, producing per-second behavioral vectors for each subject with minimal user input. We use the open-source Qwen2.5-VL model and enhance performance through prompt design, in-context learning (ICL) with labeled examples, and frame-level preprocessing. Each method contributes to improved classification, and their combination yields strong F1 scores across all behaviors, including rare classes such as freezing and fleeing, without any fine-tuning. This framework enables scalable behavioral annotation that can be integrated with neural and experimental datasets to address complex research questions.

Paper

Similar papers

© 2026 NYSGPT2525 LLC