A Hybrid Deep Learning Framework for Multi-Camera Object Tracking and Anomaly Detection in Surveillance

Modern surveillance systems face increasing challenges in real-time object tracking, behavioral analysis, and anomaly detection across distributed camera networks. Traditional single-camera solutions fail to address spatial discontinuity and occlusion issues in multi-camera environments. To overcome these challenges, this paper proposes a hybrid deep learning framework integrating Convolutional Neural Networks (CNNs) and Recurrent Neural Networks (RNNs) with attention-based data fusion for multi-camera object tracking and anomaly detection in large-scale surveillance networks. The proposed framework employs a CNN-based feature extractor to identify and encode object characteristics across multiple camera streams, followed by a spatio-temporal attention module that aligns overlapping regions and correlates object trajectories between cameras. The RNN-LSTM network models temporal dependencies to track object motion patterns, while an autoencoder-based anomaly detector identifies abnormal behaviors such as unattended baggage, unauthorized entry, or erratic movement. A multi-view consistency algorithm ensures smooth cross-camera transition tracking. Experimental evaluation on PETS 2009, TownCentre, and Avenue datasets demonstrates the framework’s superiority, achieving a tracking accuracy (MOTA) of 94.7%, an IDF1 score of 92.5%, and an anomaly detection accuracy of 96.2%, outperforming existing CNN-LSTM and transformer-based baselines. The proposed hybrid framework enhances situational awareness, robustness, and scalability, making it suitable for deployment in smart cities, public safety monitoring, and industrial surveillance.

Paper

The full text of this publication is not hosted on 44B due to licensing.

Read it at OpenAlex

Similar papers

© 2026 NYSGPT2525 LLC