A Visual-Language Agent for Summarizing Surveillance Footage Using Vision Transformers and Agentic LLM

The surveillance infrastructure system expansion leads to the creation of excessive video material which needs constant observation. The process of checking surveillance video needs too much time, creates too many mistakes, and cannot support immediate operational choices. This research introduces an automated vision-language system which transforms surveillance footage into brief human-readable text summaries that show main events. The system combines a spatio-temporal vision transformer which extracts visual elements and detects events with an agentic large language model that enhances event descriptions through repeated cognitive processes. The system processes video streams by dividing them into brief segments which allow activity and location detection before generating text summaries that nontechnical users can understand. The framework shows high performance during testing which includes its three main functions, where it extracts features and detects events and creates complete summaries. The proposed framework demonstrates reliable temporal event detection through its transformer-based modeling which achieves an F1-score of 0.60. The summarization module produces high-quality textual outputs which result in ROUGE score of 0.84. The new method enables surveillance systems to operate without human monitoring because it improves security for smart surveillance systems.

Paper

The full text of this publication is not hosted on 44B due to licensing.

Read it at OpenAlex

Similar papers

© 2026 NYSGPT2525 LLC