Video Enriched Retrieval Augmented Generation Using Aligned Video Captions

In this work, we propose the use of "aligned visual captions" as a mechanism for integrating information contained within videos into retrieval augmented generation (RAG) based chat assistant systems. These captions are able to describe the visual and audio content of videos in a large corpus while having the advantage of being in a textual format that is both easy to reason about & incorporate into large language model (LLM) prompts, but also typically require less multimedia content to be inserted into the multimodal LLM context window, where typical configurations can aggressively fill up the context window by sampling video frames from the source video. Furthermore, visual captions can be adapted to specific use cases by prompting the original foundational model / captioner for particular visual details or fine tuning. In hopes of helping advancing progress in this area, we curate a dataset and describe automatic evaluation procedures on common RAG tasks.

Paper

References (12)

082023. Video-ChatGPT: Towards Detailed Video Understanding via Large Vision and Language ModelsarXiv
092023. VideoChat: Chat-Centric Video UnderstandingarXiv preprint
102023. Retrieving-to-Answer: Zero-Shot Video Question Answering with Frozen Large Language Models2023 IEEE/CVF International Conference on Computer Vision Workshops (ICCVW)
112023. BLIP-2:Bootstrapping Language-Image Pre-training with Frozen Image Encoders and Large Language ModelsICML
122024. Retrieval-Augmented Egocentric Video Captioning2024 IEEE/CVF ConferenceonComputerVisionandPatternRecognition(CVPR)

Similar papers

© 2026 NYSGPT2525 LLC