Multimodal Memorability: Modeling Effects of Semantics and Decay on Video Memorability

A key capability of an intelligent system is deciding when events from past\nexperience must be remembered and when they can be forgotten. Towards this\ngoal, we develop a predictive model of human visual event memory and how those\nmemories decay over time. We introduce Memento10k, a new, dynamic video\nmemorability dataset containing human annotations at different viewing delays.\nBased on our findings we propose a new mathematical formulation of memorability\ndecay, resulting in a model that is able to produce the first quantitative\nestimation of how a video decays in memory over time. In contrast with previous\nwork, our model can predict the probability that a video will be remembered at\nan arbitrary delay. Importantly, our approach combines visual and semantic\ninformation (in the form of textual captions) to fully represent the meaning of\nevents. Our experiments on two video memorability benchmarks, including\nMemento10k, show that our model significantly improves upon the best prior\napproach (by 12% on average).\n

Paper

Similar papers

© 2026 NYSGPT2525 LLC