vix.ing · top · new · best · stats · spec

Multimodal Memorability: Modeling Effects of Semantics and Decay on\n Video Memorability

2020/09/05 by Anelise Newman, Camilo Fosco, Newman, Anelise +9 · 3 citations
Computer Science · #Advanced Image and Video Retrieval Techniques #Computer Vision and Pattern Recognition (cs.CV) #FOS: Computer and information sciences #Multimodal Machine Learning Applications #Visual Attention and Saliency Detection

paper · pdf · doi:10.48550/arxiv.2009.02568

openalex publication_date 2020/09/05 · openalex created_date 2022/07/25 · openalex updated_date 2026/07/28

Abstract

A key capability of an intelligent system is deciding when events from past\nexperience must be remembered and when they can be forgotten. Towards this\ngoal, we develop a predictive model of human visual event memory and how those\nmemories decay over time. We introduce Memento10k, a new, dynamic video\nmemorability dataset containing human annotations at different viewing delays.\nBased on our findings we propose a new mathematical formulation of memorability\ndecay, resulting in a model that is able to produce the first quantitative\nestimation of how a video decays in memory over time. In contrast with previous\nwork, our model can predict the probability that a video will be remembered at\nan arbitrary delay. Importantly, our approach combines visual and semantic\ninformation (in the form of textual captions) to fully represent the meaning of\nevents. Our experiments on two video memorability benchmarks, including\nMemento10k, show that our model significantly improves upon the best prior\napproach (by 12% on average).\n

Cited by

Related