MOMENTA: A Multimodal Framework for Detecting Harmful Memes and Their Targets

Internet memes have become powerful means to transmit political,\npsychological, and socio-cultural ideas. Although memes are typically humorous,\nrecent days have witnessed an escalation of harmful memes used for trolling,\ncyberbullying, and abuse. Detecting such memes is challenging as they can be\nhighly satirical and cryptic. Moreover, while previous work has focused on\nspecific aspects of memes such as hate speech and propaganda, there has been\nlittle work on harm in general. Here, we aim to bridge this gap. We focus on\ntwo tasks: (i)detecting harmful memes, and (ii)identifying the social entities\nthey target. We further extend a recently released HarMeme dataset, which\ncovered COVID-19, with additional memes and a new topic: US politics. To solve\nthese tasks, we propose MOMENTA (MultimOdal framework for detecting harmful\nMemEs aNd Their tArgets), a novel multimodal deep neural network that uses\nglobal and local perspectives to detect harmful memes. MOMENTA systematically\nanalyzes the local and the global perspective of the input meme (in both\nmodalities) and relates it to the background context. MOMENTA is interpretable\nand generalizable, and our experiments show that it outperforms several strong\nrivaling approaches.\n

Paper

References (75)

Scroll for more · 38 remaining

Similar papers

© 2026 NYSGPT2525 LLC