UMML: Layout-aware Text-Image Fusion for Unified Multilingual Multimodal Learning

Memes, a popular form of online communication, often blend textual and visual elements to convey humor, opinions, and social commentary. However, they are also increasingly used to spread harmful content, including hate speech and offensive sentiments. While significant research has been conducted on sentiment analysis and hate speech detection in high-resource languages such as English, there has been little attention given to low-resource languages like Bangla, despite its large number of speakers. To bridge this gap, our paper presents a novel framework for analyzing Bangla memes by integrating text, image, and layout information to detect sentiment polarity and hate speech. Leveraging the Dual Contrastive Learning (DualCL) approach, our multimodal model addresses the complexities of combining textual and visual modalities. We utilize the MUTE and MemoSen datasets, focusing on hate speech detection and sentiment analysis. Experimental results demonstrate that our approach outperforms baseline models, highlighting the importance of incorporating cross-modal multi-head attention for improved performance in Bangla meme analysis. This research fills a critical gap in the NLP landscape for low-resource languages and contributes to safeguarding digital spaces by enabling more accurate detection of harmful content in Bangla-language memes.

Paper

Full text

PDF

UMML: Layout-aware Text-Image Fusion for Unified Multilingual Multimodal Learning

OpenAlex · Advanced Image and Video Retrieval Techniques · 2024

Abstract

Memes, a popular form of online communication, often blend textual and visual elements to convey humor, opinions, and social commentary. However, they are also increasingly used to spread harmful content, including hate speech and offensive sentiments. While significant research has been conducted on sentiment analysis and hate speech detection in high-resource languages such as English, there has been little attention given to low-resource languages like Bangla, despite its large number of speakers. To bridge this gap, our paper presents a novel framework for analyzing Bangla memes by integrating text, image, and layout information to detect sentiment polarity and hate speech. Leveraging the Dual Contrastive Learning (DualCL) approach, our multimodal model addresses the complexities of combining textual and visual modalities. We utilize the MUTE and MemoSen datasets, focusing on hate speech detection and sentiment analysis. Experimental results demonstrate that our approach outperforms baseline models, highlighting the importance of incorporating cross-modal multi-head attention for improved performance in Bangla meme analysis. This research fills a critical gap in the NLP landscape for low-resource languages and contributes to safeguarding digital spaces by enabling more accurate detection of harmful content in Bangla-language memes.

Similar papers

© 2026 NYSGPT2525 LLC