English
Related papers

Related papers: A CLIP-based siamese approach for meme classificat…

200 papers

State-of-the-art image and text classification models, such as Convolutional Neural Networks and Transformers, have long been able to classify their respective unimodal reasoning satisfactorily with accuracy close to or exceeding human…

Machine Learning · Computer Science 2022-12-20 Weijun Jin , Lance Wilhelm

Memes are a widely popular tool for web users to express their thoughts using visual metaphors. Understanding memes requires recognizing and interpreting visual metaphors with respect to the text inside or around the meme, often while…

Computation and Language · Computer Science 2023-05-24 EunJeong Hwang , Vered Shwartz

Image clustering is an important and open-challenging task in computer vision. Although many methods have been proposed to solve the image clustering task, they only explore images and uncover clusters according to the image features, thus…

Computer Vision and Pattern Recognition · Computer Science 2023-04-11 Shaotian Cai , Liping Qiu , Xiaojun Chen , Qin Zhang , Longteng Chen

Social media memes are a challenging domain for hate detection because they intertwine visual and textual cues into culturally nuanced messages. To tackle these challenges, we introduce TRACE, a hierarchical multimodal framework that…

Computer Vision and Pattern Recognition · Computer Science 2025-11-10 Girish A. Koushik , Helen Treharne , Aditya Joshi , Diptesh Kanojia

Unsupervised adaptation of CLIP-based vision-language models (VLMs) for fine-grained image classification requires sensitivity to microscopic local cues. While CLIP exhibits strong zero-shot transfer, its reliance on coarse global features…

Computer Vision and Pattern Recognition · Computer Science 2025-10-03 Sathira Silva , Eman Ali , Chetan Arora , Muhammad Haris Khan

The large-scale pretrained model CLIP, trained on 400 million image-text pairs, offers a promising paradigm for tackling vision tasks, albeit at the image level. Later works, such as DenseCLIP and LSeg, extend this paradigm to dense…

Computer Vision and Pattern Recognition · Computer Science 2023-10-12 Ke Jin , Wankou Yang

Hateful memes have emerged as a significant concern on the Internet. Detecting hateful memes requires the system to jointly understand the visual and textual modalities. Our investigation reveals that the embedding space of existing…

Computation and Language · Computer Science 2024-10-31 Jingbiao Mei , Jinghong Chen , Weizhe Lin , Bill Byrne , Marcus Tomalin

Recently, there has been a surge of interest in applying deep learning techniques to animal behavior recognition, particularly leveraging pre-trained visual language models, such as CLIP, due to their remarkable generalization capacity…

Computer Vision and Pattern Recognition · Computer Science 2025-05-02 Enmin Zhong , Carlos R. del-Blanco , Daniel Berjón , Fernando Jaureguizar , Narciso García

In this study, we propose feature extraction for multimodal meme classification using Deep Learning approaches. A meme is usually a photo or video with text shared by the young generation on social media platforms that expresses a…

Artificial Intelligence · Computer Science 2022-07-08 Sofiane Ouaari , Tsegaye Misikir Tashu , Tomas Horvath

Image-with-text memes combine text with imagery to achieve comedy, but in today's world, they also play a pivotal role in online communication, influencing politics, marketing, and social norms. A "meme template" is a preexisting layout or…

Computers and Society · Computer Science 2024-08-16 Levente Murgás , Marcell Nagy , Kate Barnes , Roland Molontay

Concept-based approaches, which aim to identify human-understandable concepts within a model's internal representations, are a promising method for interpreting embeddings from deep neural network models, such as CLIP. While these…

Machine Learning · Computer Science 2025-06-18 Jitian Zhao , Chenghui Li , Frederic Sala , Karl Rohe

The dissemination of hateful memes online has adverse effects on social media platforms and the real world. Detecting hateful memes is challenging, one of the reasons being the evolutionary nature of memes; new hateful memes can emerge by…

Social and Information Networks · Computer Science 2023-07-10 Yiting Qu , Xinlei He , Shannon Pierson , Michael Backes , Yang Zhang , Savvas Zannettou

Memes are one of the most popular types of content used to spread information online. They can influence a large number of people through rhetorical and psychological techniques. The task, Detection of Persuasion Techniques in Texts and…

Computation and Language · Computer Science 2021-06-02 Kshitij Gupta , Devansh Gautam , Radhika Mamidi

Warning: This paper contains memes that may be offensive to some readers. Multimodal Internet Memes are now a ubiquitous fixture in online discourse. One strand of meme-based research is the classification of memes according to various…

Machine Learning · Computer Science 2025-07-04 Muzhaffar Hazman , Susan McKeever , Josephine Griffith

This paper describes the participation of our NUAA-QMUL-AIIT team in the Memotion 3 shared task on meme emotion analysis. We propose a novel multi-modal fusion method, Squeeze-and-Excitation Fusion (SEFusion), and embed it into our system…

Computation and Language · Computer Science 2023-02-17 Xiaoyu Guo , Jing Ma , Arkaitz Zubiaga

As a critical psychological stress response, micro-expressions (MEs) are fleeting and subtle facial movements revealing genuine emotions. Automatic ME recognition (MER) holds valuable applications in fields such as criminal investigation…

Human-Computer Interaction · Computer Science 2025-05-12 Shifeng Liu , Xinglong Mao , Sirui Zhao , Peiming Li , Tong Xu , Enhong Chen

Traditional online content moderation systems struggle to classify modern multimodal means of communication, such as memes, a highly nuanced and information-dense medium. This task is especially hard in a culturally diverse society like…

Computer Vision and Pattern Recognition · Computer Science 2025-03-11 Cao Yuxuan , Wu Jiayang , Alistair Cheong Liang Chuen , Bryan Shan Guanrong , Theodore Lee Chong Jen , Sherman Chann Zhi Shen

CLIP has shown promising performance across many short-text tasks in a zero-shot manner. However, limited by the input length of the text encoder, CLIP struggles on under-stream tasks with long-text inputs ($>77$ tokens). To remedy this…

Computer Vision and Pattern Recognition · Computer Science 2025-07-30 Bingchao Wang , Zhiwei Ning , Jianyu Ding , Xuanang Gao , Yin Li , Dongsheng Jiang , Jie Yang , Wei Liu

CLIP has emerged as a powerful multimodal model capable of connecting images and text through joint embeddings, but to what extent does it 'see' the same way humans do - especially when interpreting artworks? In this paper, we investigate…

Computer Vision and Pattern Recognition · Computer Science 2025-10-29 Andrea Asperti , Leonardo Dessì , Maria Chiara Tonetti , Nico Wu

Semantic compression, a compression scheme where the distortion metric, typically MSE, is replaced with semantic fidelity metrics, tends to become more and more popular. Most recent semantic compression schemes rely on the foundation model…

Image and Video Processing · Electrical Eng. & Systems 2024-12-09 Tom Bachard , Thomas Maugey