English
Related papers

Related papers: MER-Bench: A Comprehensive Benchmark for Multimoda…

200 papers

The main task of Multimodal Emotion Recognition in Conversations (MERC) is to identify the emotions in modalities, e.g., text, audio, image and video, which is a significant development direction for realizing machine intelligence. However,…

Sound · Computer Science 2023-12-12 Tao Meng , Yuntao Shou , Wei Ai , Nan Yin , Keqin Li

With the integration of multimodal large language models (MLLMs) into robotic systems and AI applications, embedding emotional intelligence (EI) capabilities is essential for enabling these models to perceive, interpret, and respond to…

Computation and Language · Computer Science 2026-04-28 He Hu , Lianzhong You , Hongbo Xu , Qianning Wang , Fei Richard Yu , Fei Ma , Zebang Cheng , Zheng Lian , Yucheng Zhou , Laizhong Cui

Conventional Multi-modal multi-label emotion recognition (MMER) assumes complete access to visual, textual, and acoustic modalities. However, real-world multi-party settings often violate this assumption, as non-speakers frequently lack…

Computer Vision and Pattern Recognition · Computer Science 2025-09-03 Xudong Yang , Yizhang Zhu , Hanfeng Liu , Zeyi Wen , Nan Tang , Yuyu Luo

Multimodal retrieval is becoming a crucial component of modern AI applications, yet its evaluation lags behind the demands of more realistic and challenging scenarios. Existing benchmarks primarily probe surface-level semantic…

Information Retrieval · Computer Science 2025-10-01 Junjie Zhou , Ze Liu , Lei Xiong , Jin-Ge Yao , Yueze Wang , Shitao Xiao , Fenfen Lin , Miguel Hu Chen , Zhicheng Dou , Siqi Bao , Defu Lian , Yongping Xiong , Zheng Liu

While text-to-image models like DALLE-3 and Stable Diffusion are rapidly proliferating, they often encounter challenges such as hallucination, bias, and the production of unsafe, low-quality output. To effectively address these issues, it…

In the past few years, there has been a surge of interest in multi-modal problems, from image captioning to visual question answering and beyond. In this paper, we focus on hate speech detection in multi-modal memes wherein memes pose an…

Computer Vision and Pattern Recognition · Computer Science 2021-01-01 Abhishek Das , Japsimar Singh Wahi , Siyao Li

Memes are an increasingly prevalent element of online discourse in social networks, especially among young audiences. They carry ideas and messages that range from humorous to hateful, and are widely consumed. Their potentially high impact…

Interleaved multimodal comprehension and generation, enabling models to produce and interpret both images and text in arbitrary sequences, have become a pivotal area in multimodal learning. Despite significant advancements, the evaluation…

Computer Vision and Pattern Recognition · Computer Science 2025-04-01 Peng Xia , Siwei Han , Shi Qiu , Yiyang Zhou , Zhaoyang Wang , Wenhao Zheng , Zhaorun Chen , Chenhang Cui , Mingyu Ding , Linjie Li , Lijuan Wang , Huaxiu Yao

We introduce CEMTM, a context-enhanced multimodal topic model designed to infer coherent and interpretable topic structures from both short and long documents containing text and images. CEMTM builds on fine-tuned large vision language…

Computation and Language · Computer Science 2025-10-07 Amirhossein Abaskohi , Raymond Li , Chuyuan Li , Shafiq Joty , Giuseppe Carenini

Knowledge editing (KE) provides a scalable approach for updating factual knowledge in large language models without full retraining. While previous studies have demonstrated effectiveness in general domains and medical QA tasks, little…

Artificial Intelligence · Computer Science 2025-08-12 Shengtao Wen , Haodong Chen , Yadong Wang , Zhongying Pan , Xiang Chen , Yu Tian , Bo Qian , Dong Liang , Sheng-Jun Huang

Memes have gained popularity as a means to share visual ideas through the Internet and social media by mixing text, images and videos, often for humorous purposes. Research enabling automated analysis of memes has gained attention in recent…

Computer Vision and Pattern Recognition · Computer Science 2023-02-17 Xiaoyu Guo , Jing Ma , Arkaitz Zubiaga

Multimodal emotion recognition (MER) aims to identify human emotions by combining data from various modalities such as language, audio, and vision. Despite the recent advances of MER approaches, the limitations in obtaining extensive…

Computer Vision and Pattern Recognition · Computer Science 2025-06-03 Yehun Song , Sunyoung Cho

Multimodal punchlines, which involve humor or sarcasm conveyed in image-caption pairs, are a popular way of communication on online multimedia platforms. With the rapid development of multimodal large language models (MLLMs), it is…

Computer Vision and Pattern Recognition · Computer Science 2025-06-18 Kun Ouyang , Yuanxin Liu , Shicheng Li , Yi Liu , Hao Zhou , Fandong Meng , Jie Zhou , Xu Sun

Recent advances in multi-modal generative models have driven substantial improvements in image editing. However, current generative models still struggle with handling diverse and complex image editing tasks that require implicit reasoning,…

Computer Vision and Pattern Recognition · Computer Science 2025-11-25 Feng Han , Yibin Wang , Chenglin Li , Zheming Liang , Dianyi Wang , Yang Jiao , Zhipeng Wei , Chao Gong , Cheng Jin , Jingjing Chen , Jiaqi Wang

Multimodal Large Language Models (MLLMs) have demonstrated significant advances in visual understanding tasks. However, their capacity to comprehend human-centric scenes has rarely been explored, primarily due to the absence of…

Computer Vision and Pattern Recognition · Computer Science 2025-10-16 Yuansen Liu , Haiming Tang , Jinlong Peng , Jiangning Zhang , Xiaozhong Ji , Qingdong He , Wenbin Wu , Donghao Luo , Zhenye Gan , Junwei Zhu , Yunhang Shen , Chaoyou Fu , Chengjie Wang , Xiaobin Hu , Shuicheng Yan

With the increasing prevalence of multimodal content on social media, sentiment analysis faces significant challenges in effectively processing heterogeneous data and recognizing multi-label emotions. Existing methods often lack effective…

Computation and Language · Computer Science 2025-08-26 Xilai Xu , Zilin Zhao , Chengye Song , Zining Wang , Jinhe Qiang , Jiongrui Yan , Yuhuai Lin

Recent advances in creative AI have enabled the synthesis of high-fidelity images and videos conditioned on language instructions. Building on these developments, text-to-video diffusion models have evolved into embodied world models (EWMs)…

Robotics · Computer Science 2025-05-20 Hu Yue , Siyuan Huang , Yue Liao , Shengcong Chen , Pengfei Zhou , Liliang Chen , Maoqing Yao , Guanghui Ren

Multimodal conversation, a crucial form of human communication, carries rich emotional content, making the exploration of the causes of emotions within it a research endeavor of significant importance. However, existing research on the…

Computation and Language · Computer Science 2025-06-05 Lin Wang , Xiaocui Yang , Shi Feng , Daling Wang , Yifei Zhang , Zhitao Zhang

Built on the power of LLMs, numerous multimodal large language models (MLLMs) have recently achieved remarkable performance on various vision-language tasks. However, most existing MLLMs and benchmarks primarily focus on single-image input…

Computer Vision and Pattern Recognition · Computer Science 2024-10-10 Haowei Liu , Xi Zhang , Haiyang Xu , Yaya Shi , Chaoya Jiang , Ming Yan , Ji Zhang , Fei Huang , Chunfeng Yuan , Bing Li , Weiming Hu

Hateful memes often require compositional multimodal reasoning: the image and text may appear benign in isolation, yet their interaction conveys harmful intent. Although thinking-based multimodal large language models (MLLMs) have recently…

Computation and Language · Computer Science 2026-03-03 Mohamed Bayan Kmainasi , Mucahid Kutlu , Ali Ezzat Shahroor , Abul Hasnat , Firoj Alam
‹ Prev 1 4 5 6 7 8 10 Next ›