中文
相关论文

相关论文: Moodifier: MLLM-Enhanced Emotion-Driven Image Edit…

200 篇论文

Recently, Multimodal Large Language Models (MLLMs) have achieved exceptional performance across diverse tasks, continually surpassing previous expectations regarding their capabilities. Nevertheless, their proficiency in perceiving emotions…

计算机视觉与模式识别 · 计算机科学 2026-03-03 Daiqing Wu , Dongbao Yang , Sicheng Zhao , Can Ma , Yu Zhou

Understanding emotions accurately is essential for fields like human-computer interaction. Due to the complexity of emotions and their multi-modal nature (e.g., emotions are influenced by facial expressions and audio), researchers have…

计算机视觉与模式识别 · 计算机科学 2025-01-17 Qize Yang , Detao Bai , Yi-Xing Peng , Xihan Wei

In daily life, images as common affective stimuli have widespread applications. Despite significant progress in text-driven image editing, there is limited work focusing on understanding users' emotional requests. In this paper, we…

计算机视觉与模式识别 · 计算机科学 2025-12-29 Peixuan Zhang , Shuchen Weng , Chengxuan Zhu , Binghao Tang , Zijian Jia , Si Li , Boxin Shi

Multi-modal large language models (MLLMs) have achieved remarkable performance on objective multimodal perception tasks, but their ability to interpret subjective, emotionally nuanced multimodal content remains largely unexplored. Thus, it…

计算机视觉与模式识别 · 计算机科学 2024-07-02 Qu Yang , Mang Ye , Bo Du

Previous studies on visual customization primarily rely on the objective alignment between various control signals (e.g., language, layout and canny) and the edited images, which largely ignore the subjective emotional contents, and more…

计算机视觉与模式识别 · 计算机科学 2026-02-23 Jiamin Luo , Xuqian Gu , Jingjing Wang , Jiahong Lu

Recent advances in AI-generated content (AIGC) have significantly accelerated image editing techniques, driving increasing demand for diverse and fine-grained edits. Despite these advances, existing image editing methods still face…

计算机视觉与模式识别 · 计算机科学 2025-05-27 Shuyu Wang , Weiqi Li , Qian Wang , Shijie Zhao , Jian Zhang

The recent advancement of Multimodal Large Language Models (MLLMs) is transforming human-computer interaction (HCI) from surface-level exchanges into more nuanced and emotionally intelligent communication. To realize this shift, emotion…

人工智能 · 计算机科学 2026-01-06 Hyeongseop Rha , Jeong Hun Yeo , Yeonju Kim , Yong Man Ro

Affective Image Manipulation (AIM) seeks to modify user-provided images to evoke specific emotional responses. This task is inherently complex due to its twofold objective: significantly evoking the intended emotion, while preserving the…

计算机视觉与模式识别 · 计算机科学 2025-06-19 Jingyuan Yang , Jiawei Feng , Weibin Luo , Dani Lischinski , Daniel Cohen-Or , Hui Huang

This paper introduces a multi-label visual emotion analysis benchmark dataset for comprehensively evaluating the ability of multimodal large language models (MLLMs) to predict the emotions evoked by images. Recent user studies report an…

计算机视觉与模式识别 · 计算机科学 2026-05-15 Tianwei Chen , Takuya Furusawa , Yuki Hirakawa , Ryotaro Shimizu , Mo Fan , Takashi Wada

Multimodal Affective Computing (MAC) aims to recognize and interpret human emotions by integrating information from diverse modalities such as text, video, and audio. Recent advancements in Multimodal Large Language Models (MLLMs) have…

人工智能 · 计算机科学 2025-08-05 Miaosen Luo , Jiesen Long , Zequn Li , Yunying Yang , Yuncheng Jiang , Sijie Mai

Emotional information is essential for enhancing human-computer interaction and deepening image understanding. However, while deep learning has advanced image recognition, the intuitive understanding and precise control of emotional…

计算机视觉与模式识别 · 计算机科学 2025-01-06 Junjie Xu , Xingjiao Wu , Tanren Yao , Zihao Zhang , Jiayang Bei , Wu Wen , Liang He

Sentiment and emotion understanding are essential to applications such as human-computer interaction and depression detection. While Multimodal Large Language Models (MLLMs) demonstrate robust general capabilities, they face considerable…

计算与语言 · 计算机科学 2025-07-08 Ao Li , Longwei Xu , Chen Ling , Jinghui Zhang , Pengwei Wang

Understanding the multi-dimensional attributes and intensity nuances of image-evoked emotions is pivotal for advancing machine empathy and empowering diverse human-computer interaction applications. However, existing models are still…

计算机视觉与模式识别 · 计算机科学 2026-05-26 Lancheng Gao , Ziheng Jia , Zixuan Xing , Wei Sun , Huiyu Duan , Guangtao Zhai , Xiongkuo Min

Understanding human emotions from multimodal signals poses a significant challenge in affective computing and human-robot interaction. While multimodal large language models (MLLMs) have excelled in general vision-language tasks, their…

计算机视觉与模式识别 · 计算机科学 2026-02-24 Xiaojiang Peng , Jingyi Chen , Zebang Cheng , Bao Peng , Fengyi Wu , Yifei Dong , Shuyuan Tu , Qiyu Hu , Huiting Huang , Yuxiang Lin , Jun-Yan He , Kai Wang , Zheng Lian , Zhi-Qi Cheng

Accurate emotion perception is crucial for various applications, including human-computer interaction, education, and counseling. However, traditional single-modality approaches often fail to capture the complexity of real-world emotional…

Vision-language models (VLMs) show promise as tools for inferring affect from visual stimuli at scale; it is not yet clear how closely their outputs align with human affective ratings. We benchmarked nine VLMs, ranging from state-of-the-art…

Large language models and vision-language models (which we jointly call LMs) have transformed NLP and CV, demonstrating remarkable potential across various fields. However, their capabilities in affective analysis (i.e. sentiment analysis…

计算与语言 · 计算机科学 2025-06-02 Zhiwei Liu , Lingfei Qian , Qianqian Xie , Jimin Huang , Kailai Yang , Sophia Ananiadou

Recent advances in multimodal large language models (MLLMs) have catalyzed transformative progress in affective computing, enabling models to exhibit emergent emotional intelligence. Despite substantial methodological progress, current…

Recognising emotions in context involves identifying an individual's apparent emotions while considering contextual cues from the surrounding scene. Previous approaches to this task have typically designed explicit scene-encoding…

计算机视觉与模式识别 · 计算机科学 2025-07-16 Alexandros Xenos , Niki Maria Foteinopoulou , Ioanna Ntinou , Ioannis Patras , Georgios Tzimiropoulos

In this work, we introduce Mozualization, a music generation and editing tool that creates multi-style embedded music by integrating diverse inputs, such as keywords, images, and sound clips (e.g., segments from various pieces of music or…

人机交互 · 计算机科学 2025-04-22 Wanfang Xu , Lixiang Zhao , Haiwen Song , Xinheng Song , Zhaolin Lu , Yu Liu , Min Chen , Eng Gee Lim , Lingyun Yu
‹ 上一页 1 2 3 10 下一页 ›