中文
相关论文

相关论文: AIM-Bench: Benchmarking and Improving Affective Im…

200 篇论文

While real-world applications increasingly demand intricate scene manipulation, existing instruction-guided image editing benchmarks often oversimplify task complexity and lack comprehensive, fine-grained instructions. To bridge this gap,…

Text-guided human pose editing has gained significant traction in AIGC applications. However,it remains plagued by structural anomalies and generative artifacts. Existing evaluation metrics often isolate authenticity detection from quality…

计算机视觉与模式识别 · 计算机科学 2026-01-21 Ningyu Sun , Zhaolin Cai , Zitong Xu , Peihang Chen , Huiyu Duan , Yichao Yan , Xiongkuo Min , Xiaokang Yang

In cinematography, visual attributes such as color grading, contrast, and brightness are manipulated to reinforce the emotional narrative of a scene. However, conventional Image Signal Processors (ISPs) prioritize scene fidelity,…

图像与视频处理 · 电气工程与系统科学 2026-05-06 Dor Barber , Rony Zatzarinni , Hava Matichin , Noam Levy

Multimodal Large Language Models (MLLMs) are increasingly applied in Personalized Image Aesthetic Assessment (PIAA) as a scalable alternative to expert evaluations. However, their predictions may reflect subtle biases influenced by…

计算与语言 · 计算机科学 2025-09-16 Kun Li , Lai-Man Po , Hongzheng Yang , Xuyuan Xu , Kangcheng Liu , Yuzhi Zhao

The rapid advancement of educational applications, artistic creation, and AI-generated content (AIGC) technologies has substantially increased practical requirements for comprehensive Image Aesthetics Assessment (IAA), particularly…

计算机视觉与模式识别 · 计算机科学 2025-08-12 Shuo Cao , Nan Ma , Jiayang Li , Xiaohui Li , Lihao Shao , Kaiwen Zhu , Yu Zhou , Yuandong Pu , Jiarui Wu , Jiaquan Wang , Bo Qu , Wenhai Wang , Yu Qiao , Dajuin Yao , Yihao Liu

Images can convey rich semantics and induce various emotions in viewers. Recently, with the rapid advancement of emotional intelligence and the explosive growth of visual data, extensive research efforts have been dedicated to affective…

计算机视觉与模式识别 · 计算机科学 2021-07-01 Sicheng Zhao , Xingxu Yao , Jufeng Yang , Guoli Jia , Guiguang Ding , Tat-Seng Chua , Björn W. Schuller , Kurt Keutzer

Vision-Language Models (VLMs) have demonstrated strong capabilities in perception, yet holistic Affective Image Content Analysis (AICA), which integrates perception, reasoning, and generation into a unified framework, remains underexplored.…

计算机视觉与模式识别 · 计算机科学 2026-04-08 Dong She , Xianrong Yao , Liqun Chen , Jinghe Yu , Yang Gao , Zhanpeng Jin

Simultaneously optimizing multiple, frequently conflicting, molecular properties is a key bottleneck in the development of novel therapeutics. Although a promising approach, the efficacy of multi-task learning is often compromised by…

机器学习 · 计算机科学 2025-10-01 Mason Minot , Gisbert Schneider

Recent breakthroughs in large multimodal models (LMMs) have significantly advanced both text-to-image (T2I) generation and image-to-text (I2T) interpretation. However, many generated images still suffer from issues related to perceptual…

计算机视觉与模式识别 · 计算机科学 2025-04-14 Jiarui Wang , Huiyu Duan , Yu Zhao , Juntong Wang , Guangtao Zhai , Xiongkuo Min

Understanding the multi-dimensional attributes and intensity nuances of image-evoked emotions is pivotal for advancing machine empathy and empowering diverse human-computer interaction applications. However, existing models are still…

计算机视觉与模式识别 · 计算机科学 2026-05-26 Lancheng Gao , Ziheng Jia , Zixuan Xing , Wei Sun , Huiyu Duan , Guangtao Zhai , Xiongkuo Min

Recent advances in multi-modal generative models have enabled significant progress in instruction-based image editing. However, while these models produce visually plausible outputs, their capacity for knowledge-based reasoning editing…

计算机视觉与模式识别 · 计算机科学 2025-05-23 Yongliang Wu , Zonghui Li , Xinting Hu , Xinyu Ye , Xianfang Zeng , Gang Yu , Wenbo Zhu , Bernt Schiele , Ming-Hsuan Yang , Xu Yang

Recent advances in image manipulation have enabled highly photorealistic content generation, but also lowered the barrier to arbitrary editing, raising concerns about multimedia authenticity and security. Existing Image Manipulation…

计算机视觉与模式识别 · 计算机科学 2026-05-18 Haozhen Yan , Yan Hong , Jiahui Zhan , Suning Lang , Yikun Ji , Huijia Zhu , Jun Lan , Jianfu Zhang

This paper introduces a multi-label visual emotion analysis benchmark dataset for comprehensively evaluating the ability of multimodal large language models (MLLMs) to predict the emotions evoked by images. Recent user studies report an…

计算机视觉与模式识别 · 计算机科学 2026-05-15 Tianwei Chen , Takuya Furusawa , Yuki Hirakawa , Ryotaro Shimizu , Mo Fan , Takashi Wada

The ability to represent emotion plays a significant role in human cognition and social interaction, yet the high-dimensional geometry of this affective space and its neural underpinnings remain debated. A key challenge, the…

人机交互 · 计算机科学 2025-09-30 Changde Du , Yizhuo Lu , Zhongyu Huang , Yi Sun , Zisen Zhou , Shaozheng Qin , Huiguang He

Modern automotive infotainment systems necessitate intelligent and adaptive solutions to manage frequent User Interface (UI) updates and diverse design variations. This work introduces a vision-language framework to facilitate the…

计算机视觉与模式识别 · 计算机科学 2025-08-06 Benjamin Raphael Ernhofer , Daniil Prokhorov , Jannica Langner , Dominik Bollmann

Large Multi-modality Models (LMMs) have made significant progress in visual understanding and generation, but they still face challenges in General Visual Editing, particularly in following complex instructions, preserving appearance…

计算机视觉与模式识别 · 计算机科学 2025-05-28 Xiangyu Zhao , Peiyuan Zhang , Kexian Tang , Xiaorong Zhu , Hao Li , Wenhao Chai , Zicheng Zhang , Renqiu Xia , Guangtao Zhai , Junchi Yan , Hua Yang , Xue Yang , Haodong Duan

Psychological research results have confirmed that people can have different emotional reactions to different visual stimuli. Several papers have been published on the problem of visual emotion analysis. In particular, attempts have been…

人工智能 · 计算机科学 2016-05-10 Quanzeng You , Jiebo Luo , Hailin Jin , Jianchao Yang

Instruction-guided image editing has seen remarkable progress with models like FLUX.2 and Qwen-Image-Edit, yet they still struggle with complex scenarios with multiple similar instances each requiring individual edits. We observe that…

计算机视觉与模式识别 · 计算机科学 2026-04-08 Ziqian Liu , Stephan Alaniz

Recent multimodal image generators such as GPT-4o, Gemini 2.0 Flash, and Gemini 2.5 Pro excel at following complex instructions, editing images and maintaining concept consistency. However, they are still evaluated by disjoint toolkits:…

计算机视觉与模式识别 · 计算机科学 2025-05-29 Hang Hua , Ziyun Zeng , Yizhi Song , Yunlong Tang , Liu He , Daniel Aliaga , Wei Xiong , Jiebo Luo

In personalized machine learning, the aim of personalization is to train a model that caters to a specific individual or group of individuals by optimizing one or more performance metrics and adhering to specific constraints. In this paper,…

人机交互 · 计算机科学 2026-01-15 Jialin Li , Maha Elgarf , Alia Waleed , Hanan Salam