English
Related papers

Related papers: Enhancing Zero-shot Personalized Image Aesthetics …

200 papers

Hateful meme detection is a challenging multimodal task that requires comprehension of both vision and language, as well as cross-modal interactions. Recent studies have tried to fine-tune pre-trained vision-language models (PVLMs) for this…

Computer Vision and Pattern Recognition · Computer Science 2023-08-17 Rui Cao , Ming Shan Hee , Adriel Kuek , Wen-Haw Chong , Roy Ka-Wei Lee , Jing Jiang

Multimodal large language models (MLLMs) achieve strong performance on vision-language tasks, yet their visual processing is opaque. Most black-box evaluations measure task accuracy, but reveal little about underlying mechanisms. Drawing on…

Computer Vision and Pattern Recognition · Computer Science 2025-10-23 John Burden , Jonathan Prunty , Ben Slater , Matthieu Tehenan , Greg Davis , Lucy Cheke

Recently, Multimodal Large Language Models (MLLMs) have demonstrated exceptional capabilities in visual understanding and reasoning across various vision-language tasks. However, we found that MLLMs cannot process effectively from…

Computer Vision and Pattern Recognition · Computer Science 2025-09-26 Bangyan Li , Wenxuan Huang , Zhenkun Gao , Yeqiang Wang , Yunhang Shen , Jingzhong Lin , Ling You , Yuxiang Shen , Shaohui Lin , Wanli Ouyang , Yuling Sun

Image quality assessment (IQA) focuses on the perceptual visual quality of images, playing a crucial role in downstream tasks such as image reconstruction, compression, and generation. The rapid advancement of multi-modal large language…

Computer Vision and Pattern Recognition · Computer Science 2025-05-26 Weiqi Li , Xuanyu Zhang , Shijie Zhao , Yabin Zhang , Junlin Li , Li Zhang , Jian Zhang

Zero-shot anomaly detection (ZSAD) recognizes and localizes anomalies in previously unseen objects by establishing feature mapping between textual prompts and inspection images, demonstrating excellent research value in flexible industrial…

Computer Vision and Pattern Recognition · Computer Science 2026-04-02 Huilin Deng , Hongchen Luo , Wei Zhai , Yang Cao , Yu Kang

The growing reliance on artificial intelligence (AI) in customer support has significantly improved operational efficiency and user experience. However, traditional machine learning (ML) approaches, which require extensive local training on…

Machine Learning · Computer Science 2025-01-03 Anant Prakash Awasthi , Girdhar Gopal Agarwal , Chandraketu Singh , Rakshit Varma , Sanchit Sharma

Vision-Language Models (VLMs) like CLIP have demonstrated remarkable generalization in zero- and few-shot settings, but adapting them efficiently to decentralized, heterogeneous data remains a challenge. While prompt tuning has emerged as a…

Computer Vision and Pattern Recognition · Computer Science 2026-03-02 Sajjad Ghiasvand , Mahnoosh Alizadeh , Ramtin Pedarsani

Recently, few-shot learning (FSL) has become a popular task that aims to recognize new classes from only a few labeled examples and has been widely applied in fields such as natural science, remote sensing, and medical images. However, most…

Computer Vision and Pattern Recognition · Computer Science 2026-02-12 Liwen Wu , Wei Wang , Lei Zhao , Zhan Gao , Qika Lin , Shaowen Yao , Zuozhu Liu , Bin Pu

Robotic search of people in human-centered environments, including healthcare settings, is challenging as autonomous robots need to locate people without complete or any prior knowledge of their schedules, plans or locations. Furthermore,…

Robotics · Computer Science 2024-12-03 Angus Fung , Aaron Hao Tan , Haitong Wang , Beno Benhabib , Goldie Nejat

Typical LLM responses tend to follow a default style, even though users often have distinct preferences regarding tone, verbosity, and formality that they do not explicitly state in their prompts. Evaluating whether personalization methods…

Computation and Language · Computer Science 2026-05-21 Philipp Spohn , Leander Girrbach , Zeynep Akata

Reinforcement Learning (RL) has empowered Multimodal Large Language Models (MLLMs) to achieve superior human preference alignment in Image Quality Assessment (IQA). However, existing RL-based IQA models typically rely on coarse-grained…

Image and Video Processing · Electrical Eng. & Systems 2026-05-11 Xiang Li , Xueheng Li , Yu Wang , Xuanhua He , Zhangchi Hu , Weiwei Yu , Chengjun Xie

Recent advancements in zero-shot commonsense reasoning have empowered Pre-trained Language Models (PLMs) to acquire extensive commonsense knowledge without requiring task-specific fine-tuning. Despite this progress, these models frequently…

Artificial Intelligence · Computer Science 2026-03-06 Hyuntae Park , Yeachan Kim , SangKeun Lee

Personalized Large Language Models (LLMs) have been shown to be an effective way to create more engaging and enjoyable user-AI interactions. While previous studies have explored using prompts to elicit specific personality traits in LLMs,…

Computation and Language · Computer Science 2025-11-26 Shi-Wei Dai , Yan-Wei Shie , Tsung-Huan Yang , Lun-Wei Ku , Yung-Hui Li

The Multimodal Large Language Models (MLLMs) have activated the capabilitiesof Large Language Models (LLMs) in solving visual-language tasks by integratingvisual information. The prevailing approach in existing MLLMs involvesemploying an…

Computer Vision and Pattern Recognition · Computer Science 2024-10-31 Tianxiang Wu , Minxin Nie , Ziqiang Cao

Large Multimodal Models (LMMs) have recently enabled considerable advances in the realm of image and video quality assessment, but this progress has yet to be fully explored in the domain of 3D assets. We are interested in using these…

Computer Vision and Pattern Recognition · Computer Science 2025-10-10 Shashank Gupta , Gregoire Phillips , Alan C. Bovik

The rapidly developing field of large multimodal models (LMMs) has led to the emergence of diverse models with remarkable capabilities. However, existing benchmarks fail to comprehensively, objectively and accurately evaluate whether LMMs…

We introduce a Depicted image Quality Assessment method (DepictQA), overcoming the constraints of traditional score-based methods. DepictQA allows for detailed, language-based, human-like evaluation of image quality by leveraging…

Computer Vision and Pattern Recognition · Computer Science 2024-07-16 Zhiyuan You , Zheyuan Li , Jinjin Gu , Zhenfei Yin , Tianfan Xue , Chao Dong

Traditional supervised methods for detecting AI-generated images depend on large, curated datasets for training and fail to generalize to novel, out-of-domain image generators. As an alternative, we explore pre-trained Vision-Language…

Machine Learning · Computer Science 2026-01-27 Zoher Kachwala , Danishjeet Singh , Danielle Yang , Filippo Menczer

Text-based person re-identification (TBPReID) aims to retrieve person images represented by a given textual query. In this task, how to effectively align images and texts globally and locally is a crucial challenge. Recent works have…

Computer Vision and Pattern Recognition · Computer Science 2023-09-12 Takuro Fujii , Shuhei Tarashima

Visual preference alignment involves training Large Vision-Language Models (LVLMs) to predict human preferences between visual inputs. This is typically achieved by using labeled datasets of chosen/rejected pairs and employing optimization…

Computer Vision and Pattern Recognition · Computer Science 2024-10-24 Ziyu Liu , Yuhang Zang , Xiaoyi Dong , Pan Zhang , Yuhang Cao , Haodong Duan , Conghui He , Yuanjun Xiong , Dahua Lin , Jiaqi Wang
‹ Prev 1 4 5 6 7 8 10 Next ›