English
Related papers

Related papers: Through Their Eyes: Fixation-aligned Tuning for Pe…

200 papers

This paper proposes the User Viewing Flow Modeling (SINGLE) method for the article recommendation task, which models the user constant preference and instant interest from user-clicked articles. Specifically, we first employ a user constant…

Information Retrieval · Computer Science 2024-03-08 Zhenghao Liu , Zulong Chen , Moufeng Zhang , Shaoyang Duan , Hong Wen , Liangyue Li , Nan Li , Yu Gu , Ge Yu

As wearable devices like smart glasses integrate Large Multimodal Models (LMMs) into the continuous first-person visual streams of individual users, the evolution of these models into true personal assistants hinges on visual…

Computer Vision and Pattern Recognition · Computer Science 2026-05-12 Zihui Xue , Ami Baid , Sangho Kim , Mi Luo , Kristen Grauman

Vision-language models (VLMs) excel in various multimodal tasks but frequently suffer from poor calibration, resulting in misalignment between their verbalized confidence and response correctness. This miscalibration undermines user trust,…

Computer Vision and Pattern Recognition · Computer Science 2025-04-22 Yunpu Zhao , Rui Zhang , Junbin Xiao , Ruibo Hou , Jiaming Guo , Zihao Zhang , Yifan Hao , Yunji Chen

Vision-language models (VLMs) have rapidly evolved into general-purpose multimodal reasoners with strong zero-shot generalization. In this context, VLMs could greatly benefit the analysis of human gaze and attention, a central task in human…

Computer Vision and Pattern Recognition · Computer Science 2026-05-25 Hengfei Wang , Anshul Gupta , Pierre Vuillecard , Jean-Marc Odobez

Preference-based reinforcement learning (RL) offers a promising approach for aligning policies with human intent but is often constrained by the high cost of human feedback. In this work, we introduce PrefVLM, a framework that integrates…

Machine Learning · Computer Science 2025-02-04 Udita Ghosh , Dripta S. Raychaudhuri , Jiachen Li , Konstantinos Karydis , Amit Roy-Chowdhury

Recent advancements in multimodal large language models (MLLMs) have demonstrated significant progress; however, these models exhibit a notable limitation, which we refer to as "face blindness". Specifically, they can engage in general…

Computer Vision and Pattern Recognition · Computer Science 2024-10-10 Renjie Pi , Jianshu Zhang , Tianyang Han , Jipeng Zhang , Rui Pan , Tong Zhang

Vision-language models (VLMs) can describe urban scenes in rich detail, yet consistently fail to produce reliable human preference labels in domain-specific tasks such as safety assessment and aesthetic evaluation. The standard fix,…

Computer Vision and Pattern Recognition · Computer Science 2026-03-12 Yecheng Zhang , Rong Zhao , Zhizhou Sha , Yong Li , Lei Wang , Ce Hou , Wen Ji , Hao Huang , Yunshan Wan , Jian Yu , Junhao Xia , Yuru Zhang , Chunlei Shi

The goal of vision-language modeling is to allow models to tie language understanding with visual inputs. The aim of this paper is to evaluate and align the Visual Language Model (VLM) called Multimodal Augmentation of Generative Models…

Computer Vision and Pattern Recognition · Computer Science 2022-10-26 Jean-Charles Layoun , Alexis Roger , Irina Rish

Effective recommender systems demand dynamic user understanding, especially in complex, evolving environments. Traditional user profiling often fails to capture the nuanced, temporal contextual factors of user preferences, such as transient…

Information Retrieval · Computer Science 2025-08-13 Milad Sabouri , Masoud Mansoury , Kun Lin , Bamshad Mobasher

Large Language Models (LLMs) have achieved remarkable success, where instruction tuning is the critical step in aligning LLMs with user intentions. In this work, we investigate how the instruction tuning adjusts pre-trained models with a…

Computation and Language · Computer Science 2024-04-05 Xuansheng Wu , Wenlin Yao , Jianshu Chen , Xiaoman Pan , Xiaoyang Wang , Ninghao Liu , Dong Yu

As large language models (LLMs) become integral to intelligent user interfaces (IUIs), their role as decision-making agents raises critical concerns about alignment. Although extensive research has addressed issues such as factuality, bias,…

Artificial Intelligence · Computer Science 2025-04-23 Anna Karnysheva , Christian Drescher , Dietrich Klakow

Accurately modeling user preferences is crucial for improving the performance of content-based recommender systems. Existing approaches often rely on simplistic user profiling methods, such as averaging or concatenating item embeddings,…

Information Retrieval · Computer Science 2025-08-13 Milad Sabouri , Masoud Mansoury , Kun Lin , Bamshad Mobasher

Personalized driving refers to an autonomous vehicle's ability to adapt its driving behavior or control strategies to match individual users' preferences and driving styles while maintaining safety and comfort standards. However, existing…

Recommender systems have become integral to our digital experiences, from online shopping to streaming platforms. Still, the rationale behind their suggestions often remains opaque to users. While some systems employ a graph-based approach,…

The advent of immersive Virtual Reality applications has transformed various domains, yet their integration with advanced artificial intelligence technologies like Visual Language Models remains underexplored. This study introduces a…

Robotics · Computer Science 2024-08-06 Mikhail Konenkov , Artem Lykov , Daria Trinitatova , Dzmitry Tsetserukou

Recently, much effort has been devoted to modeling users' multi-interests based on their behaviors or auxiliary signals. However, existing methods often rely on heuristic assumptions, e.g., co-occurring items indicate the same interest of…

Information Retrieval · Computer Science 2025-07-18 Ziyan Wang , Yingpeng Du , Zhu Sun , Jieyi Bi , Haoyan Chua , Tianjun Wei , Jie Zhang

Vision-Language-Action (VLA) models have demonstrated strong performance across a wide range of robotic manipulation tasks. Despite the success, extending large pretrained Vision-Language Models (VLMs) to the action space can induce…

Computer Vision and Pattern Recognition · Computer Science 2026-02-06 Yiye Chen , Yanan Jian , Xiaoyi Dong , Shuxin Cao , Jing Wu , Patricio Vela , Benjamin E. Lundell , Dongdong Chen

Vision-Language Models (VLMs) offer the ability to generate high-level, interpretable descriptions of complex activities from images and videos, making them valuable for situational awareness (SA) applications. In such settings, the focus…

Computer Vision and Pattern Recognition · Computer Science 2026-01-19 Pavana Pradeep , Krishna Kant , Suya Yu

As large language models (LLMs) continue to advance, aligning these models with human preferences has emerged as a critical challenge. Traditional alignment methods, relying on human or LLM annotated datasets, are limited by their…

With the large language model showing human-like logical reasoning and understanding ability, whether agents based on the large language model can simulate the interaction behavior of real users, so as to build a reliable virtual…

Information Retrieval · Computer Science 2024-03-05 Chenwei Zhang , Wenran Lu , Chunhe Ni , Hongbo Wang , Jiang Wu
‹ Prev 1 4 5 6 7 8 10 Next ›