中文
相关论文

相关论文: Egocentric Bias in Vision-Language Models

200 篇论文

Vision-language models (VLMs) have made strong progress on high-level image-text alignment, yet their ability to perceive subtle visual differences remains limited. We study this problem in rendered web interfaces, where localized visual…

计算机视觉与模式识别 · 计算机科学 2026-05-29 Linhao Zhang , Aiwei Liu , Yuan Liu , Xiao Zhou

Driven by the great success of Large Language Models (LLMs) in the 2D image domain, their applications in 3D scene understanding has emerged as a new trend. A key difference between 3D and 2D is that the situation of an egocentric observer…

计算机视觉与模式识别 · 计算机科学 2025-04-01 Zhihao Yuan , Yibo Peng , Jinke Ren , Yinghong Liao , Yatong Han , Chun-Mei Feng , Hengshuang Zhao , Guanbin Li , Shuguang Cui , Zhen Li

Embodied cognition argues that intelligence arises from sensorimotor interaction rather than passive observation. It raises an intriguing question: do modern vision-language models (VLMs), trained largely in a disembodied manner, exhibit…

人工智能 · 计算机科学 2025-11-27 Qineng Wang , Wenlong Huang , Yu Zhou , Hang Yin , Tianwei Bao , Jianwen Lyu , Weiyu Liu , Ruohan Zhang , Jiajun Wu , Li Fei-Fei , Manling Li

Vision-language models (VLMs) achieve strong performance on spatial reasoning benchmarks, yet it remains unclear whether this reflects structured 3D understanding or reliance on statistical shortcuts in natural images. We introduce a…

计算机视觉与模式识别 · 计算机科学 2026-05-29 Cheolhong Min , Jaeyun Jung , Daeun Lee , Hyeonseong Jeon , Yu Su , Jonathan Tremblay , Chan Hee Song , Jaesik Park

People continuously perceive and interact with their surroundings based on underlying intentions that drive their exploration and behaviors. While research in egocentric user and scene understanding has focused primarily on motion and…

计算机视觉与模式识别 · 计算机科学 2025-11-25 Heeseung Yun , Joonil Na , Jaeyeon Kim , Calvin Murdock , Gunhee Kim

In human imitation learning, the imitator typically take the egocentric view as a benchmark, naturally transferring behaviors observed from an exocentric view to their owns, which provides inspiration for researching how robots can more…

计算机视觉与模式识别 · 计算机科学 2025-04-01 Heqian Qiu , Zhaofeng Shi , Lanxiao Wang , Huiyu Xiong , Xiang Li , Hongliang Li

Vision-Language Models (VLMs) demonstrate impressive capabilities across multimodal tasks, yet exhibit systematic spatial reasoning failures, achieving only 49% (CLIP) to 54% (BLIP-2) accuracy on basic directional relationships. For safe…

计算机视觉与模式识别 · 计算机科学 2026-01-21 Muhammad Imran , Yugyung Lee

Recent advancements in Large Language Models (LLMs) and Vision-Language Models (VLMs) have made them powerful tools in embodied navigation, enabling agents to leverage commonsense and spatial reasoning for efficient exploration in…

Recent advances in Vision-Language Models (VLMs) have achieved impressive progress in multimodal mathematical reasoning. Yet, how much visual information truly contributes to reasoning remains unclear. Existing benchmarks report strong…

计算机视觉与模式识别 · 计算机科学 2025-12-01 Yuandong Wang , Yao Cui , Yuxin Zhao , Zhen Yang , Yangfu Zhu , Zhenzhou Shao

Vision-language models (VLMs) have advanced rapidly, but their ability to capture spatial relationships remains a blindspot. Current VLMs are typically built with contrastive language-image pretraining (CLIP) style image encoders. The…

Vision-language models (VLMs) can respond to queries about images in many languages. However, beyond language, culture affects how we see things. For example, individuals from Western cultures focus more on the central figure in an image…

计算与语言 · 计算机科学 2025-03-04 Amith Ananthram , Elias Stengel-Eskin , Mohit Bansal , Kathleen McKeown

Traditional Visual Grounding (VG) predominantly relies on textual descriptions to localize objects, a paradigm that inherently struggles with linguistic ambiguity and often ignores non-verbal deictic cues prevalent in real-world…

计算机视觉与模式识别 · 计算机科学 2026-03-30 Ling Li , Bowen Liu , Zinuo Zhan , Peng Jie , Jianhui Zhong , Kenglun Chang , Zhidong Deng

Vision Language Models (VLMs) have demonstrated significant potential in various downstream tasks, including Image/Video Generation, Visual Question Answering, Multimodal Chatbots, and Video Understanding. However, these models often…

计算机视觉与模式识别 · 计算机科学 2025-03-17 Ahmad Mustafa Anis , Hasnain Ali , Saquib Sarfraz

Recently, many bias detection methods have been proposed to determine the level of bias a large language model captures. However, tests to identify which parts of a large language model are responsible for bias towards specific groups…

计算与语言 · 计算机科学 2025-08-12 Keshav Varadarajan , Tananun Songdechakraiwut

As Vision-Language Models (VLMs) become integral to educational decision-making, ensuring their fairness is paramount. However, current text-centric evaluations neglect the visual modality, leaving an unregulated channel for latent social…

人工智能 · 计算机科学 2026-04-15 Ruijia Li , Mingzi Zhang , Zengyi Yu , Yuang Wei , Bo Jiang

As Vision-Language Models (VLMs) are increasingly deployed as autonomous cognitive cores for embodied assistants, evaluating their privacy awareness in physical environments becomes critical. Unlike digital chatbots, these agents operate in…

密码学与安全 · 计算机科学 2026-05-11 Junran Wang , Xinjie Shen , Zehao Jin , Pan Li

In order to engage in complex social interaction, humans learn at a young age to infer what others see and cannot see from a different point-of-view, and learn to predict others' plans and behaviors. These abilities have been mostly lacking…

机器人学 · 计算机科学 2021-05-12 Boyuan Chen , Yuhang Hu , Robert Kwiatkowski , Shuran Song , Hod Lipson

Vision-language (VL) models, pretrained on colossal image-text datasets, have attained broad VL competence that is difficult to evaluate. A common belief is that a small number of VL skills underlie the variety of VL tests. In this paper,…

计算机视觉与模式识别 · 计算机科学 2024-04-04 Anthony Meng Huat Tiong , Junqi Zhao , Boyang Li , Junnan Li , Steven C. H. Hoi , Caiming Xiong

In this paper, we propose to investigate the problem of out-of-domain visio-linguistic pretraining, where the pretraining data distribution differs from that of downstream data on which the pretrained model will be fine-tuned. Existing…

计算机视觉与模式识别 · 计算机科学 2020-10-05 Shengyu Zhang , Tan Jiang , Tan Wang , Kun Kuang , Zhou Zhao , Jianke Zhu , Jin Yu , Hongxia Yang , Fei Wu

Large Vision-Language Models (LVLMs) evolve rapidly as Large Language Models (LLMs) was equipped with vision modules to create more human-like models. However, we should carefully evaluate their applications in different domains, as they…

计算与语言 · 计算机科学 2024-09-24 Yuhang Xiao , Yudi Lin , Ming-Chang Chiu