中文
相关论文

相关论文: Egocentric Bias in Vision-Language Models

200 篇论文

Vision Language Models (VLMs) have achieved strong performance across diverse video understanding tasks. However, their viewpoint invariant training limits their ability to understand egocentric properties (e.g., human object interactions)…

计算机视觉与模式识别 · 计算机科学 2025-12-17 Dominick Reilly , Manish Kumar Govind , Le Xue , Srijan Das

Visual Language Models (VLMs) show remarkable performance in visual reasoning tasks, successfully tackling college-level challenges that require high-level understanding of images. However, some recent reports of VLMs struggling to reason…

计算机视觉与模式识别 · 计算机科学 2025-04-17 Gene Tangtartharakul , Katherine R. Storrs

This paper addresses the daily challenges encountered by visually impaired individuals, such as limited access to information, navigation difficulties, and barriers to social interaction. To alleviate these challenges, we introduce a novel…

计算机视觉与模式识别 · 计算机科学 2024-05-31 Inpyo Song , Minjun Joo , Joonhyung Kwon , Jangwon Lee

We present SpinBench, a cognitively grounded diagnostic benchmark for evaluating spatial reasoning in vision language models (VLMs). SpinBench is designed around the core challenge of spatial reasoning: perspective taking, the ability to…

计算机视觉与模式识别 · 计算机科学 2026-03-03 Yuyou Zhang , Radu Corcodel , Chiori Hori , Anoop Cherian , Ding Zhao

Transferring and integrating knowledge across first-person (egocentric) and third-person (exocentric) viewpoints is intrinsic to human intelligence, enabling humans to learn from others and convey insights from their own experiences.…

计算机视觉与模式识别 · 计算机科学 2025-07-25 Yuping He , Yifei Huang , Guo Chen , Baoqi Pei , Jilan Xu , Tong Lu , Jiangmiao Pang

The ability to read, understand and find important information from written text is a critical skill in our daily lives for our independence, comfort and safety. However, a significant part of our society is affected by partial vision…

计算机视觉与模式识别 · 计算机科学 2024-04-16 Wiktor Mucha , Florin Cuconasu , Naome A. Etori , Valia Kalokyri , Giovanni Trappolini

While Vision-Language Models (VLMs) have advanced highlevel reasoning in autonomous driving, their ability to ground this reasoning in the underlying physics of ego-motion remains poorly understood. We introduce EgoDyn-Bench, a diagnostic…

计算机视觉与模式识别 · 计算机科学 2026-04-28 Finn Rasmus Schäfer , Yuan Gao , Dingrui Wang , Thomas Stauner , Stephan Günnemann , Mattia Piccinini , Sebastian Schmidt , Johannes Betz

Can VLMs predict how each camera move changes the view, and plan many such moves ahead? We call this capability view planning, requiring (1)understanding how a single action transforms the view, and (2)composing many such transformations…

We present a solution to egocentric 3D body pose estimation from monocular images captured from downward looking fish-eye cameras installed on the rim of a head mounted VR device. This unusual viewpoint leads to images with unique visual…

计算机视觉与模式识别 · 计算机科学 2020-11-04 Denis Tome , Thiemo Alldieck , Patrick Peluse , Gerard Pons-Moll , Lourdes Agapito , Hernan Badino , Fernando De la Torre

Recent advancements in Multimodal Large Language Models (MLLMs) have demonstrated impressive performance on standard visual reasoning benchmarks. However, there is growing concern that these models rely excessively on linguistic shortcuts…

计算与语言 · 计算机科学 2026-01-09 Ziteng Wang , Yujie He , Guanliang Li , Siqi Yang , Jiaqi Xiong , Songxiang Liu

Recent advances in multimodal large language models (MLLMs) offer a promising approach for natural language-based scene change queries in virtual reality (VR). Prior work on applying MLLMs for object state understanding has focused on…

计算机视觉与模式识别 · 计算机科学 2026-03-10 Shiyi Ding , Shaoen Wu , Ying Chen

Vision-language models (VLMs) have gained widespread adoption in both industry and academia. In this study, we propose a unified framework for systematically evaluating gender, race, and age biases in VLMs with respect to professions. Our…

计算机视觉与模式识别 · 计算机科学 2024-06-18 Ashutosh Sathe , Prachi Jain , Sunayana Sitaram

Vision--language models (VLMs) achieve strong performance on many multimodal benchmarks but remain brittle on spatial reasoning tasks that require aligning abstract overhead representations with egocentric views. We introduce m2sv, a…

计算机视觉与模式识别 · 计算机科学 2026-01-28 Yosub Shin , Michael Buriek , Igor Molybog

Understanding social interactions from egocentric views is crucial for many applications, ranging from assistive robotics to AR/VR. Key to reasoning about interactions is to understand the body pose and motion of the interaction partner…

计算机视觉与模式识别 · 计算机科学 2022-08-17 Siwei Zhang , Qianli Ma , Yan Zhang , Zhiyin Qian , Taein Kwon , Marc Pollefeys , Federica Bogo , Siyu Tang

Knowing others' intentions and taking others' perspectives are two core components of human intelligence that are considered to be instantiations of theory-of-mind. Infiltrating machines with these abilities is an important step towards…

人工智能 · 计算机科学 2025-08-13 Qingying Gao , Yijiang Li , Haiyun Lyu , Haoran Sun , Dezhi Luo , Hokin Deng

As Vision Language Models (VLMs) gain widespread use, their fairness remains under-explored. In this paper, we analyze demographic biases across five models and six datasets. We find that portrait datasets like UTKFace and CelebA are the…

计算与语言 · 计算机科学 2025-04-01 Kuleen Sasse , Shan Chen , Jackson Pond , Danielle Bitterman , John Osborne

Vision-Language Models (VLMs) are increasingly deployed in socially consequential settings, raising concerns about social bias driven by demographic cues. A central challenge in measuring such social bias is attribution under visual…

计算机视觉与模式识别 · 计算机科学 2026-04-21 Haodong Chen , Qiang Huang , Jiaqi Zhao , Qiuping Jiang , Xiaojun Chang , Jun Yu

The ability to forecast human-environment collisions from egocentric observations is vital to enable collision avoidance in applications such as VR, AR, and wearable assistive robotics. In this work, we introduce the challenging problem of…

计算机视觉与模式识别 · 计算机科学 2023-03-28 Boxiao Pan , Bokui Shen , Davis Rempe , Despoina Paschalidou , Kaichun Mo , Yanchao Yang , Leonidas J. Guibas

Vision language models (VLMs) are designed to extract relevant visuospatial information from images. Some research suggests that VLMs can exhibit humanlike scene understanding, while other investigations reveal difficulties in their ability…

Many datasets represent a combination of different ways of looking at the same data that lead to different generalizations. For example, a corpus with examples generated by different people may be mixtures of many perspectives and can be…

机器学习 · 计算机科学 2022-01-25 Karthik Dinakar , Henry Lieberman