中文
相关论文

相关论文: RoboView-Bias: Benchmarking Visual Bias in Embodie…

200 篇论文

Multimodal Large Language Models (MLLMs) have shown significant advancements, providing a promising future for embodied agents. Existing benchmarks for evaluating MLLMs primarily utilize static images or videos, limiting assessments to…

计算机视觉与模式识别 · 计算机科学 2025-04-14 Zhili Cheng , Yuge Tu , Ran Li , Shiqi Dai , Jinyi Hu , Shengding Hu , Jiahao Li , Yang Shi , Tianyu Yu , Weize Chen , Lei Shi , Maosong Sun

We present EmbodiedMAE, a unified 3D multi-modal representation for robot manipulation. Current approaches suffer from significant domain gaps between training datasets and robot manipulation tasks, while also lacking model architectures…

机器人学 · 计算机科学 2025-05-16 Zibin Dong , Fei Ni , Yifu Yuan , Yinchuan Li , Jianye Hao

Does everyone equally benefit from computer vision systems? Answers to this question become more and more important as computer vision systems are deployed at large scale, and can spark major concerns when they exhibit vast performance…

计算机视觉与模式识别 · 计算机科学 2022-02-16 Priya Goyal , Adriana Romero Soriano , Caner Hazirbas , Levent Sagun , Nicolas Usunier

Current vision-based robotics simulation benchmarks have significantly advanced robotic manipulation research. However, robotics is fundamentally a real-world problem, and evaluation for real-world applications has lagged behind in…

机器人学 · 计算机科学 2025-08-18 Xuning Yang , Clemens Eppner , Jonathan Tremblay , Dieter Fox , Stan Birchfield , Fabio Ramos

Vision-Language Models (VLMs) are increasingly deployed in socially consequential settings, raising concerns about social bias driven by demographic cues. A central challenge in measuring such social bias is attribution under visual…

计算机视觉与模式识别 · 计算机科学 2026-04-21 Haodong Chen , Qiang Huang , Jiaqi Zhao , Qiuping Jiang , Xiaojun Chang , Jun Yu

We introduce RoboEval, a structured evaluation framework and benchmark for robotic manipulation that augments binary success with principled behavioral and outcome metrics. Existing evaluations often collapse performance into outcome…

Recent progress in video-to-video (V2V) translation has enabled realistic resimulation of embodied AI demonstrations, a capability that allows pretrained robot policies to be transferable to new environments without additional data…

计算机视觉与模式识别 · 计算机科学 2026-03-27 George Eskandar , Fengyi Shen , Mohammad Altillawi , Dong Chen , Yang Bai , Liudi Yang , Ziyuan Liu

Machine learning systems are increasingly deployed in high-stakes domains, yet they remain vulnerable to bias systematic disparities that disproportionately impact specific demographic groups. Traditional bias detection methods often depend…

机器学习 · 计算机科学 2025-06-16 Chirudeep Tupakula , Rittika Shamsuddin

Large Vision-Language Models (LVLMs) have been widely adopted in various applications; however, they exhibit significant gender biases. Existing benchmarks primarily evaluate gender bias at the demographic group level, neglecting individual…

计算机视觉与模式识别 · 计算机科学 2024-07-02 Yisong Xiao , Aishan Liu , QianJia Cheng , Zhenfei Yin , Siyuan Liang , Jiapeng Li , Jing Shao , Xianglong Liu , Dacheng Tao

Recent Vision-Language-Action (VLA) models report impressive success rates on standard robotic benchmarks, fueling optimism about general-purpose physical intelligence. However, recent evidence suggests a systematic misalignment between…

机器人学 · 计算机科学 2026-04-21 Haiweng Xu , Sipeng Zheng , Hao Luo , Wanpeng Zhang , Ziheng Xi , Zongqing Lu

In this paper we present an approach and a benchmark for visual reasoning in robotics applications, in particular small object grasping and manipulation. The approach and benchmark are focused on inferring object properties from visual and…

计算机视觉与模式识别 · 计算机科学 2020-04-07 Michal Nazarczuk , Krystian Mikolajczyk

Autonomous robots operating in dynamic environments should identify and report anomalies. Embodying proactive mitigation improves safety and operational continuity. This paper presents a multimodal anomaly detection and mitigation system…

机器人学 · 计算机科学 2025-09-09 Oluwadamilola Sotomi , Devika Kodi , Kiruthiga Chandra Shekar , Aliasghar Arab

Embodied visual tracking is to follow a target object in dynamic 3D environments using an agent's egocentric vision. This is a vital and challenging skill for embodied agents. However, existing methods suffer from inefficient training and…

计算机视觉与模式识别 · 计算机科学 2024-07-23 Fangwei Zhong , Kui Wu , Hai Ci , Churan Wang , Hao Chen

Recently, large vision-language models (LVLMs) have emerged as the preferred tools for judging text-image alignment, yet their robustness along the visual modality remains underexplored. This work is the first study to address a key…

计算与语言 · 计算机科学 2025-11-18 Yerin Hwang , Dongryeol Lee , Kyungmin Min , Taegwan Kang , Yong-il Kim , Kyomin Jung

Vision-Language-Action (VLA) models empower robots to understand and execute tasks described by natural language instructions. However, a key challenge lies in their ability to generalize beyond the specific environments and conditions they…

Embodied cognition argues that intelligence arises from sensorimotor interaction rather than passive observation. It raises an intriguing question: do modern vision-language models (VLMs), trained largely in a disembodied manner, exhibit…

人工智能 · 计算机科学 2025-11-27 Qineng Wang , Wenlong Huang , Yu Zhou , Hang Yin , Tianwei Bao , Jianwen Lyu , Weiyu Liu , Ruohan Zhang , Jiajun Wu , Li Fei-Fei , Manling Li

Current approaches to embodied AI tend to learn policies from expert demonstrations. However, without a mechanism to evaluate the quality of demonstrated actions, they are limited to learning from optimal behaviour, or they risk replicating…

计算与语言 · 计算机科学 2025-10-14 Sabrina McCallum , Amit Parekh , Alessandro Suglia

Robotic manipulation policies often fail to generalize because they must simultaneously learn where to attend, what actions to take, and how to execute them. We argue that high-level reasoning about where and what can be offloaded to…

机器人学 · 计算机科学 2025-09-24 Jesse Zhang , Marius Memmel , Kevin Kim , Dieter Fox , Jesse Thomason , Fabio Ramos , Erdem Bıyık , Abhishek Gupta , Anqi Li

Recent advanced vision-language models(VLMs) have demonstrated strong performance on passive, offline image and video understanding tasks. However, their effectiveness in embodied settings, which require online interaction and active scene…

计算机视觉与模式识别 · 计算机科学 2025-07-15 Mingxian Lin , Wei Huang , Yitang Li , Chengjie Jiang , Kui Wu , Fangwei Zhong , Shengju Qian , Xin Wang , Xiaojuan Qi

As AI agents increasingly operate in open, real-world environments, they require a deep synergy of multimodal perception, tool invocation with multi-hop reasoning, and dynamic interaction with users. However, existing benchmarks fail to…

人工智能 · 计算机科学 2026-05-28 Yunqi Liu , Tong Niu , Zitong Wang , Zhenlong Dai , Yuqi Qing , Weiqiang Wang , Jian Liu