English
Related papers

Related papers: WorldAfford: Affordance Grounding based on Natural…

200 papers

Object placement aims to determine the appropriate placement (\emph{e.g.}, location and size) of a foreground object when placing it on the background image. Most previous works are limited by small-scale labeled dataset, which hinders the…

Computer Vision and Pattern Recognition · Computer Science 2025-04-17 Bingjie Gao , Bo Zhang , Li Niu

Reasoning the human-object interactions (HOI) is essential for deeper scene understanding, while object affordances (or functionalities) are of great importance for human to discover unseen HOIs with novel objects. Inspired by this, we…

Computer Vision and Pattern Recognition · Computer Science 2022-03-30 Zhi Hou , Baosheng Yu , Yu Qiao , Xiaojiang Peng , Dacheng Tao

GUI grounding is a critical capability for vision-language models (VLMs) that enables automated interaction with graphical user interfaces by locating target elements from natural language instructions. However, grounding on GUI screenshots…

Computer Vision and Pattern Recognition · Computer Science 2026-03-19 Siqi Pei , Liang Tang , Tiaonan Duan , Long Chen , Shuxian Li , Kaer Huang , Yanzhe Jing , Yiqiang Yan , Bo Zhang , Chenghao Jiang , Borui Zhang , Jiwen Lu

Assistive robots operating in unstructured environments must understand not only what objects are, but what they can be used for. This requires grounding language-based action queries to objects that both afford the requested function and…

Robotics · Computer Science 2025-12-05 Zhou Chen , Joe Lin , Carson Bulgin , Sathyanarayanan N. Aakur

In the field of visual affordance learning, previous methods mainly used abundant images or videos that delineate human behavior patterns to identify action possibility regions for object manipulation, with a variety of applications in…

Computer Vision and Pattern Recognition · Computer Science 2024-04-09 Zhipeng Zhang , Zhimin Wei , Guolei Sun , Peng Wang , Luc Van Gool

The concept of 3D scene graphs is increasingly recognized as a powerful semantic and hierarchical representation of the environment. Current approaches often address this at a coarse, object-level resolution. In contrast, our goal is to…

Computer Vision and Pattern Recognition · Computer Science 2025-08-12 Dennis Rotondi , Fabio Scaparro , Hermann Blum , Kai O. Arras

Road attributes understanding is extensively researched to support vehicle's action for autonomous driving, whereas current works mainly focus on urban road nets and rely much on traffic signs. This paper generalizes the same issue to the…

Robotics · Computer Science 2019-11-28 Huifang Ma , Yue Wang , Rong Xiong , Sarath Kodagoda , Qianhui Luo

Autonomous navigation based on precise localization has been widely developed in both academic research and practical applications. The high demand for localization accuracy has been essential for safe robot planing and navigation while it…

Robotics · Computer Science 2019-06-07 Huifang Ma , Yue Wang , Li Tang , Sarath Kodagoda , Rong Xiong

Language-guided active sensing is a robotics subtask where a robot with an onboard sensor interacts efficiently with the environment via object manipulation to maximize perceptual information, following given language instructions. These…

Robotics · Computer Science 2024-02-06 Weihan Chen , Hanwen Ren , Ahmed H. Qureshi

As robots begin to cohabit with humans in semi-structured environments, the need arises to understand instructions involving rich variability---for instance, learning to ground symbols in the physical world. Realistically, this task must…

Artificial Intelligence · Computer Science 2017-06-02 Yordan Hristov , Svetlin Penkov , Alex Lascarides , Subramanian Ramamoorthy

Continual learning for video--language understanding is increasingly important as models face non-stationary data, domains, and query styles, yet prevailing solutions blur what should stay stable versus what should adapt, rely on static…

Computer Vision and Pattern Recognition · Computer Science 2025-12-02 Mengzhu Xu , Hanzhi Liu , Ningkang Peng , Qianyu Chen , Canran Xiao

Learning robot manipulation from human videos is appealing due to the scale and diversity of human demonstrations, but transferring such demonstrations to executable robot behavior remains challenging. Prior work either relies on robot data…

Robotics · Computer Science 2026-05-05 Yifan Han , Jianxiang Liu , Haoyu Zhang , Yuqi Gu , Yunhan Guo , Wenzhao Lian

This paper introduces an innovative application of foundation models, enabling Unmanned Ground Vehicles (UGVs) equipped with an RGB-D camera to navigate to designated destinations based on human language instructions. Unlike learning-based…

Robotics · Computer Science 2024-10-15 Chanhoe Ryu , Hyunki Seong , Daegyu Lee , Seongwoo Moon , Sungjae Min , D. Hyunchul Shim

Many everyday objects are difficult to directly grasp (e.g., a flat iPad) or manipulate functionally (e.g., opening the cap of a pen lying on a desk). Such tasks require sequential, asymmetric coordination between two arms, where one arm…

Robotics · Computer Science 2026-03-24 Yan Shen , Feng Jiang , Zichen He , Xiaoqi Li , Yuchen Liu , Zhiyu Li , Ruihai Wu , Hao Dong

Dexterous robotic hands are appealing for their agility and human-like morphology, yet their high degree of freedom makes learning to manipulate challenging. We introduce an approach for learning dexterous grasping. Our key idea is to embed…

Robotics · Computer Science 2021-06-18 Priyanka Mandikal , Kristen Grauman

Recent development in autonomous driving involves high-level computer vision and detailed road scene understanding. Today, most autonomous vehicles are using mediated perception approach for path planning and control, which highly rely on…

Computer Vision and Pattern Recognition · Computer Science 2019-03-22 Chen Sun , Jean M. Uwabeza Vianney , Dongpu Cao

Short-Term object-interaction Anticipation consists of detecting the location of the next-active objects, the noun and verb categories of the interaction, and the time to contact from the observation of egocentric video. This ability is…

Computer Vision and Pattern Recognition · Computer Science 2024-06-06 Lorenzo Mur-Labadia , Ruben Martinez-Cantin , Josechu Guerrero , Giovanni Maria Farinella , Antonino Furnari

Robotic non-destructive disassembly of mating parts remains challenging due to the need for flexible manipulation and the limited visibility of internal structures. This study presents an affordance-guided teleoperation system that enables…

Robotics · Computer Science 2025-08-11 Gen Sako , Takuya Kiyokawa , Kensuke Harada , Tomoki Ishikura , Naoya Miyaji , Genichiro Matsuda

For effective human-robot collaboration, it is crucial for robots to understand requests from users and ask reasonable follow-up questions when there are ambiguities. While comprehending the users' object descriptions in the requests,…

Robotics · Computer Science 2021-07-13 Fethiye Irmak Dogan , Gaspar I. Melsion , Iolanda Leite

The human language is one of the most natural interfaces for humans to interact with robots. This paper presents a robot system that retrieves everyday objects with unconstrained natural language descriptions. A core issue for the system is…

Robotics · Computer Science 2017-07-19 Mohit Shridhar , David Hsu
‹ Prev 1 8 9 10 Next ›