中文
相关论文

相关论文: Ges3ViG: Incorporating Pointing Gestures into Lang…

200 篇论文

Referring Expression Comprehension (REC) is a popular multimodal task that aims to accurately detect target objects within a single image based on a given textual expression. However, due to the limitations of earlier models, traditional…

机器学习 · 计算机科学 2025-08-21 Guanghao Jin , Jingpei Wu , Tianpei Guo , Yiyi Niu , Weidong Zhou , Guoyang Liu

Visual grounding associates textual descriptions with objects in an image. Conventional methods target third-person image inputs and named object queries. In applications such as AI assistants, the perspective shifts -- inputs are…

计算机视觉与模式识别 · 计算机科学 2025-04-21 Pengzhan Sun , Junbin Xiao , Tze Ho Elden Tse , Yicong Li , Arjun Akula , Angela Yao

Understanding 3D scenes goes beyond simply recognizing objects; it requires reasoning about the spatial and semantic relationships between them. Current 3D scene-language models often struggle with this relational understanding,…

计算机视觉与模式识别 · 计算机科学 2025-12-09 Jintang Xue , Ganning Zhao , Jie-En Yao , Hong-En Chen , Yue Hu , Meida Chen , Suya You , C. -C. Jay Kuo

Open-Vocabulary 3D object affordance grounding aims to anticipate ``action possibilities'' regions on 3D objects with arbitrary instructions, which is crucial for robots to generically perceive real scenarios and respond to operational…

计算机视觉与模式识别 · 计算机科学 2025-04-01 Yawen Shao , Wei Zhai , Yuhang Yang , Hongchen Luo , Yang Cao , Zheng-Jun Zha

Recent advancements in 3D reconstruction and neural rendering have enhanced the creation of high-quality digital assets, yet existing methods struggle to generalize across varying object shapes, textures, and occlusions. While Next Best…

机器人学 · 计算机科学 2024-09-25 Zhenghao Qi , Shenghai Yuan , Fen Liu , Haozhi Cao , Tianchen Deng , Jianfei Yang , Lihua Xie

Most existing 3D referring expression segmentation (3DRES) methods rely on dense, high-quality point clouds, while real-world agents such as robots and mobile phones operate with only a few sparse RGB views and strict latency constraints.…

计算机视觉与模式识别 · 计算机科学 2026-04-01 Changli Wu , Haodong Wang , Jiayi Ji , Yutian Yao , Chunsai Du , Jihua Kang , Yanwei Fu , Liujuan Cao

3D visual grounding aims to locate the referred target object in 3D point cloud scenes according to a free-form language description. Previous methods mostly follow a two-stage paradigm, i.e., language-irrelevant detection and cross-modal…

计算机视觉与模式识别 · 计算机科学 2023-10-13 Junyu Luo , Jiahui Fu , Xianghao Kong , Chen Gao , Haibing Ren , Hao Shen , Huaxia Xia , Si Liu

3D Visual Grounding (3DVG) seeks to locate target objects in 3D scenes using natural language descriptions, enabling downstream applications such as augmented reality and robotics. Existing approaches typically rely on labeled 3D data and…

计算机视觉与模式识别 · 计算机科学 2025-05-29 Rong Li , Shijie Li , Lingdong Kong , Xulei Yang , Junwei Liang

3D visual grounding is the task of localizing the object in a 3D scene which is referred by a description in natural language. With a wide range of applications ranging from autonomous indoor robotics to AR/VR, the task has recently risen…

计算机视觉与模式识别 · 计算机科学 2024-07-17 Ozan Unal , Christos Sakaridis , Suman Saha , Luc Van Gool

Embodied reference understanding is crucial for intelligent agents to predict referents based on human intention through gesture signals and language descriptions. This paper introduces the Attention-Dynamic DINO, a novel framework designed…

计算机视觉与模式识别 · 计算机科学 2024-11-14 Hao Guo , Wei Fan , Baichun Wei , Jianfei Zhu , Jin Tian , Chunzhi Yi , Feng Jiang

Visual search in 3D environments requires embodied agents to actively explore their surroundings and acquire task-relevant evidence. However, existing visual search and embodied AI benchmarks, including EQA, typically rely on static…

计算机视觉与模式识别 · 计算机科学 2026-04-24 Koya Sakamoto , Taiki Miyanishi , Daichi Azuma , Shuhei Kurita , Shu Morikuni , Naoya Chiba , Motoaki Kawanabe , Yusuke Iwasawa , Yutaka Matsuo

The 3D visual grounding task has been explored with visual and language streams comprehending referential language to identify target objects in 3D scenes. However, most existing methods devote the visual stream to capturing the 3D visual…

计算机视觉与模式识别 · 计算机科学 2022-11-28 Eslam Mohamed Bakr , Yasmeen Alsaedy , Mohamed Elhoseiny

Visual-language grounding aims to establish semantic correspondences between natural language and visual entities, enabling models to accurately identify and localize target objects based on textual instructions. Existing VLG approaches…

计算机视觉与模式识别 · 计算机科学 2026-03-17 Linfei Li , Lin Zhang , Ying Shen

The rapid development of Large Multimodal Models (LMMs) has led to remarkable progress in 2D visual understanding; however, extending these capabilities to 3D scene understanding remains a significant challenge. Existing approaches…

计算机视觉与模式识别 · 计算机科学 2025-09-05 Hongpei Zheng , Lintao Xiang , Qijun Yang , Qian Lin , Hujun Yin

This paper presents INGRESS, a robot system that follows human natural language instructions to pick and place everyday objects. The core issue here is the grounding of referring expressions: infer objects and their relationships from input…

机器人学 · 计算机科学 2018-06-12 Mohit Shridhar , David Hsu

Reconstructing and semantically interpreting 3D scenes from sparse 2D views remains a fundamental challenge in computer vision. Conventional methods often decouple semantic understanding from reconstruction or necessitate costly per-scene…

计算机视觉与模式识别 · 计算机科学 2026-03-25 Xiangyu Sun , Haoyi Jiang , Liu Liu , Seungtae Nam , Gyeongjin Kang , Xinjie Wang , Wei Sui , Zhizhong Su , Wenyu Liu , Xinggang Wang , Eunbyung Park

As robots enter human workspaces, there is a crucial need for them to comprehend embodied human instructions, enabling intuitive and fluent human-robot interaction (HRI). However, accurate comprehension is challenging due to a lack of…

机器人学 · 计算机科学 2025-12-09 Md Mofijul Islam , Alexi Gladstone , Sujan Sarker , Ganesh Nanduru , Md Fahim , Keyan Du , Aman Chadha , Tariq Iqbal

In real-life scenarios, humans seek out objects in the 3D world to fulfill their daily needs or intentions. This inspires us to introduce 3D intention grounding, a new task in 3D object detection employing RGB-D, based on human intention,…

计算机视觉与模式识别 · 计算机科学 2025-02-19 Weitai Kang , Mengxue Qu , Jyoti Kini , Yunchao Wei , Mubarak Shah , Yan Yan

Recently, 3D vision-and-language tasks have attracted increasing research interest. Compared to other vision-and-language tasks, the 3D visual question answering (VQA) task is less exploited and is more susceptible to language priors and…

计算机视觉与模式识别 · 计算机科学 2022-09-27 Lichen Zhao , Daigang Cai , Jing Zhang , Lu Sheng , Dong Xu , Rui Zheng , Yinjie Zhao , Lipeng Wang , Xibo Fan

Segment matching is an important intermediate task in computer vision that establishes correspondences between semantically or geometrically coherent regions across images. Unlike keypoint matching, which focuses on localized features,…

计算机视觉与模式识别 · 计算机科学 2025-10-27 Rohit Jayanti , Swayam Agrawal , Vansh Garg , Siddharth Tourani , Muhammad Haris Khan , Sourav Garg , Madhava Krishna